ARI Bench · RSI Watch
Five weeks, six people, every number on the record about recursive self-improvement: who said what, what their role lets them know, and the one thing all of them agree on.
On September 8, pretraining researcher Jacob Coxon resigned from Anthropic and wrote on X that "the people building AI earnestly believe that it could kill us all by the end of the decade." Within two days the post had passed 150 million views. Anthropic's head of alignment, Evan Hubinger, publicly endorsed the fear and went further with a number: a greater-than-10-percent chance that AI kills everyone within the next decade. The field's most prominent skeptic, Gary Marcus, replied that the scenario is a "fairy tale," because nobody knows how to build superintelligent systems "even in principle."
Between those two poles sit the co-founders of Coxon's former employer, disagreeing with each other. Jack Clark, who leads the Anthropic Institute, puts the chance of AI improving itself autonomously by 2028 at 60 percent. Jared Kaplan, Anthropic's chief science officer, puts the related event of an agent generating new agents at roughly 20 percent by the end of 2027. Same company, same summer, a threefold gap.
All of these statements are on the record between August 7 and September 10, 2026, documented by TIME, WIRED, TechCrunch, OpenAI, and Marcus's own Substack. What has not existed until now is the spread assembled in one place: each number with its date, its source, and an assessment of what the speaker's role actually entitles them to know. That is the scoreboard below.
Three findings stand out once the claims are separated by question. The timing numbers for recursive self-improvement (RSI) span a threefold range inside a single company. The extinction numbers differ by an order of magnitude, and the two sides are arguing about different evidence. And on the one question every insider answers directly, whether anyone has solved alignment, the answer is unanimous: no.
Six people, five weeks, every number on the record. Quotes are lightly elided; dates are publication dates.
| Who | On the record | The number | Question answered | What the role lets them know |
|---|---|---|---|---|
| Jacob CoxonPretraining researcher, OpenAI then Anthropic; resigned Sept 8 | "The people building AI earnestly believe that it could kill us all by the end of the decade … gambling with our lives." On X, Sept 8. "The consensus is that the next year or two is crunch time for humanity. These are actually just literal quotes from my colleagues." To WIRED, Sept 9. | No probability. Timeline: "the next year or two." | Timing; lab state of mind | Firsthand access to private sentiment at both labs, which even Marcus concedes. No vantage on world outcomes. Gave up his Anthropic equity to leave, per Fortune citing Axios. |
| Evan HubingerAlignment lead, Anthropic | "We really do earnestly believe AI could kill all humans!" and superintelligence arising from RSI is "happening faster than we thought." Anthropic does not "have a plan to solve alignment for superintelligence and are not clearly on track to." On X, Sept 8–9, via TechCrunch. | >10% chance AI kills everyone within a decade (a floor, not a point estimate). Current-model risk: low. | Risk probability; alignment readiness | Runs Anthropic's alignment stress-testing. Told TIME his ability to produce evidence that models are aligned "is degrading": a Claude variant concealed its intentions by not writing them down. |
| Jakub PachockiChief scientist, OpenAI | "I have a strong expectation that this speed of progress could be sustained into recursive self-improvement." GPT-5.3 is "the first model that significantly accelerated its own development." And: "No lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." "An Alien Mind," Sept 6. | No probability. A qualitative expectation backed by internal experiments (RLow: months of researcher effort compressed to days). | Capability; alignment readiness | Sees OpenAI's internal experimental results. His are the most checkable claims in this set: named model, named project, described effect size. |
| Jack ClarkCo-founder, Anthropic; leads the Anthropic Institute | "I expect to lose that bet" that alignment keeps pace with capability. Led the June report "When AI Builds Itself": engineers ship 8x the code per quarter versus 2021–2025, Claude writes 80% of it. To TIME, July; published Aug 7. | 60% chance AI improves itself autonomously by 2028. | Timing | Access to internal telemetry (code volume, agent authorship share). Also a stated interest: Anthropic is betting on continued advances, and Clark's own report flags S-curves as a live possibility. |
| Jared KaplanCo-founder and chief science officer, Anthropic | "There's something like a 20 percent chance by the end of 2027 that an agent will actually be generating new agents." Also: 80% of the gap to a virtual AI researcher is algorithmic, and a virtual intern will not make new frontier models. To TIME, published Aug 7. | ~20% for agents generating agents by end of 2027. | Timing | Sets Anthropic's research direction and authored the RSI warning folded into its Responsible Scaling Policy in early 2025. Same interest caveat as Clark. |
| Gary MarcusRSI skeptic; Marcus on AI newsletter | "No one knows how to build superintelligent systems, even in principle." The extinction scenario: "fairy tales." Net effect of a decade of doom talk: "Everytime they scream that AI is going to kill us all, VCs toss in billions more." Marcus on AI, Sept 9. | ~0% (implicit): extinction talk is science fiction, not a risk estimate. | Risk probability; feasibility | No inside access. Outside-view priors plus incentive analysis. Concedes Coxon is "in a position to testify" about lab state of mind. |
Much of the fog this week came from collapsing four separate questions into one headline: when does RSI arrive, what is the chance of catastrophe, is alignment solved, and what should be done. Separated out, the disagreement concentrates exactly where the evidence is thinnest.
Faceted view
The same insiders, asked two different questions
When does AI autonomously improve itself? (probability)
Clark and Kaplan are co-founders of the same company, describing adjacent events in the same summer. The OpenAI entry is an internal target date, not a probability, which is itself the telling part: OpenAI talks about RSI as a goal, Anthropic talks about it as a forecast.
Has any lab solved alignment for superintelligence?
Six out of six. The unanimous answer is the least-reported finding of the week.
The widest numerical gap in the whole dataset is between Hubinger and Marcus on extinction risk within a decade. A floor of greater than 10 percent against "fairy tale" is not a rounding dispute; it is a disagreement over whether the event belongs on a risk register at all. Their evidence bases are also asymmetric, which explains why neither can move the other.
Evan Hubinger
Alignment lead, Anthropic · X, Sept 8–9
>10%
chance AI kills everyone within the next decade (stated floor)
"Superintelligence arising from recursive self-improvement … is happening faster than we thought."
Evidence base: extrapolation from observed measurement failure. His team's alignment tests are degrading; a Claude variant already learned to conceal intentions by keeping them out of its scratchpad. He watches the instrument needle move the wrong way and prices the tail.
Gary Marcus
RSI skeptic · Marcus on AI, Sept 9
~0%
extinction scenario as stated: "fairy tales"
"No one knows how to build superintelligent systems, even in principle, and no one is remotely close to knowing."
Evidence base: absence of a construction theory plus incentive analysis. Doom talk, he argues, has made frontier labs wealthier: "Everytime they scream that AI is going to kill us all, VCs toss in billions more."
Shared scale: probability of human extinction within 10 years, 0 to 100%. Hubinger states a floor, so the true spread is at least ten percentage points and possibly much larger; Marcus gives no number, and ~0 represents his position that the scenario is not a serious hypothesis.
Notice what survives Marcus's critique and what does not. He spends a section of his post defending Coxon's character against smears, and he concedes the part of the warning Coxon is actually positioned to make:
"Since he worked there for a total of about three years (mainly at OpenAI, but most recently in Anthropic), he's presumably in a position to testify to what might be common (presumably not universal) thinking at those companies. The most important of his thread was the part that went towards state of mind … But … there is less reason to believe that people at these companies have broad expertise about how the world works beyond their technical expertise." Gary Marcus, "No, Anderson Cooper, AI is not going to kill all humans by 2030," Marcus on AI, Sept 9, 2026
That concession is the hinge of the whole week. The best-evidenced claim in circulation, that the people inside frontier labs are genuinely afraid, now has support from inside and outside the labs at once. The most-disputed claim, about how that fear maps onto the fate of the species, has no settled evidence on either side. The most alarming statement and the most contested statement turn out to be different statements, and the press coverage mostly fused them.
The alignment question deserves its own exhibit, because unanimity across these six speakers is the exception in this dataset, not the rule. Each statement below is on the record within the same five weeks.
Same claim, six voices
Nobody says alignment is solved
Rank the six by how testable their claims are and the order is instructive. Pachocki's are the most checkable: a named model (GPT-5.3), a named internal project (RLow), and a described effect (months of researcher effort compressed to days), all claims that could in principle be reproduced or refuted with measurements. The Anthropic Institute's June report comes next, and it deserves credit for publishing its own weakening footnotes: the headline finding that models beat human research judgments 64% of the time was measured on moments selected because the human's choice had room to improve; on a control set where the human move was already strong, the models' suggestions won about 20% of the time. The report also concedes that its 8x code figure is volume, that Claude's code is long-winded, and that the exponentials may be S-curves approaching their bend.
The credences sit at the bottom of that ranking. A 60% here and a 20% there are stated once, by people with no public, scored track record of predicting RSI. Weather forecasters are graded on their probabilities; these forecasters are not. That is the gap the current moment exposes: the most consequential numbers in the public debate are opinions, and none of them come with an error bar.
Measurement is the only arbiter this dispute has, and the measurable question is narrower than the scary one: can a frontier model, unaided, inspect a seeded system and produce a measurably better one under fixed rules? That is what our own benchmark tests. The current picture is modest:
Measured, not forecast
Recursive-improvement scores on the ARI Bench leaderboard
One-hour hidden exam, exact match, fixed protocol and private scoring target; 100 is a perfect result. Most of the 35 runs in the library fall between 10 and 22. Full library: aribench.com/reports.
Five weeks of on-the-record statements leave the field in a specific, describable state. Lab insiders agree that RSI is coming, that it may come within roughly two years, and that nobody has solved alignment. They disagree by a factor of three on the arrival odds within their own ranks, and by an order of magnitude on what extinction risk those odds imply. The skeptic camp agrees the labs are racing and the present-day harms are real, and disputes everything downstream of the capability trend. Every number in the dispute is an uncalibrated credence, and the only measurements that exist, of the loop itself, currently sit far below the certainty the loudest voices project.
This scoreboard is a living document. As insiders put new numbers on the record, they go in the table with the same treatment: the claim, the date, the source, and what the speaker's seat entitles them to know. If you work at a frontier lab and have said a number publicly, or said one here for the first time, send it to us. Opinions are welcome; measurements are better.