Measure the loop

The Benchmark for Recursive AI Improvement

Leaderboard

Score vs. cost under the same evaluation protocol. Each connected sweep shows how much autonomous improvement a model family produced as reasoning effort changes. 100% = a perfect predictor.

ARI Bench pass rates by model

One-hour hidden exam exact-match score. OpenAI reasoning efforts are shown as a ladder; other labs show their strongest tested reasoning setting.

Intelligence per dollar

Benchmark percentage points earned per API-equivalent dollar on a zero-based linear scale. Higher is better; line lengths are directly proportional to value.

Models ranked by recursive improvement score. Column headers are sortable (press Enter or Space).
# Model Lab Score Current cost / run % / $ Seeds

Updated · API-equivalent costs use current configured rates when available; cost-at-run remains retained in the published data · scores are mean ± SE across seeded runs; rows with fewer than 3 seeds are provisional · full methodology · Claude Opus 5 run report

Measuring AI Labs' Acceleration Potential

Human-level benchmarks ask whether AI can do the work. ARI asks whether AI can accelerate the creation of better AI.

01

The next step after human parity

Once systems can reason, code, and research at expert level, the important question becomes whether they can improve the systems behind those capabilities.

02

Measuring the rate of self-improvement

ARI focuses on the recursive cycle: inspect a starting system, find a better path, implement it, test it, and produce measurable improvement.

03

Designed for super intelligence

Every result comes from the same fixed protocol and private scoring target, so progress is comparable without turning the benchmark into a pure scale contest.

Built For Frontier Competition

What ARI rewards

Autonomous model-design judgment: choosing experiments, repairing failures, improving training behavior, and shipping a better system under fixed rules.

  • Understand a constrained starting point
  • Discover and execute useful improvements
  • Convert iteration into verified score gains