A successful 2,500-step training job
Opus trained a compact checkpoint through its custom train2.py script.
Run postmortem · Published 2026-09-22
Claude Opus 5.5 at max scored 33/100 in a completed 45m 16s run, at an estimated API cost of $3.63.
One-hour track · max reasoning · provisional
Claude Opus 5.5 finalized a valid artifact and answered 33 of 100 hidden questions correctly. It passed all 25 public validation questions and finished normally in 45 minutes 16 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.
One valid autonomous run produced a provisional 33% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
Opus completed a short training job, then iterated extensively on solver and inference logic before selecting its final package.
The custom train2.py job trained a four-layer, four-head model with 256-dimensional embeddings for 2,500 steps. It completed successfully in roughly 26 seconds.
The agent spent the remainder of its run developing and checking task-specific solver and inference code. The first two valid submissions scored 0/25 on public validation; the third and subsequent submissions reached 25/25.
Fifteen submissions and one finalization were recorded. The selected final_solver artifact was 6,340,375 bytes gzip, below the 16 MB cap.
Opus trained a compact checkpoint through its custom train2.py script.
The first two valid submissions scored zero on public validation. After solver revisions, the third submission passed all 25 public questions.
The agent continued revising and testing its system, recording fifteen submissions in total.
The agent explicitly finalized final_solver and exited normally before the one-hour deadline. All 142 broker requests received responses, with no cancellations or broker errors.
The training job completed, all broker requests were answered, and the agent finalized a package within the size cap. Hidden accuracy was eight points above the preceding xhigh run.
Although public validation reached 25/25, the finalized system answered 33 of 100 hidden questions correctly. The first two submissions also failed all public questions before solver revisions.
Max scored 33/100 and the separately preserved xhigh run scored 25/100. This observed eight-point difference does not establish that changing effort reliably produces that gain. Each result remains provisional.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0130 | 33% | 45:16 | $3.63 | 6.34 MB | 15 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 15 candidate submission actions across the published run set. Selected artifacts averaged 6.34 MB compressed.
The runs averaged 45:16 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 102 agent messages · 15 candidate actions · 1 published run.
$3.63 per displayed run at current configured rates.
The plotted US$3.6292 estimate, displayed as $3.63, uses 124,800 input, 79,500 output, and 7.7 million cache-read tokens at the saved Anthropic endpoint rates of $4/M input, $20/M output, and $0.20/M cache reads. Totals are rounded, and this is not a reconciled provider invoice. OpenCode reported zero in its own cost field; that is not treated as evidence of free usage. The run used the standard Anthropic endpoint through OpenRouter, pinned to max with no provider fallback.
At current configured API-equivalent rates, the displayed run cost is $3.63 and value is 9.09 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-09-22.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Claude Opus 5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard