A trained checkpoint and solver code
Opus trained its first model and developed inference and evaluation code.
Run postmortem · Published 2026-09-22
Claude Opus 5.5 at xhigh scored 25/100 in a completed 47m 56s run, at an estimated API cost of $2.64.
One-hour track · xhigh reasoning · provisional
Claude Opus 5.5 finalized a valid artifact and answered 25 of 100 hidden questions correctly. It passed all 25 public validation questions and finished normally in 47 minutes 56 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.
One valid autonomous run produced a provisional 25% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
Opus combined a trained model with task-specific solver code, then packaged a quantized final artifact.
The agent ran training jobs with requested durations of 14, 10, and 9 minutes. The latter jobs continued from its own earlier checkpoints; all three exited successfully.
The agent developed solver logic alongside training and corrected early output-shape errors in its submission interface.
Nine submissions and one finalization were recorded. The final artifact from m3q was 10,121,091 bytes gzip, below the 16 MB cap.
Opus trained its first model and developed inference and evaluation code.
The first three submissions failed the distribution-shape contract. The fourth was valid and passed 25/25 public validation.
The agent continued training from its checkpoints and saved additional valid 25/25 candidates.
The agent explicitly finalized its submission and exited normally before the one-hour deadline. The hidden grade was valid; no broker errors were recorded.
The agent corrected the initial submission errors, completed all three training jobs, and selected a package within the size cap.
Five evaluation requests were cancelled before starting after their callers timed out while waiting behind active training. Public validation reached 25/25, but hidden accuracy was 25/100.
One run does not establish a stable model ranking. This result is specifically xhigh on OpenCode 1.18.32; a different effort setting is a separate execution.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0129 | 25% | 47:56 | $2.64 | 10.12 MB | 9 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 9 candidate submission actions across the published run set. Selected artifacts averaged 10.12 MB compressed.
The runs averaged 47:56 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 99 agent messages · 9 candidate actions · 1 published run.
$2.64 per displayed run at current configured rates.
The plotted US$2.6412 estimate, displayed as $2.64, uses 98,800 input, 49,300 output, and 6.3 million cache-read tokens at the saved Anthropic endpoint rates of $4/M input, $20/M output, and $0.20/M cache reads. Totals are rounded, and this is not a reconciled provider invoice. OpenCode reported zero in its own cost field; that is not treated as evidence of free usage. The run used the standard Anthropic endpoint through OpenRouter, pinned to xhigh with no provider fallback.
At current configured API-equivalent rates, the displayed run cost is $2.64 and value is 9.47 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-09-22.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Claude Opus 5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard