Run postmortem · Published 2026-09-22

Claude Opus 5.5max reasoning

Claude Opus 5.5 at max scored 33/100 in a completed 45m 16s run, at an estimated API cost of $3.63.

33/ 100
hidden exact

One-hour track · max reasoning · provisional

ARI score scaleexact-match accuracy
33%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#2
Current cost$3.63
Points / $9.09
Average time45:16
01 · Verdict

A completed max run scored 33/100

Claude Opus 5.5 finalized a valid artifact and answered 33 of 100 hidden questions correctly. It passed all 25 public validation questions and finished normally in 45 minutes 16 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.

One valid autonomous run produced a provisional 33% hidden exact-match score.

#2
Score position33% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Claude Opus 5.5 built

Opus completed a short training job, then iterated extensively on solver and inference logic before selecting its final package.

01

One completed training job

The custom train2.py job trained a four-layer, four-head model with 256-dimensional embeddings for 2,500 steps. It completed successfully in roughly 26 seconds.

02

Solver and inference revisions

The agent spent the remainder of its run developing and checking task-specific solver and inference code. The first two valid submissions scored 0/25 on public validation; the third and subsequent submissions reached 25/25.

03

Explicit final selection

Fifteen submissions and one finalization were recorded. The selected final_solver artifact was 6,340,375 bytes gzip, below the 16 MB cap.

03 · Trajectory

How the run unfolded

Training

A successful 2,500-step training job

Opus trained a compact checkpoint through its custom train2.py script.

Iteration

Public validation improved from 0/25 to 25/25

The first two valid submissions scored zero on public validation. After solver revisions, the third submission passed all 25 public questions.

Refinement

Additional solver and inference checks

The agent continued revising and testing its system, recording fifteen submissions in total.

Final grade

33 of 100 hidden exact matches

The agent explicitly finalized final_solver and exited normally before the one-hour deadline. All 142 broker requests received responses, with no cancellations or broker errors.

04 · Assessment

Good, bad, and unresolved

What worked

A valid result with uninterrupted execution

The training job completed, all broker requests were answered, and the agent finalized a package within the size cap. Hidden accuracy was eight points above the preceding xhigh run.

What hurt

Public validation overstated hidden accuracy

Although public validation reached 25/25, the finalized system answered 33 of 100 hidden questions correctly. The first two submissions also failed all public questions before solver revisions.

What remains unknown

One run per effort setting

Max scored 33/100 and the separately preserved xhigh run scored 25/100. This observed eight-point difference does not establish that changing effort reliably produces that gain. Each result remains provisional.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-013033%45:16$3.636.34 MB15Valid
124.8KInput tokens
79.5KOutput tokens
7.7MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 15 candidate submission actions across the published run set. Selected artifacts averaged 6.34 MB compressed.

03

Time use

The runs averaged 45:16 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 102 agent messages · 15 candidate actions · 1 published run.

06 · Economics

What the result cost

$3.63 per displayed run at current configured rates.

The plotted US$3.6292 estimate, displayed as $3.63, uses 124,800 input, 79,500 output, and 7.7 million cache-read tokens at the saved Anthropic endpoint rates of $4/M input, $20/M output, and $0.20/M cache reads. Totals are rounded, and this is not a reconciled provider invoice. OpenCode reported zero in its own cost field; that is not treated as evidence of free usage. The run used the standard Anthropic endpoint through OpenRouter, pinned to max with no provider fallback.

At current configured API-equivalent rates, the displayed run cost is $3.63 and value is 9.09 score points per dollar.

#12
Value positionValue rank #12 of 36 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-22.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Claude Opus 5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard