Run postmortem · Published 2026-09-22

Claude Opus 5.5xhigh reasoning

Claude Opus 5.5 at xhigh scored 25/100 in a completed 47m 56s run, at an estimated API cost of $2.64.

25/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
25%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#4
Current cost$2.64
Points / $9.47
Average time47:56
01 · Verdict

A completed xhigh run scored 25/100

Claude Opus 5.5 finalized a valid artifact and answered 25 of 100 hidden questions correctly. It passed all 25 public validation questions and finished normally in 47 minutes 56 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.

One valid autonomous run produced a provisional 25% hidden exact-match score.

#4
Score position25% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Claude Opus 5.5 built

Opus combined a trained model with task-specific solver code, then packaged a quantized final artifact.

01

Three completed training jobs

The agent ran training jobs with requested durations of 14, 10, and 9 minutes. The latter jobs continued from its own earlier checkpoints; all three exited successfully.

02

Solver and inference revisions

The agent developed solver logic alongside training and corrected early output-shape errors in its submission interface.

03

Explicit final selection

Nine submissions and one finalization were recorded. The final artifact from m3q was 10,121,091 bytes gzip, below the 16 MB cap.

03 · Trajectory

How the run unfolded

Initial training

A trained checkpoint and solver code

Opus trained its first model and developed inference and evaluation code.

Submission repair

Three malformed submissions were corrected

The first three submissions failed the distribution-shape contract. The fourth was valid and passed 25/25 public validation.

Further training

Two more training stages completed

The agent continued training from its checkpoints and saved additional valid 25/25 candidates.

Final grade

25 of 100 hidden exact matches

The agent explicitly finalized its submission and exited normally before the one-hour deadline. The hidden grade was valid; no broker errors were recorded.

04 · Assessment

Good, bad, and unresolved

What worked

Recovered and finalized a valid trained system

The agent corrected the initial submission errors, completed all three training jobs, and selected a package within the size cap.

What hurt

Queued checks expired during training

Five evaluation requests were cancelled before starting after their callers timed out while waiting behind active training. Public validation reached 25/25, but hidden accuracy was 25/100.

What remains unknown

One provisional run

One run does not establish a stable model ranking. This result is specifically xhigh on OpenCode 1.18.32; a different effort setting is a separate execution.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-012925%47:56$2.6410.12 MB9Valid
98.8KInput tokens
49.3KOutput tokens
6.3MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 9 candidate submission actions across the published run set. Selected artifacts averaged 10.12 MB compressed.

03

Time use

The runs averaged 47:56 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 99 agent messages · 9 candidate actions · 1 published run.

06 · Economics

What the result cost

$2.64 per displayed run at current configured rates.

The plotted US$2.6412 estimate, displayed as $2.64, uses 98,800 input, 49,300 output, and 6.3 million cache-read tokens at the saved Anthropic endpoint rates of $4/M input, $20/M output, and $0.20/M cache reads. Totals are rounded, and this is not a reconciled provider invoice. OpenCode reported zero in its own cost field; that is not treated as evidence of free usage. The run used the standard Anthropic endpoint through OpenRouter, pinned to xhigh with no provider fallback.

At current configured API-equivalent rates, the displayed run cost is $2.64 and value is 9.47 score points per dollar.

#10
Value positionValue rank #10 of 35 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-22.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Claude Opus 5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard