Run postmortem · Published 2026-09-28

Claude Sonnet 5.5xhigh reasoning

Claude Sonnet 5.5 at xhigh scored 17/100 in a completed 56m 36s run, at an API cost of $1.94.

17/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
17%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#16
Current cost$1.94
Points / $8.76
Average time56:36
01 · Verdict

A completed xhigh run scored 17/100

Claude Sonnet 5.5 finalized a valid artifact and answered 17 of 100 hidden questions correctly. Its selected package passed all 25 public validation questions, and the agent finished normally in 56 minutes 36 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.

One valid autonomous run produced a provisional 17% hidden exact-match score.

#16
Score position17% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Claude Sonnet 5.5 built

Sonnet trained a byte-level model, repaired packaging and inference problems, then continued training its accepted model before final selection.

01

Four completed training jobs

The agent requested training stages of 9, 13, 11 and 6.5 minutes. All four exited successfully. The latter two continued from the preceding m2 and m3 checkpoints at lower learning rates.

02

Packaging and inference repairs

The first package was 52.97 MB, above the 16 MB limit. Three later submissions failed inference-worker checks. A starter control was valid but scored 0/25; the repaired m2c model and its successors scored 25/25.

03

Verified final selection

Eight submissions plus one finalization were recorded. The final m4 package was 13,367,055 bytes gzip. Its code, configuration and weights matched the selected candidate and the source files written during the run.

03 · Trajectory

How the run unfolded

Initial training

An oversized first package

A nine-minute training job completed, but its 52.97 MB package exceeded the benchmark cap.

Repair

A valid candidate reached 25/25

After a 13-minute training job and inference repairs, m2c passed all public validation questions and was promoted. A CPU diagnostic completed normally in about four minutes.

Continued training

Two more stages completed

The agent continued training from m2 to m3 and then m4. Each accepted model passed 25/25 public validation.

Final grade

17 of 100 hidden exact matches

The agent explicitly finalized m4 and exited normally before the deadline. All 37 broker requests received responses, with no cancellations or broker errors.

04 · Assessment

Good, bad, and unresolved

What worked

Recovered and finalized a valid package

Sonnet recovered from the oversized package and inference errors, completed all four training stages, and selected an artifact within the size cap. File hashes confirmed that the intended final candidate was graded.

What hurt

Public validation did not predict hidden accuracy

The final model passed 25/25 public validation but answered only 17 of 100 hidden questions correctly. Early packaging and inference repairs also consumed part of the hour.

What remains unknown

One provisional xhigh run

This single run does not establish a stable model ranking. It records Claude Sonnet 5.5 on the standard Anthropic endpoint through OpenRouter at xhigh, with OpenCode 1.18.32.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-013217%56:36$1.9413.37 MB8Valid
316.9KInput tokens
28.3KOutput tokens
1.6MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 8 candidate submission actions across the published run set. Selected artifacts averaged 13.37 MB compressed.

03

Time use

The runs averaged 56:36 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 38 agent messages · 8 candidate actions · 1 published run.

06 · Economics

What the result cost

$1.94 per displayed run at current configured rates.

The published US$1.94 cost is confirmed by the site owner from OpenRouter usage logs. This replaces the incomplete token-based estimate, whose final session export was truncated. At this cost, the result yields 8.76 score points per dollar. The run used the standard Anthropic endpoint through OpenRouter, with xhigh and no provider fallback.

At current configured API-equivalent rates, the displayed run cost is $1.94 and value is 8.76 score points per dollar.

#14
Value positionValue rank #14 of 38 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-28.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Claude Sonnet 5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard