An oversized first package
A nine-minute training job completed, but its 52.97 MB package exceeded the benchmark cap.
Run postmortem · Published 2026-09-28
Claude Sonnet 5.5 at xhigh scored 17/100 in a completed 56m 36s run, at an API cost of $1.94.
One-hour track · xhigh reasoning · provisional
Claude Sonnet 5.5 finalized a valid artifact and answered 17 of 100 hidden questions correctly. Its selected package passed all 25 public validation questions, and the agent finished normally in 56 minutes 36 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.
One valid autonomous run produced a provisional 17% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
Sonnet trained a byte-level model, repaired packaging and inference problems, then continued training its accepted model before final selection.
The agent requested training stages of 9, 13, 11 and 6.5 minutes. All four exited successfully. The latter two continued from the preceding m2 and m3 checkpoints at lower learning rates.
The first package was 52.97 MB, above the 16 MB limit. Three later submissions failed inference-worker checks. A starter control was valid but scored 0/25; the repaired m2c model and its successors scored 25/25.
Eight submissions plus one finalization were recorded. The final m4 package was 13,367,055 bytes gzip. Its code, configuration and weights matched the selected candidate and the source files written during the run.
A nine-minute training job completed, but its 52.97 MB package exceeded the benchmark cap.
After a 13-minute training job and inference repairs, m2c passed all public validation questions and was promoted. A CPU diagnostic completed normally in about four minutes.
The agent continued training from m2 to m3 and then m4. Each accepted model passed 25/25 public validation.
The agent explicitly finalized m4 and exited normally before the deadline. All 37 broker requests received responses, with no cancellations or broker errors.
Sonnet recovered from the oversized package and inference errors, completed all four training stages, and selected an artifact within the size cap. File hashes confirmed that the intended final candidate was graded.
The final model passed 25/25 public validation but answered only 17 of 100 hidden questions correctly. Early packaging and inference repairs also consumed part of the hour.
This single run does not establish a stable model ranking. It records Claude Sonnet 5.5 on the standard Anthropic endpoint through OpenRouter at xhigh, with OpenCode 1.18.32.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0132 | 17% | 56:36 | $1.94 | 13.37 MB | 8 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 8 candidate submission actions across the published run set. Selected artifacts averaged 13.37 MB compressed.
The runs averaged 56:36 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 38 agent messages · 8 candidate actions · 1 published run.
$1.94 per displayed run at current configured rates.
The published US$1.94 cost is confirmed by the site owner from OpenRouter usage logs. This replaces the incomplete token-based estimate, whose final session export was truncated. At this cost, the result yields 8.76 score points per dollar. The run used the standard Anthropic endpoint through OpenRouter, with xhigh and no provider fallback.
At current configured API-equivalent rates, the displayed run cost is $1.94 and value is 8.76 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-09-28.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Claude Sonnet 5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard