A valid baseline scored 0/25
The 500-step starter package was valid but answered none of the public validation questions correctly.
Run postmortem · Published 2026-09-29
GPT-6.1 Sol at xhigh scored 27/100 in a completed 58m 48s run, at an estimated API cost of $2.39.
One-hour track · xhigh reasoning · provisional
GPT-6.1 Sol explicitly finalized a valid artifact and answered 27 of 100 hidden questions correctly. Its selected package passed all 25 public validation questions, and the agent finished normally in 58 minutes 48 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.
One valid autonomous run produced a provisional 27% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
GPT-6.1 Sol trained a byte-level model, expanded it, continued training with additional context examples, and revised inference code through repeated validation checks.
A 500-step starter run was followed by an 18,000-step task-model run, a larger training stage with a nine-minute budget, and two continuation stages with 7.5-minute and six-minute budgets. All five exited successfully.
The larger model used six layers, width 320 and five attention heads. The final two jobs resumed from earlier checkpoints with lower learning rates and context augmentation.
Eleven submissions and one explicit finalization were recorded. The completion_final package was 14,540,991 bytes gzip, within the 16 MB cap. Its code, configuration and weights matched the saved selected candidate and the source files written before finalization.
The 500-step starter package was valid but answered none of the public validation questions correctly.
The task model and larger model passed all public validation questions. The agent then continued training and revised inference behavior while preserving accepted candidates.
The later stages resumed from earlier checkpoints. Interface, composition and artifact checks completed, with no failed or cancelled compute jobs.
The agent explicitly finalized completion_final and exited normally before the deadline. All 90 broker requests received responses, with no broker errors. The graded artifact matched the selected candidate.
The agent completed all five training jobs and its validation checks without compute failures or cancellations. It saved a package within the size cap and explicitly selected it before exiting. Artifact hashes confirmed that the intended candidate was graded.
The selected package passed 25/25 public cases but answered 27/100 hidden questions correctly. The public validation score did not establish equivalent accuracy on unseen questions.
This single run does not establish a stable model ranking. It records GPT-6.1 Sol on the standard OpenAI endpoint through OpenRouter at xhigh, with OpenCode 1.18.32.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0133 | 27% | 58:48 | $2.39 | 14.54 MB | 11 | Valid |
Reasoning tokens: 18,544, billed at the output rate.
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 11 candidate submission actions across the published run set. Selected artifacts averaged 14.54 MB compressed.
The runs averaged 58:48 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 99 agent messages · 11 candidate actions · 1 published run.
$2.39 estimated for this run, including reasoning.
The estimated API cost is US$2.39, calculated from complete metadata for all 98 model responses. It includes 187,386 uncached input tokens, 80,895 visible output tokens, 18,544 reasoning tokens and 10,219,079 cached input tokens. Rates were $2 per million input tokens, $10 per million output and reasoning tokens, and $0.10 per million cached input tokens. No request crossed the 272,000-token long-context threshold. This is a token-based estimate; the provider-billed dollar amount has not been independently reconciled. The result yields 11.29 score points per dollar at these rates.
At current configured API-equivalent rates, the displayed run cost is $2.39 and value is 11.29 score points per dollar.
Estimated API cost, including reasoning / run · pricing snapshot 2026-09-29.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare GPT-6.1 Sol with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard