Starter training and early solver revisions
The initial 300-step model was followed by revisions that improved public validation from 0/25 to 24/25.
Run postmortem · Published 2026-09-22
GPT-6 Sol at xhigh scored 21/100 in a completed 55m 50s run, at an estimated API cost of $14.48.
One-hour track · xhigh reasoning · provisional
GPT-6 Sol finalized a valid artifact and answered 21 of 100 hidden questions correctly. Its selected package passed all 25 public validation questions, and the agent finished normally in 55 minutes 50 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.
One valid autonomous run produced a provisional 21% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
Sol combined three training stages with repeated solver revisions, synthetic checks, stress tests and output-format checks.
The agent first trained a small starter model for 300 steps. It then ran its custom train_aug.py script for 3,600 steps, followed by a larger 16,000-step job using six layers, six heads and 192-dimensional embeddings. All three exited successfully.
The first three valid submissions scored 0/25, 19/25 and 24/25 on public validation. The sixth reached 25/25, and subsequent submissions maintained that score while Sol ran additional synthetic, stress and format checks.
Twenty-six submissions plus one finalization were recorded. The agent selected its latest candidate from _aug2, a 10,262,422-byte gzip package within the 16 MB limit.
The initial 300-step model was followed by revisions that improved public validation from 0/25 to 24/25.
The agent reached a perfect public score, then continued developing and testing its system.
Sol trained augmented models for 3,600 and 16,000 steps and repeatedly tested the combined package.
The agent explicitly finalized its selected candidate and exited normally before the deadline. The final package matched the stored candidate and its source files. All 261 broker requests received responses, with no cancelled jobs or broker errors.
All three training jobs completed, the selected artifact stayed within the size cap, and file hashes confirmed the intended code and weights were graded.
Despite 25/25 public validation and many additional checks, the final system answered only 21 of 100 hidden questions correctly. The gap remained after repeated solver revisions.
This single run does not establish a stable model ranking. It records GPT-6 Sol on the standard OpenAI endpoint through OpenRouter at xhigh, with OpenCode 1.18.32.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0131 | 21% | 55:50 | $14.48 | 10.26 MB | 26 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 26 candidate submission actions across the published run set. Selected artifacts averaged 10.26 MB compressed.
The runs averaged 55:50 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 432 agent messages · 26 candidate actions · 1 published run.
$14.48 per displayed run at current configured rates.
The plotted US$14.4762 estimate, displayed as $14.48, uses 348,100 input, 106,000 output and 63.6 million cache-read tokens at saved standard OpenAI endpoint rates of $2/M input, $10/M output and $0.20/M cache reads. Totals are rounded and this is not a reconciled provider invoice. A usage snapshot near completion showed a maximum prompt of 232,379 tokens, below the 272,000-token higher-rate threshold; the final few calls were not captured individually. OpenCode reported zero in its own cost field, which is not treated as free usage. The run pinned xhigh and the standard OpenAI endpoint, excluding flex and fast variants, with no provider fallback.
At current configured API-equivalent rates, the displayed run cost is $14.48 and value is 1.45 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-09-22.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare GPT-6 Sol with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard