Run postmortem · Published 2026-09-22

GPT-6 Solxhigh reasoning

GPT-6 Sol at xhigh scored 21/100 in a completed 55m 50s run, at an estimated API cost of $14.48.

21/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
21%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#9
Current cost$14.48
Points / $1.45
Average time55:50
01 · Verdict

A completed xhigh run scored 21/100

GPT-6 Sol finalized a valid artifact and answered 21 of 100 hidden questions correctly. Its selected package passed all 25 public validation questions, and the agent finished normally in 55 minutes 50 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.

One valid autonomous run produced a provisional 21% hidden exact-match score.

#9
Score position21% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GPT-6 Sol built

Sol combined three training stages with repeated solver revisions, synthetic checks, stress tests and output-format checks.

01

Three successful training jobs

The agent first trained a small starter model for 300 steps. It then ran its custom train_aug.py script for 3,600 steps, followed by a larger 16,000-step job using six layers, six heads and 192-dimensional embeddings. All three exited successfully.

02

Repeated solver and validation checks

The first three valid submissions scored 0/25, 19/25 and 24/25 on public validation. The sixth reached 25/25, and subsequent submissions maintained that score while Sol ran additional synthetic, stress and format checks.

03

Explicit selection of the final package

Twenty-six submissions plus one finalization were recorded. The agent selected its latest candidate from _aug2, a 10,262,422-byte gzip package within the 16 MB limit.

03 · Trajectory

How the run unfolded

Initial work

Starter training and early solver revisions

The initial 300-step model was followed by revisions that improved public validation from 0/25 to 24/25.

Public validation

The sixth submission reached 25/25

The agent reached a perfect public score, then continued developing and testing its system.

Further training

Two custom training jobs completed

Sol trained augmented models for 3,600 and 16,000 steps and repeatedly tested the combined package.

Final grade

21 of 100 hidden exact matches

The agent explicitly finalized its selected candidate and exited normally before the deadline. The final package matched the stored candidate and its source files. All 261 broker requests received responses, with no cancelled jobs or broker errors.

04 · Assessment

Good, bad, and unresolved

What worked

A valid, explicitly finalized package

All three training jobs completed, the selected artifact stayed within the size cap, and file hashes confirmed the intended code and weights were graded.

What hurt

Public success did not transfer to the hidden exam

Despite 25/25 public validation and many additional checks, the final system answered only 21 of 100 hidden questions correctly. The gap remained after repeated solver revisions.

What remains unknown

One provisional xhigh run

This single run does not establish a stable model ranking. It records GPT-6 Sol on the standard OpenAI endpoint through OpenRouter at xhigh, with OpenCode 1.18.32.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-013121%55:50$14.4810.26 MB26Valid
348.1KInput tokens
106KOutput tokens
63.6MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 26 candidate submission actions across the published run set. Selected artifacts averaged 10.26 MB compressed.

03

Time use

The runs averaged 55:50 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 432 agent messages · 26 candidate actions · 1 published run.

06 · Economics

What the result cost

$14.48 per displayed run at current configured rates.

The plotted US$14.4762 estimate, displayed as $14.48, uses 348,100 input, 106,000 output and 63.6 million cache-read tokens at saved standard OpenAI endpoint rates of $2/M input, $10/M output and $0.20/M cache reads. Totals are rounded and this is not a reconciled provider invoice. A usage snapshot near completion showed a maximum prompt of 232,379 tokens, below the 272,000-token higher-rate threshold; the final few calls were not captured individually. OpenCode reported zero in its own cost field, which is not treated as free usage. The run pinned xhigh and the standard OpenAI endpoint, excluding flex and fast variants, with no provider fallback.

At current configured API-equivalent rates, the displayed run cost is $14.48 and value is 1.45 score points per dollar.

#29
Value positionValue rank #29 of 37 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-22.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GPT-6 Sol with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard