Run postmortem · Published 2026-09-29

GPT-6.1 Solxhigh reasoning

GPT-6.1 Sol at xhigh scored 27/100 in a completed 58m 48s run, at an estimated API cost of $2.39.

27/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
27%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#4
Estimated cost$2.39
Points / $11.29
Average time58:48
01 · Verdict

A completed xhigh run scored 27/100

GPT-6.1 Sol explicitly finalized a valid artifact and answered 27 of 100 hidden questions correctly. Its selected package passed all 25 public validation questions, and the agent finished normally in 58 minutes 48 seconds. This is one valid run in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the compact system built by the agent, rather than asking the controller to answer the hidden questions directly.

One valid autonomous run produced a provisional 27% hidden exact-match score.

#4
Score position27% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GPT-6.1 Sol built

GPT-6.1 Sol trained a byte-level model, expanded it, continued training with additional context examples, and revised inference code through repeated validation checks.

01

Five successful training jobs

A 500-step starter run was followed by an 18,000-step task-model run, a larger training stage with a nine-minute budget, and two continuation stages with 7.5-minute and six-minute budgets. All five exited successfully.

02

Larger model and continued training

The larger model used six layers, width 320 and five attention heads. The final two jobs resumed from earlier checkpoints with lower learning rates and context augmentation.

03

Verified final selection

Eleven submissions and one explicit finalization were recorded. The completion_final package was 14,540,991 bytes gzip, within the 16 MB cap. Its code, configuration and weights matched the saved selected candidate and the source files written before finalization.

03 · Trajectory

How the run unfolded

Starter

A valid baseline scored 0/25

The 500-step starter package was valid but answered none of the public validation questions correctly.

Training and validation

Accepted candidates reached 25/25

The task model and larger model passed all public validation questions. The agent then continued training and revised inference behavior while preserving accepted candidates.

Final checks

Five training jobs finished cleanly

The later stages resumed from earlier checkpoints. Interface, composition and artifact checks completed, with no failed or cancelled compute jobs.

Final grade

27 of 100 hidden exact matches

The agent explicitly finalized completion_final and exited normally before the deadline. All 90 broker requests received responses, with no broker errors. The graded artifact matched the selected candidate.

04 · Assessment

Good, bad, and unresolved

What worked

Completed, checked and finalized a valid package

The agent completed all five training jobs and its validation checks without compute failures or cancellations. It saved a package within the size cap and explicitly selected it before exiting. Artifact hashes confirmed that the intended candidate was graded.

What hurt

Perfect public validation did not transfer to the hidden exam

The selected package passed 25/25 public cases but answered 27/100 hidden questions correctly. The public validation score did not establish equivalent accuracy on unseen questions.

What remains unknown

One provisional xhigh result

This single run does not establish a stable model ranking. It records GPT-6.1 Sol on the standard OpenAI endpoint through OpenRouter at xhigh, with OpenCode 1.18.32.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-013327%58:48$2.3914.54 MB11Valid
187.4KInput tokens
80.9KVisible output tokens
10.22MCache-read tokens

Reasoning tokens: 18,544, billed at the output rate.

01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 11 candidate submission actions across the published run set. Selected artifacts averaged 14.54 MB compressed.

03

Time use

The runs averaged 58:48 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 99 agent messages · 11 candidate actions · 1 published run.

06 · Economics

What the result cost

$2.39 estimated for this run, including reasoning.

The estimated API cost is US$2.39, calculated from complete metadata for all 98 model responses. It includes 187,386 uncached input tokens, 80,895 visible output tokens, 18,544 reasoning tokens and 10,219,079 cached input tokens. Rates were $2 per million input tokens, $10 per million output and reasoning tokens, and $0.10 per million cached input tokens. No request crossed the 272,000-token long-context threshold. This is a token-based estimate; the provider-billed dollar amount has not been independently reconciled. The result yields 11.29 score points per dollar at these rates.

At current configured API-equivalent rates, the displayed run cost is $2.39 and value is 11.29 score points per dollar.

#10
Value positionValue rank #10 of 39 valid priced entries. Higher score points per dollar indicate better value.

Estimated API cost, including reasoning / run · pricing snapshot 2026-09-29.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GPT-6.1 Sol with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard