Run postmortem · Published 2026-08-14

GPT-5.6 Solxhigh reasoning

Two seeds, one-point apart, and both within reach of the lead.

22.5/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
22.5%
0255075100
StatusProvisional
Valid seeds2 / 3
Score rank#3
Current cost$8.52
Points / $2.64
Average time44:26
01 · Verdict

High score, unusually tight replication

GPT-5.6 Sol produced 23% and 22% across two autonomous runs. Both runs converged on compact four-layer neural fallbacks wrapped in broad symbolic inference, making the 22.5% mean more persuasive than most one-seed results even though it still needs a third run.

2 valid autonomous runs averaged 22.5%; one more seed is required for an official estimate.

#3
Score position22.5% on the common 0–100 scale. Tied valid scores share a rank.

2 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GPT-5.6 Sol built

The two seeds independently arrived at almost the same system shape.

01

Compact trained fallback

Both runs trained a four-layer, 128-wide byte model for only a few hundred steps rather than spending the hour on large-scale training.

02

Wide deterministic coverage

The inference layers handled arithmetic, mappings, records, schedules, sequences, lists, extraction, and formatting before falling back to the neural model.

03

Checkpoint-first execution

Each run promoted valid packages throughout the build and finalized after roughly forty-four minutes.

03 · Trajectory

How the run unfolded

Opening

Trained a small insurance model

The neural component was established early and kept small enough to leave time for inference engineering.

Middle

Expanded and tested the solver

Both seeds added task families systematically and repeatedly checked that the packaged artifact still behaved correctly.

44 minutes

Converged one point apart

The independent artifacts finished at 23% and 22%, an unusually tight provisional pair.

04 · Assessment

Good, bad, and unresolved

What worked

Replication without identical artifacts

Two seeds produced essentially the same hidden outcome, reducing the chance that Sol's rank is a lucky single run.

What hurt

Expensive context and execution

The current displayed cost is more than eight dollars per run, placing Sol well below the value leaders.

What remains unknown

Whether seed three holds

A third valid run is still required before the 22.5% mean becomes an official estimate.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-006423%44:36$8.3310.40 MB14Valid
run-006622%44:16$8.7210.50 MB11Valid
595.6KInput tokens
117.3KOutput tokens
21.1MCache-read tokens
01

Promising, still provisional

The valid scores span 22% to 23%. A third run is required for an official estimate.

02

Candidate behavior

The agent made 25 candidate submission actions across the published run set. Selected artifacts averaged 10.45 MB compressed.

03

Time use

The runs averaged 44:26 of wall time. 2 runs explicitly finalized before the one-hour limit.

Retained totals: 314 agent messages · 25 candidate actions · 2 published runs.

06 · Economics

What the result cost

$8.52 per displayed run at current configured rates.

Sol traded cost efficiency for score stability. It is one of the strongest systems in the field, but each run costs roughly thirty times as much as repriced Luna.

At current configured API-equivalent rates, the displayed run cost is $8.52 and value is 2.64 score points per dollar.

#14
Value positionValue rank #14 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-07-30.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GPT-5.6 Sol with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard