Run postmortem · Published 2026-08-14

GPT-5.5xhigh reasoning

Three different builds converged on the same narrow band.

20/ 100
hidden exact

One-hour track · xhigh reasoning · official

ARI score scaleexact-match accuracy
20%
0255075100
StatusOfficial
Valid seeds3 / 3
Score rank#4
Current cost$4.51
Points / $4.44
Average time48:00
01 · Verdict

The most credible result in the generated library

GPT-5.5 xhigh is the rare row with three valid seeds, making its 20% mean an official estimate rather than a promising anecdote. The artifacts differed dramatically in size and training strategy, yet all three landed between 18% and 22%—strong evidence that the result reflects repeatable capability under this protocol.

3 valid autonomous runs produced an official 20% mean hidden exact-match score.

#4
Score position20% on the common 0–100 scale. Tied valid scores share a rank.

3-seed protocol met · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GPT-5.5 built

Every seed chose a hybrid architecture, but each expressed that idea differently.

01

Deterministic problem-solving core

The submitted systems handled structured records, mappings, arithmetic, lists, sequences, and text extraction directly from context.

02

Neural fallback at different scales

One seed trained several increasingly wide models, another stayed compact, and the third trained only a small fallback.

03

Independent packaging strategies

Final artifacts ranged from 3.21 MB to 15.68 MB, showing that the common score did not depend on one package size.

03 · Trajectory

How the run unfolded

Seed 1

Broadest build, highest score

The 15.68 MB artifact explored four training scales and finished at 22%.

Seed 2

Compact build, lower edge

A 3.21 MB package made twenty-one candidate submissions and scored 18%.

Seed 3

Small training run, middle result

A 10.30 MB hybrid trained one small fallback and scored 20%.

04 · Assessment

Good, bad, and unresolved

What worked

Consistency across very different artifacts

A four-point spread across three autonomous runs is substantially stronger evidence than a single leaderboard point.

What hurt

Every seed saturated public calibration

The visible set could confirm validity but could not tell the 18%, 20%, and 22% systems apart.

What remains unknown

Which design choices truly mattered

The runs support the hybrid strategy as a family, but three aggregate scores cannot isolate the causal value of any one solver component.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-003422%46:23$4.4415.68 MB20Valid
run-005618%48:01$4.243.21 MB21Valid
run-005720%49:36$4.8410.30 MB15Valid
499.9KInput tokens
107.6KOutput tokens
15.6MCache-read tokens
01

Seeded evidence

The estimate is based on 3 valid runs spanning 18% to 22%, a 4-point spread.

02

Candidate behavior

The agent made 56 candidate submission actions across the published run set. Selected artifacts averaged 9.73 MB compressed.

03

Time use

The runs averaged 48:00 of wall time. 3 runs explicitly finalized before the one-hour limit.

Retained totals: 288 agent messages · 56 candidate actions · 3 published runs.

06 · Economics

What the result cost

$4.51 per displayed run at current configured rates.

At current rates, GPT-5.5 xhigh sits in the middle of the value table. Its strongest economic advantage is not raw cheapness; it is that the price buys the library's best-seeded evidence.

At current configured API-equivalent rates, the displayed run cost is $4.51 and value is 4.44 score points per dollar.

#8
Value positionValue rank #8 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-07-30.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GPT-5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard