Run postmortem · Published 2026-08-14

GPT-5.5high reasoning

It built the biggest neural candidates of the GPT-5.5 sweep—and gained only one point over medium.

11/ 100
hidden exact

One-hour track · high reasoning · provisional

ARI score scaleexact-match accuracy
11%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#13
Current cost$2.52
Points / $4.36
Average time46:43
01 · Verdict

Bigger training was not the missing ingredient

GPT-5.5 high invested most of its run in progressively larger neural candidates, including eight- and ten-layer variants. The agent made a sensible model-selection decision from sampled loss, but the final 11% hidden score was only one point above medium effort.

One valid autonomous run produced a provisional 11% hidden exact-match score.

#13
Score position11% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GPT-5.5 built

This run leaned harder into model scale while retaining a substantial deterministic inference layer.

01

Progressive depth search

Five training attempts moved from a small three-layer model to much deeper candidates near the artifact limit.

02

Loss-based fallback selection

The agent compared the largest models and deliberately returned to the smaller eight-layer candidate when it looked safer on sampled loss.

03

Broad deterministic wrapper

Hundreds of lines of parsing logic handled records, mappings, schedules, transformations, and extraction before neural fallback.

03 · Trajectory

How the run unfolded

Opening

Scaled aggressively

The run treated capacity as the primary lever and trained successively deeper models.

Late middle

Rejected the largest candidate

The ten-layer model fit, but the eight-layer package looked smaller, faster, and slightly better on sampled loss.

46:43

Finalized below the hidden frontier

Despite perfect public calibration and seventeen submissions, the hidden exam settled at 11%.

04 · Assessment

Good, bad, and unresolved

What worked

Disciplined model selection

The agent did not equate the largest network with the best candidate and explicitly chose the safer fallback.

What hurt

Capacity barely moved exact-match accuracy

A much larger artifact and longer run produced only a one-point gain over medium effort.

What remains unknown

Neural versus solver contribution

Aggregate grading cannot separate which hidden answers came from the trained fallback and which came from deterministic logic.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-003211%46:43$2.5212.22 MB17Valid
84KInput tokens
21.8KOutput tokens
2.9MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 17 candidate submission actions across the published run set. Selected artifacts averaged 12.22 MB compressed.

03

Time use

The runs averaged 46:43 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 80 agent messages · 17 candidate actions · 1 published run.

06 · Economics

What the result cost

$2.52 per displayed run at current configured rates.

High effort used almost the full hour and cost more than medium, but its marginal gain was only one point. In this run, extra model scale was a poor return on spend.

At current configured API-equivalent rates, the displayed run cost is $2.52 and value is 4.36 score points per dollar.

#11
Value positionValue rank #11 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-07-30.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GPT-5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard