Run postmortem · Published 2026-08-14

GPT-5.5medium reasoning

More iteration doubled the score, but public saturation still arrived too soon.

10/ 100
hidden exact

One-hour track · medium reasoning · provisional

ARI score scaleexact-match accuracy
10%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#15
Current cost$2.28
Points / $4.39
Average time27:55
01 · Verdict

A better hybrid, still calibrated to the visible surface

GPT-5.5 medium improved sharply over the low-effort run, reaching 10% while exploring several model sizes and a much richer inference layer. It again reached perfect public calibration and finalized early, leaving the hidden exam to expose the remaining coverage gap.

One valid autonomous run produced a provisional 10% hidden exact-match score.

#15
Score position10% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GPT-5.5 built

The medium run spent more of its budget comparing model scales and packaging a compact hybrid system.

01

Five training scales

It moved from a small baseline through wider and deeper candidates rather than betting on a single training configuration.

02

Structured local solver

The inference path added record parsing, forced answers, mapping logic, and local context extraction.

03

Frequent packaging

Fifteen packaging actions turned the evolving solver into gradeable checkpoints instead of treating packaging as an end-of-run task.

03 · Trajectory

How the run unfolded

Opening

Tested the neural baseline

The run established a small model and used it as a fallback while evaluating broader configurations.

Middle

Shifted toward synthetic and local logic

The system combined task-shaped training with explicit parsing and repeatedly checked packaged candidates.

27:53

Stopped at a perfect public score

The final candidate had saturated calibration, yet its hidden score plateaued at 10%.

04 · Assessment

Good, bad, and unresolved

What worked

A real gain over low effort

The hidden score rose from 4% to 10%, showing that additional planning and iteration produced measurable benefit.

What hurt

The selector failed after saturation

Once public calibration reached 100%, later candidate choices could not be ordered by the only visible accuracy signal.

What remains unknown

Whether the right candidate was chosen

Eighteen submissions show active search, but the archive cannot prove that the final public-perfect package was the best hidden generalizer.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-003110%27:55$2.284.91 MB18Valid
95.3KInput tokens
21.7KOutput tokens
2.3MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 18 candidate submission actions across the published run set. Selected artifacts averaged 4.91 MB compressed.

03

Time use

The runs averaged 27:55 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 65 agent messages · 18 candidate actions · 1 published run.

06 · Economics

What the result cost

$2.28 per displayed run at current configured rates.

Medium effort cost about two and a half times the low run and delivered six additional score points. It was a meaningful improvement, but not yet an efficient frontier result.

At current configured API-equivalent rates, the displayed run cost is $2.28 and value is 4.39 score points per dollar.

#9
Value positionValue rank #9 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-07-30.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GPT-5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard