Run postmortem · Published 2026-08-14

GPT-5.5low reasoning

It found the right idea, then stopped before it was broad enough.

4/ 100
hidden exact

One-hour track · low reasoning · provisional

ARI score scaleexact-match accuracy
4%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#17
Current cost$0.91
Points / $4.39
Average time19:06
01 · Verdict

A useful sketch, not a finished system

GPT-5.5 low quickly recognized that a neural checkpoint alone would not be enough and moved toward a hybrid predictor. The run produced a valid artifact and perfect public calibration, but its 4% hidden score shows that the early solver remained narrow.

One valid autonomous run produced a provisional 4% hidden exact-match score.

#17
Score position4% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GPT-5.5 built

The run combined a trained byte model with task-aware inference logic instead of relying on either one alone.

01

Timeout-safe neural checkpoint

It trained and promoted a valid model almost immediately, protecting the run from ending without a gradeable artifact.

02

Hybrid inference layer

It added contextual copying, extraction, arithmetic, sequence, and record-oriented rules around the neural fallback.

03

Model scaling under a small budget

Six training attempts explored larger configurations, but the final artifact still had to fit the fixed 16 MB cap.

03 · Trajectory

How the run unfolded

Opening

Secured a checkpoint

The first trained candidate was promoted as insurance before the more ambitious inference work began.

Middle

Broadened the solver

The agent layered deterministic completions over the trained model and repeatedly resubmitted candidates.

19:05

Finalized very early

The run ended with roughly forty minutes unused. Public calibration was saturated, but hidden coverage remained thin.

04 · Assessment

Good, bad, and unresolved

What worked

Fast recognition of the real task

The run did not stay trapped in ordinary next-byte training; it pivoted toward contextual problem solving.

What hurt

Public success hid a narrow solver

A 25/25 public result became 4/100 on the hidden exam, the clearest sign that the visible task family was overfit.

What remains unknown

What forty more minutes could buy

More time might have widened the solver, but the public set no longer offered a useful ranking signal.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-00304%19:06$0.9114.03 MB11Valid
54.4KInput tokens
10.5KOutput tokens
649.2KCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 11 candidate submission actions across the published run set. Selected artifacts averaged 14.03 MB compressed.

03

Time use

The runs averaged 19:06 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 34 agent messages · 11 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.91 per displayed run at current configured rates.

The low-effort run was inexpensive and fast, but the saved spend came with a large accuracy tradeoff. Its value ratio looks ordinary only because both cost and score were small.

At current configured API-equivalent rates, the displayed run cost is $0.91 and value is 4.39 score points per dollar.

#10
Value positionValue rank #10 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-07-30.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GPT-5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard