Secured a checkpoint
The first trained candidate was promoted as insurance before the more ambitious inference work began.
Run postmortem · Published 2026-08-14
It found the right idea, then stopped before it was broad enough.
One-hour track · low reasoning · provisional
GPT-5.5 low quickly recognized that a neural checkpoint alone would not be enough and moved toward a hybrid predictor. The run produced a valid artifact and perfect public calibration, but its 4% hidden score shows that the early solver remained narrow.
One valid autonomous run produced a provisional 4% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The run combined a trained byte model with task-aware inference logic instead of relying on either one alone.
It trained and promoted a valid model almost immediately, protecting the run from ending without a gradeable artifact.
It added contextual copying, extraction, arithmetic, sequence, and record-oriented rules around the neural fallback.
Six training attempts explored larger configurations, but the final artifact still had to fit the fixed 16 MB cap.
The first trained candidate was promoted as insurance before the more ambitious inference work began.
The agent layered deterministic completions over the trained model and repeatedly resubmitted candidates.
The run ended with roughly forty minutes unused. Public calibration was saturated, but hidden coverage remained thin.
The run did not stay trapped in ordinary next-byte training; it pivoted toward contextual problem solving.
A 25/25 public result became 4/100 on the hidden exam, the clearest sign that the visible task family was overfit.
More time might have widened the solver, but the public set no longer offered a useful ranking signal.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0030 | 4% | 19:06 | $0.91 | 14.03 MB | 11 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 11 candidate submission actions across the published run set. Selected artifacts averaged 14.03 MB compressed.
The runs averaged 19:06 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 34 agent messages · 11 candidate actions · 1 published run.
$0.91 per displayed run at current configured rates.
The low-effort run was inexpensive and fast, but the saved spend came with a large accuracy tradeoff. Its value ratio looks ordinary only because both cost and score were small.
At current configured API-equivalent rates, the displayed run cost is $0.91 and value is 4.39 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-07-30.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare GPT-5.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard