Run postmortem · Published 2026-08-14

GPT-5.6 Lunaxhigh reasoning

A compact 27-minute build became the value outlier after repricing.

17/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
17%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#6
Current cost$0.27
Points / $62.52
Average time27:18
01 · Verdict

The efficiency story is real—but still one seed

GPT-5.6 Luna built a compact hybrid system, finalized in 27 minutes, and scored 17%. The result was already economical at run time; after the model's price cut, the same observed score reprices to 62.52 points per dollar, far ahead of the field.

One valid autonomous run produced a provisional 17% hidden exact-match score.

#6
Score position17% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GPT-5.6 Luna built

Luna kept both the trained model and the inference layer deliberately small.

01

One compact training run

It trained a three-layer, 128-wide fallback for 1,200 steps instead of exploring a large architecture sweep.

02

Practical hybrid solver

The final inference path combined arithmetic, records, extraction, list operations, and exact formatting with neural fallback.

03

Repeated packaging, stable artifact

The agent repackaged the same approximately 2.47 MB system several times while hardening the inference path.

03 · Trajectory

How the run unfolded

Opening

Chose a compact model

The only training run established a small fallback and preserved most of the hour for inference work.

Middle

Reached perfect calibration

The hybrid solved every visible calibration item, then received no further public signal about hidden breadth.

27:18

Finalized with half the hour unused

The 17% hidden result came from a short, valid run rather than a full-hour search.

04 · Assessment

Good, bad, and unresolved

What worked

Exceptional score per current dollar

No other published row combines a double-digit score with a displayed cost below thirty cents.

What hurt

The run left exploration budget unused

Luna stopped after 27 minutes despite having no hidden-facing evidence that the solver was complete.

What remains unknown

Whether the value survives replication

Two more seeds could confirm a genuine efficiency advantage—or reveal that the 17% run was unusually favorable.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-007417%27:18$1.362.47 MB13Valid
198.1KInput tokens
33.6KOutput tokens
9.6MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 13 candidate submission actions across the published run set. Selected artifacts averaged 2.47 MB compressed.

03

Time use

The runs averaged 27:18 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 148 agent messages · 13 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.27 per displayed run at current configured rates.

The archived run cost was $1.36. The current $0.27 figure is a transparent repricing of the same token usage, not a cheaper rerun or a changed benchmark score.

At current configured API-equivalent rates, the displayed run cost is $0.27 and value is 62.52 score points per dollar. The archived cost at run time was $1.36; the difference is repricing, not a changed score.

#1
Value positionValue rank #1 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-07-30.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GPT-5.6 Luna with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard