Run postmortem · Published 2026-08-14

Claude Fable 5xhigh reasoning

A near-cap neural artifact passed public calibration and still scored 7% hidden.

7/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
7%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#16
Current cost$8.14
Points / $0.86
Average time56:06
01 · Verdict

The pure-neural bet did not transfer

Claude Fable 5 spent fifty-six minutes building and packaging a large conventional byte model. Its retained checkpoint reached perfect public calibration, but the final inference path remained minimal and the hidden exam returned 7%.

One valid autonomous run produced a provisional 7% hidden exact-match score.

#16
Score position7% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Claude Fable 5 built

Fable committed to trained-model capacity rather than a broad deterministic solver.

01

Large neural package

The final artifact was 13.62 MB, close to the fixed size ceiling after compression.

02

Minimal inference wrapper

The submitted inference path mostly loaded the model and returned neural next-byte probabilities.

03

Five retained candidates

The run iterated cautiously compared with the solver-heavy systems that submitted dozens of packages.

03 · Trajectory

How the run unfolded

Opening

Committed to model training

The run treated neural capacity as the primary path to solving the benchmark.

Late run

Reached perfect calibration

The selected candidate fit the visible set despite lacking a large direct-solving layer.

56:06

Hidden accuracy fell to 7%

The public-perfect checkpoint transferred poorly to the private exam.

04 · Assessment

Good, bad, and unresolved

What worked

A valid near-cap artifact

The run successfully trained, packaged, and finalized a large model within the strict artifact limit.

What hurt

Public fit did not equal contextual reasoning

Without broad inference logic, the large model reproduced only seven hidden answers exactly.

What remains unknown

What a hybrid Fable run would do

This result tests one mostly neural strategy, not the model's ability to build the solver-oriented systems that lead the board.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-00477%56:06$8.1413.62 MB5Valid
106Input tokens
46.7KOutput tokens
2.7MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 13.62 MB compressed.

03

Time use

The runs averaged 56:06 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 54 agent messages · 5 candidate actions · 1 published run.

06 · Economics

What the result cost

$8.14 per displayed run at current configured rates.

At more than eight dollars for seven points, Fable is one of the least efficient valid results. Most of the spend supported neural training and long-context agent work.

At current configured API-equivalent rates, the displayed run cost is $8.14 and value is 0.86 score points per dollar.

#18
Value positionValue rank #18 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Claude Fable 5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard