Committed to model training
The run treated neural capacity as the primary path to solving the benchmark.
Run postmortem · Published 2026-08-14
A near-cap neural artifact passed public calibration and still scored 7% hidden.
One-hour track · xhigh reasoning · provisional
Claude Fable 5 spent fifty-six minutes building and packaging a large conventional byte model. Its retained checkpoint reached perfect public calibration, but the final inference path remained minimal and the hidden exam returned 7%.
One valid autonomous run produced a provisional 7% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
Fable committed to trained-model capacity rather than a broad deterministic solver.
The final artifact was 13.62 MB, close to the fixed size ceiling after compression.
The submitted inference path mostly loaded the model and returned neural next-byte probabilities.
The run iterated cautiously compared with the solver-heavy systems that submitted dozens of packages.
The run treated neural capacity as the primary path to solving the benchmark.
The selected candidate fit the visible set despite lacking a large direct-solving layer.
The public-perfect checkpoint transferred poorly to the private exam.
The run successfully trained, packaged, and finalized a large model within the strict artifact limit.
Without broad inference logic, the large model reproduced only seven hidden answers exactly.
This result tests one mostly neural strategy, not the model's ability to build the solver-oriented systems that lead the board.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0047 | 7% | 56:06 | $8.14 | 13.62 MB | 5 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 13.62 MB compressed.
The runs averaged 56:06 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 54 agent messages · 5 candidate actions · 1 published run.
$8.14 per displayed run at current configured rates.
At more than eight dollars for seven points, Fable is one of the least efficient valid results. Most of the spend supported neural training and long-context agent work.
At current configured API-equivalent rates, the displayed run cost is $8.14 and value is 0.86 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Claude Fable 5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard