Run postmortem · Published 2026-09-01

Claude Fable 5.1xhigh reasoning

A 96% public calibration result became 19% on the hidden exam.

19/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
19%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#5
Current cost$10.55
Points / $1.8
Average time52:45
01 · Verdict

The hybrid improved sharply, but public calibration still overstated transfer

Claude Fable 5.1 built a near-cap hybrid that combined a trained byte transformer with a broad deterministic solver. The finalized artifact answered 24 of 25 public calibration questions but only 19 of 100 hidden questions, showing both a meaningful gain over the prior Fable 5 run and a severe limit in the visible feedback signal.

One valid autonomous run produced a provisional 19% hidden exact-match score.

#5
Score position19% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Claude Fable 5.1 built

Fable 5.1 split the hour between synthetic training and direct contextual problem solving.

01

15.55M-parameter byte model

After short smoke tests, the run trained two substantial checkpoints and retained a compressed artifact of 14,485,395 bytes.

02

Broad solver overlay

The inference layer added handlers for arithmetic, sequences, mappings, records, schedules, text transformations, and odd-one-out tasks before neural fallback.

03

Exact packaged-candidate checks

The retained candidate was finalized explicitly, and its executable files match the source and stored candidate byte for byte.

03 · Trajectory

How the run unfolded

Opening

Validated the training path

The run used short smoke attempts to establish a working 15.55M-parameter byte model before committing to longer training.

Middle

Expanded the hybrid

Two longer training runs were paired with increasingly broad deterministic inference and repeated candidate packaging.

52:45

Finalized at 96% public and 19% hidden

The selected 14.49 MB artifact finished early after seven candidate submissions; the private exam exposed a 77-point public-to-hidden gap.

04 · Assessment

Good, bad, and unresolved

What worked

A competitive provisional improvement

The 19% hidden score is twelve points above the earlier Fable 5 seed and came from a valid, reproducibly packaged hybrid.

What hurt

Calibration was not a reliable selector

Solving 96% of the visible set did not distinguish a system that still missed 81% of the hidden questions.

What remains unknown

Whether 19% survives replication

This is one autonomous seed. Two more valid runs are required before treating the result as an official model estimate.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-010819%52:45$10.5514.49 MB7Valid
133.6KInput tokens
69.7KOutput tokens
9.9MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 7 candidate submission actions across the published run set. Selected artifacts averaged 14.49 MB compressed.

03

Time use

The runs averaged 52:45 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 112 agent messages · 7 candidate actions · 1 published run.

06 · Economics

What the result cost

$10.55 per displayed run at current configured rates.

OpenRouter billed $10.55 to the dedicated benchmark key, including a $0.00039 preflight; the run-local token estimate was $7.30. That makes the observed value 1.80 hidden-score points per billed dollar.

At current configured API-equivalent rates, the displayed run cost is $10.55 and value is 1.8 score points per dollar.

#20
Value positionValue rank #20 of 27 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-01.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Claude Fable 5.1 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard