Run postmortem · Published 2026-08-14

Claude Opus 4.8xhigh reasoning

A 24-minute recovery turned early tool friction into a 16% hybrid.

16/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
16%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#8
Current cost$7.96
Points / $2.01
Average time23:54
01 · Verdict

Strong recovery, costly early stop

Claude Opus 4.8 encountered friction accessing the working data path, recovered, and still produced a valid 16% hybrid in under 24 minutes. The result is competitive for a single seed, but the run stopped with more than half of the hour unused and carried a high API-equivalent cost.

One valid autonomous run produced a provisional 16% hidden exact-match score.

#8
Score position16% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Claude Opus 4.8 built

Once oriented, Opus combined a conventional trained model with a broad induction and extraction layer.

01

Two same-shape training passes

It trained the same four-layer, 256-wide architecture for 1,500 and then 2,500 steps.

02

Context-first solver

The inference path parsed mappings and records, induced answers from nearby examples, and fell back to the neural model.

03

Frequent valid checkpoints

Thirteen candidate submissions kept the evolving system gradeable despite the rushed schedule.

03 · Trajectory

How the run unfolded

Opening

Lost time to tool friction

The agent initially struggled to access the intended data and command path.

Recovery

Built the hybrid quickly

It trained two checkpoints and expanded the inference layer once the environment was understood.

23:54

Finalized early

The valid artifact scored 16% with roughly thirty-six minutes still available.

04 · Assessment

Good, bad, and unresolved

What worked

Rapid recovery and a competitive score

The run converted a confused opening into a valid upper-midfield result without approaching the deadline.

What hurt

High cost for one seed

The displayed run cost is close to Sol's despite a materially lower score and no replication.

What remains unknown

Whether more time would help

The early stop leaves unusually large uncertainty about the ceiling of the chosen approach.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-003616%23:54$7.9612.24 MB13Valid
312Input tokens
70.6KOutput tokens
11.2MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 13 candidate submission actions across the published run set. Selected artifacts averaged 12.24 MB compressed.

03

Time use

The runs averaged 23:54 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 163 agent messages · 13 candidate actions · 1 published run.

06 · Economics

What the result cost

$7.96 per displayed run at current configured rates.

Opus 4.8 paid a premium for a short run. Its score was solid, but the combination of high cost and one seed makes its value ranking fragile.

At current configured API-equivalent rates, the displayed run cost is $7.96 and value is 2.01 score points per dollar.

#16
Value positionValue rank #16 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Claude Opus 4.8 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard