Run postmortem · Published 2026-08-14

Claude Sonnet 5xhigh reasoning

Ten submissions preserved the starter shape—and the hidden score stayed at 1%.

1/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
1%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#19
Current cost$2.01
Points / $0.5
Average time50:00
01 · Verdict

Valid execution, almost no task adaptation

Claude Sonnet 5 produced a valid 9.25 MB artifact after fifty minutes, but the final inference path remained close to the baseline neural pipeline. With little task-specific solving and only 4% public calibration, the resulting 1% hidden score is consistent with the submitted system rather than a grading surprise.

One valid autonomous run produced a provisional 1% hidden exact-match score.

#19
Score position1% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Claude Sonnet 5 built

The final package relied overwhelmingly on ordinary neural next-byte prediction.

01

Conventional byte transformer

The artifact kept the standard GPT-style model and a minimal inference wrapper.

02

No broad deterministic layer

Unlike the higher-scoring hybrids, the final inference path did not encode substantial arithmetic, record, mapping, or extraction logic.

03

Repeated packaging

Ten candidate submissions maintained validity but did not materially change the artifact's problem-solving shape.

03 · Trajectory

How the run unfolded

Opening

Stayed near the starter architecture

The run focused on training parameters rather than reframing the benchmark as contextual problem solving.

Middle

Public accuracy remained near zero

The retained best calibration score was only 4%, leaving clear evidence that the system was not ready.

50:00

Finalized a valid but weak artifact

The hidden grader found one exact answer out of 100.

04 · Assessment

Good, bad, and unresolved

What worked

Operational validity

The run respected the artifact boundary, submitted repeatedly, and produced a gradeable final package.

What hurt

The core strategy did not match the exam

A plain byte model was a poor fit for exact contextual completions that rewarded direct inference.

What remains unknown

Whether a tool-competent rerun would pivot

This seed does not show what Sonnet 5 could do after recognizing and implementing the hybrid strategy used by stronger models.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-00411%50:00$2.019.25 MB10Valid
128Input tokens
46.2KOutput tokens
3.1MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 10 candidate submission actions across the published run set. Selected artifacts averaged 9.25 MB compressed.

03

Time use

The runs averaged 50:00 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 69 agent messages · 10 candidate actions · 1 published run.

06 · Economics

What the result cost

$2.01 per displayed run at current configured rates.

Two dollars bought a valid run but almost no score. The low efficiency reflects the artifact's strategy, not merely expensive token pricing.

At current configured API-equivalent rates, the displayed run cost is $2.01 and value is 0.5 score points per dollar.

#20
Value positionValue rank #20 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Claude Sonnet 5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard