Run postmortem · Published 2026-09-04

Muse Spark 1.3xhigh reasoning

The intended artifact was hidden behind a packaging failure; an exact hash-matched regrade recovered a 13% result.

13/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
13%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#18
Current cost$1.13
Points / $11.46
Average time58:07
01 · Verdict

The original zero was a harness artifact, not the model's result

Muse Spark 1.3 built a trained byte-model fallback with a broad deterministic inference layer. The main run accidentally packaged the starter inference code instead of the promoted candidate and recorded 0/100. A post-run audit found the mismatch, repackaged the candidate that had been promoted within the time limit, verified its substantive files against the source by hash, and graded that corrected bundle in the same no-network Docker evaluator at 13/100. ARI Bench publishes the corrected regrade and excludes the packaging-induced zero from the leaderboard.

One valid autonomous run produced a provisional 13% hidden exact-match score.

#18
Score position13% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Muse Spark 1.3 built

Muse spent nearly the full hour combining a learned fallback with direct contextual problem solving, while keeping multiple valid candidates available for timeout recovery.

01

Trained byte-model fallback

The retained package included a compressed neural checkpoint and its matching configuration as a fallback for cases outside the direct solver.

02

Broad deterministic solver

The intended inference path added direct handlers for arithmetic, mappings, structured records, sequences, transformations, and exact output formatting.

03

Checkpoint-first candidate promotion

Six candidate submissions were recorded, and the candidate used for the corrected regrade had been promoted before the run timed out.

03 · Trajectory

How the run unfolded

Opening

Established a gradeable checkpoint

The run trained and packaged an early candidate before widening its deterministic coverage.

Iteration

Expanded and repeatedly promoted the hybrid

Later candidates retained the same checkpoint while the source inference layer grew substantially; local synthetic validation reached 280/300 and boundary checks reached 75/75.

Post-run audit

Recovered the intended candidate

The automatic package contained the starter inference file. Repackaging the within-time promoted source and rerunning the unchanged hidden grader produced the corrected 13% result.

04 · Assessment

Good, bad, and unresolved

What worked

A recoverable within-time artifact

The promoted candidate existed before timeout, stayed under the 16 MB cap, passed import and autoregressive checks, and matched its substantive source files by hash after correction.

What hurt

The harness selected the wrong editable code

The packaging path replaced the candidate's inference implementation with starter code, creating a false zero despite a viable retained artifact.

What remains unknown

Repeatability after one corrected seed

The 13% score is a defensible corrected regrade, but two additional valid autonomous runs are still required for an official estimate.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-011113%58:07$1.139.49 MB6Valid
278.9KInput tokens
43.6KOutput tokens
4MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 6 candidate submission actions across the published run set. Selected artifacts averaged 9.49 MB compressed.

03

Time use

The runs averaged 58:07 of wall time. Runtime alone cannot show whether more time would have improved the result.

Retained totals: 86 agent messages · 6 candidate actions · 1 published run.

06 · Economics

What the result cost

$1.13 per displayed run at current configured rates.

The scored run's archived token buckets imply $1.13 at the September 4 OpenRouter rates. A separate network-lost attempt cost about $0.05 and is excluded from the per-run leaderboard metric, bringing total campaign spend to about $1.18.

At current configured API-equivalent rates, the displayed run cost is $1.13 and value is 11.46 score points per dollar.

#7
Value positionValue rank #7 of 29 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-04.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Muse Spark 1.3 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard