Run postmortem · Published 2026-08-14

Qwen 3.8 Maxxhigh reasoning

Two full-hour attempts produced 2% and 0%—and the final artifacts explain why.

1/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
1%
0255075100
StatusProvisional
Valid seeds2 / 3
Score rank#19
Current cost$2.51
Points / $0.4
Average time61:10
01 · Verdict

The model did not ship the strategy it was trying to build

Qwen 3.8 Max attempted synthetic-data generation and task-aware fixes, but both final artifacts retained the minimal baseline inference wrapper. The two valid hidden scores—2% and 0%—are therefore consistent with what was submitted: mostly ordinary neural next-byte prediction, not the broad contextual solver used by higher-scoring models.

2 valid autonomous runs averaged 1%; one more seed is required for an official estimate.

#19
Score position1% on the common 0–100 scale. Tied valid scores share a rank.

2 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Qwen 3.8 Max built

The run logs show ambitious work, but the selected artifacts remained conventional.

01

Synthetic-data experiments

The agent tried to create task-shaped examples and repair generator bugs during both long runs.

02

Small neural training

One seed recorded a short four-layer training pass; the other retained no training trajectory beyond the packaged model.

03

Baseline final inference

Both submitted infer.py files simply loaded the model and returned neural probabilities, without a broad deterministic solver.

03 · Trajectory

How the run unfolded

Seed 1

Tooling interrupted the synthetic pipeline

A planned generator command failed in the working environment; the run still produced a valid 10.56 MB artifact and scored 2%.

Seed 2

Patched visible task logic

The agent debugged indexing and exact-output cases, but the selected final package did not contain that broad solver work.

Outcome

Two valid near-zero scores

The independent 2% and 0% results make a grader anomaly unlikely and point back to artifact selection and execution.

04 · Assessment

Good, bad, and unresolved

What worked

The rerun resolved the ambiguity

Two valid seeds now show that the initial low score was not just a single unlucky grade.

What hurt

Work did not reach the submitted artifact

The decisive failure was operational: the final packages stayed near baseline despite substantial task-aware experimentation.

What remains unknown

Qwen with a preserved hybrid

These runs do not answer how the model would score if its attempted solver and synthetic-data work were correctly integrated and selected.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-00822%60:25$2.4710.56 MB6Valid
run-00830%61:56$2.553.32 MB6Valid
1.76MInput tokens
78KOutput tokens
4.1MCache-read tokens
01

Promising, still provisional

The valid scores span 0% to 2%. A third run is required for an official estimate.

02

Candidate behavior

The agent made 12 candidate submission actions across the published run set. Selected artifacts averaged 6.94 MB compressed.

03

Time use

The runs averaged 61:10 of wall time. 2 runs explicitly finalized before the one-hour limit.

Retained totals: 129 agent messages · 12 candidate actions · 2 published runs.

06 · Economics

What the result cost

$2.51 per displayed run at current configured rates.

The two attempts averaged $2.51 each and more than sixty-one minutes of wall time. The poor value reflects failed conversion of agent work into the official artifact, not merely model pricing.

At current configured API-equivalent rates, the displayed run cost is $2.51 and value is 0.4 score points per dollar.

#21
Value positionValue rank #21 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Qwen 3.8 Max with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard