Tooling interrupted the synthetic pipeline
A planned generator command failed in the working environment; the run still produced a valid 10.56 MB artifact and scored 2%.
Run postmortem · Published 2026-08-14
Two full-hour attempts produced 2% and 0%—and the final artifacts explain why.
One-hour track · xhigh reasoning · provisional
Qwen 3.8 Max attempted synthetic-data generation and task-aware fixes, but both final artifacts retained the minimal baseline inference wrapper. The two valid hidden scores—2% and 0%—are therefore consistent with what was submitted: mostly ordinary neural next-byte prediction, not the broad contextual solver used by higher-scoring models.
2 valid autonomous runs averaged 1%; one more seed is required for an official estimate.
2 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The run logs show ambitious work, but the selected artifacts remained conventional.
The agent tried to create task-shaped examples and repair generator bugs during both long runs.
One seed recorded a short four-layer training pass; the other retained no training trajectory beyond the packaged model.
Both submitted infer.py files simply loaded the model and returned neural probabilities, without a broad deterministic solver.
A planned generator command failed in the working environment; the run still produced a valid 10.56 MB artifact and scored 2%.
The agent debugged indexing and exact-output cases, but the selected final package did not contain that broad solver work.
The independent 2% and 0% results make a grader anomaly unlikely and point back to artifact selection and execution.
Two valid seeds now show that the initial low score was not just a single unlucky grade.
The decisive failure was operational: the final packages stayed near baseline despite substantial task-aware experimentation.
These runs do not answer how the model would score if its attempted solver and synthetic-data work were correctly integrated and selected.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0082 | 2% | 60:25 | $2.47 | 10.56 MB | 6 | Valid |
| run-0083 | 0% | 61:56 | $2.55 | 3.32 MB | 6 | Valid |
The valid scores span 0% to 2%. A third run is required for an official estimate.
The agent made 12 candidate submission actions across the published run set. Selected artifacts averaged 6.94 MB compressed.
The runs averaged 61:10 of wall time. 2 runs explicitly finalized before the one-hour limit.
Retained totals: 129 agent messages · 12 candidate actions · 2 published runs.
$2.51 per displayed run at current configured rates.
The two attempts averaged $2.51 each and more than sixty-one minutes of wall time. The poor value reflects failed conversion of agent work into the official artifact, not merely model pricing.
At current configured API-equivalent rates, the displayed run cost is $2.51 and value is 0.4 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Qwen 3.8 Max with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard