Chose solver breadth over training
No training run was retained; effort went into direct inference and packaging from the beginning.
Run postmortem · Published 2026-08-14
Nineteen candidates, no recorded training run, and a 19% solver built almost entirely in inference code.
One-hour track · xhigh reasoning · provisional
DeepSeek V4 Flash spent nearly the full hour expanding and hardening a broad deterministic inference system rather than training a large fallback. The 19% hidden score is the strongest generated one-seed result below Sol and GPT-5.5 xhigh, and its $1.58 cost gives it excellent current value.
One valid autonomous run produced a provisional 19% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The artifact treated each prefix as a structured reasoning problem and solved it directly when possible.
The solver handled copying, arithmetic, sequences, mappings, records, lists, text transformations, and passage extraction.
Later patches deliberately searched recent and full context to avoid missing repeated records or mappings.
Thirteen packaging actions and nineteen submissions kept the solver under continuous artifact-level validation.
No training run was retained; effort went into direct inference and packaging from the beginning.
The agent identified repeated-record and mapping misses, broadened the search window, and reached perfect calibration.
The resulting 5.35 MB artifact scored 19% hidden.
DeepSeek converted almost the entire hour into broader solver coverage and a near-frontier one-seed result.
Once the public set reached 100%, further special-case additions could not be ranked for hidden generalization.
A second and third seed are needed to determine whether 19% is stable or a favorable solver path.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0078 | 19% | 57:38 | $1.58 | 5.35 MB | 19 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 19 candidate submission actions across the published run set. Selected artifacts averaged 5.35 MB compressed.
The runs averaged 57:38 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 232 agent messages · 19 candidate actions · 1 published run.
$1.58 per displayed run at current configured rates.
DeepSeek's $1.58 run delivered 12.04 points per dollar, placing it among the best current values without relying on later repricing.
At current configured API-equivalent rates, the displayed run cost is $1.58 and value is 12.04 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-08-01.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare DeepSeek V4 Flash 0731 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard