Built a checkpoint, then continued training
The first main training phase ran for 15 minutes and reached step 47,743. An 11-minute continuation reached step 35,001; its checkpoint was the one selected for the final artifact.
Run postmortem · Published 2026-09-11
Perfect public calibration, 24% hidden accuracy, and a complete run bill under forty cents.
One-hour track · xhigh reasoning · provisional
DeepSeek V4.1 Flash completed a valid 55:53 retest and scored 24/100 on the hidden exam. Its public calibration improved from 20/25 to 25/25 as the agent revised its inference code. This page reports the selected September 11 retest, not an average across all attempts. It remains provisional and is not a direct test of DeepSeek answering the exam itself.
One valid autonomous run produced a provisional 24% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The agent built a small trained byte-level model and surrounded it with deterministic task solvers.
The selected six-layer, 256-wide model contains 4,870,144 parameters. The complete compressed artifact is 9,062,942 bytes, below the 16 MB cap.
The agent embedded its deterministic solver into the final inference file and combined it with the trained model. A separate exact-package audit confirmed that both the solver and learned fallback were present.
OpenRouter routed the run to Novita FP8 with xhigh explicitly requested and no provider fallbacks. The endpoint's internal reasoning-effort mapping is not independently observable; this run is not labelled max.
The first main training phase ran for 15 minutes and reached step 47,743. An 11-minute continuation reached step 35,001; its checkpoint was the one selected for the final artifact.
Actual submitted candidates improved from 20/25 to 25/25 after inference revisions. Eleven submissions and two finalizations were recorded. The agent also tested newly generated synthetic examples; these were not the private exam.
The run finished before its deadline with no provider-error messages in the complete session export. The original hidden Docker grade was valid. No operator code fixes, candidate hints, packaging correction, or regrade were applied.
The final code, configuration and weights matched the selected candidate and source. The exact package loaded and exercised its trained checkpoint and passed distribution, determinism and short autoregressive contract checks.
The selected artifact solved all 25 public calibration questions but only 24 of 100 hidden questions. That gap is observed; it does not by itself identify which solver rules or model components caused the remaining errors.
This selected run does not measure run-to-run variance or isolate the benefit of either training phase. No neural-only or solver-only hidden ablation was performed. The displayed result is provisional, not an official three-seed estimate.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0119 | 24% | 55:53 | $0.40 | 9.06 MB | 13 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 13 candidate submission actions across the published run set. Selected artifacts averaged 9.06 MB compressed.
The runs averaged 55:53 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 208 agent messages · 13 candidate actions · 1 published run.
$0.40 per displayed run at current configured rates.
The complete run-window charge was $0.396131160 USD, including reasoning and auxiliary calls, measured from the provider-reported usage delta on the benchmark key. The final two post-grade snapshots were unchanged roughly one minute apart. The table displays $0.40 and 60.6 score points per dollar. The complete primary-model usage records contain 73,710 visible output tokens plus 129,354 reasoning tokens; the output-token strip above lists visible output only. Primary-model API-equivalent usage is $0.393488160; the $0.002643 residual is not independently itemized. No other benchmark ran on the rig, though off-rig key use cannot be ruled out by this counter alone. Preflights and other attempts are excluded from this completed-run cost.
At current configured API-equivalent rates, the displayed run cost is $0.40 and value is 60.59 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-09-11.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare DeepSeek V4.1 Flash with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard