Run postmortem · Published 2026-09-11

DeepSeek V4.1 Flashxhigh reasoning

Perfect public calibration, 24% hidden accuracy, and a complete run bill under forty cents.

24/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
24%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#3
Current cost$0.40
Points / $60.59
Average time55:53
01 · Verdict

A stronger public solver still left a large hidden gap

DeepSeek V4.1 Flash completed a valid 55:53 retest and scored 24/100 on the hidden exam. Its public calibration improved from 20/25 to 25/25 as the agent revised its inference code. This page reports the selected September 11 retest, not an average across all attempts. It remains provisional and is not a direct test of DeepSeek answering the exam itself.

One valid autonomous run produced a provisional 24% hidden exact-match score.

#3
Score position24% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What DeepSeek V4.1 Flash built

The agent built a small trained byte-level model and surrounded it with deterministic task solvers.

01

A trained neural fallback

The selected six-layer, 256-wide model contains 4,870,144 parameters. The complete compressed artifact is 9,062,942 bytes, below the 16 MB cap.

02

Task solvers packaged with the model

The agent embedded its deterministic solver into the final inference file and combined it with the trained model. A separate exact-package audit confirmed that both the solver and learned fallback were present.

03

Explicit hosted xhigh request

OpenRouter routed the run to Novita FP8 with xhigh explicitly requested and no provider fallbacks. The endpoint's internal reasoning-effort mapping is not independently observable; this run is not labelled max.

03 · Trajectory

How the run unfolded

Training

Built a checkpoint, then continued training

The first main training phase ran for 15 minutes and reached step 47,743. An 11-minute continuation reached step 35,001; its checkpoint was the one selected for the final artifact.

Iteration

Public calibration reached 25/25

Actual submitted candidates improved from 20/25 to 25/25 after inference revisions. Eleven submissions and two finalizations were recorded. The agent also tested newly generated synthetic examples; these were not the private exam.

55:53

Finalized early with 24/100 hidden accuracy

The run finished before its deadline with no provider-error messages in the complete session export. The original hidden Docker grade was valid. No operator code fixes, candidate hints, packaging correction, or regrade were applied.

04 · Assessment

Good, bad, and unresolved

What worked

Packaging and execution were verified

The final code, configuration and weights matched the selected candidate and source. The exact package loaded and exercised its trained checkpoint and passed distribution, determinism and short autoregressive contract checks.

What hurt

Perfect calibration did not mean broad generalization

The selected artifact solved all 25 public calibration questions but only 24 of 100 hidden questions. That gap is observed; it does not by itself identify which solver rules or model components caused the remaining errors.

What remains unknown

One selected retest is not a stable estimate

This selected run does not measure run-to-run variance or isolate the benefit of either training phase. No neural-only or solver-only hidden ablation was performed. The displayed result is provisional, not an official three-seed estimate.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-011924%55:53$0.409.06 MB13Valid
145.1KInput tokens
73.7KOutput tokens
17.71MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 13 candidate submission actions across the published run set. Selected artifacts averaged 9.06 MB compressed.

03

Time use

The runs averaged 55:53 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 208 agent messages · 13 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.40 per displayed run at current configured rates.

The complete run-window charge was $0.396131160 USD, including reasoning and auxiliary calls, measured from the provider-reported usage delta on the benchmark key. The final two post-grade snapshots were unchanged roughly one minute apart. The table displays $0.40 and 60.6 score points per dollar. The complete primary-model usage records contain 73,710 visible output tokens plus 129,354 reasoning tokens; the output-token strip above lists visible output only. Primary-model API-equivalent usage is $0.393488160; the $0.002643 residual is not independently itemized. No other benchmark ran on the rig, though off-rig key use cannot be ruled out by this counter alone. Preflights and other attempts are excluded from this completed-run cost.

At current configured API-equivalent rates, the displayed run cost is $0.40 and value is 60.59 score points per dollar.

#3
Value positionValue rank #3 of 31 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-11.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare DeepSeek V4.1 Flash with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard