Spent most of the budget on training
Repeated data-generation and transformer-training experiments delayed the first official candidate until roughly nineteen minutes remained.
Run postmortem · Published 2026-08-14
A local 25/25, a harness-measured 12%, and a decision to finalize the mismatch.
One-hour track · max reasoning · provisional
DeepSeek V4 Pro did not fail because the run crashed or the grader rejected its work. Its first submitted candidates reached 16% public calibration, later revisions fell to 12%, and the agent finalized the regression because its own local evaluator still reported 25/25. The resulting artifact scored 4% hidden—far below V4 Flash's 19%—making this a concrete candidate-engineering and selection failure rather than evidence of a broken grade.
One valid autonomous run produced a provisional 4% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The run pursued an ambitious hybrid, but the expensive pieces never converged into a reliably packaged solver.
The agent trained a six-layer byte transformer on ordinary text and synthetic task examples, producing a final artifact just over 10 MB compressed.
It wrapped the model with parsers intended to answer structured completions directly when a recognizable pattern appeared.
A custom local test said 25/25 while the runner's packaged-artifact calibration remained between 12% and 16%. The discrepancy was investigated but never resolved.
Repeated data-generation and transformer-training experiments delayed the first official candidate until roughly nineteen minutes remained.
The first two valid packaged candidates each answered 4 of 25 public calibration questions exactly.
Contract changes reduced measured calibration to 3 of 25. Despite the unresolved conflict with its local 25/25 result, the agent promoted and finalized that branch; it scored 4 of 100 hidden.
The session stayed pinned, tools remained usable, nine candidate actions completed, and the selected artifact passed size and grading checks without recorded broker errors.
The decisive mistake was trusting a custom local evaluator after the actual packaged artifact had visibly regressed. V4 Flash encountered a similar mismatch but kept debugging until both paths agreed.
This is one seed, and the allowed OpenRouter pool did not retain per-request provider attribution. Additional seeds are needed before treating 4% as the model's stable capability.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0091 | 4% | 60:48 | $3.38 | 10.09 MB | 9 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 9 candidate submission actions across the published run set. Selected artifacts averaged 10.09 MB compressed.
The runs averaged 60:48 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 285 agent messages · 9 candidate actions · 1 published run.
$3.38 per displayed run at current configured rates.
The run cost an estimated $3.38, or 1.18 score points per dollar. V4 Flash produced nearly five times the score for less than half the cost because it converted more of the hour into verified inference hardening.
At current configured API-equivalent rates, the displayed run cost is $3.38 and value is 1.18 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-07-30.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare DeepSeek V4 Pro 0813 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard