Run postmortem · Published 2026-08-14

DeepSeek V4 Pro 0813max reasoning

A local 25/25, a harness-measured 12%, and a decision to finalize the mismatch.

4/ 100
hidden exact

One-hour track · max reasoning · provisional

ARI score scaleexact-match accuracy
4%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#17
Current cost$3.38
Points / $1.18
Average time60:48
01 · Verdict

It finalized an artifact the harness had already disproved

DeepSeek V4 Pro did not fail because the run crashed or the grader rejected its work. Its first submitted candidates reached 16% public calibration, later revisions fell to 12%, and the agent finalized the regression because its own local evaluator still reported 25/25. The resulting artifact scored 4% hidden—far below V4 Flash's 19%—making this a concrete candidate-engineering and selection failure rather than evidence of a broken grade.

One valid autonomous run produced a provisional 4% hidden exact-match score.

#17
Score position4% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What DeepSeek V4 Pro 0813 built

The run pursued an ambitious hybrid, but the expensive pieces never converged into a reliably packaged solver.

01

Large neural fallback

The agent trained a six-layer byte transformer on ordinary text and synthetic task examples, producing a final artifact just over 10 MB compressed.

02

Direct inference rules

It wrapped the model with parsers intended to answer structured completions directly when a recognizable pattern appeared.

03

Competing evaluation paths

A custom local test said 25/25 while the runner's packaged-artifact calibration remained between 12% and 16%. The discrepancy was investigated but never resolved.

03 · Trajectory

How the run unfolded

Opening 40m

Spent most of the budget on training

Repeated data-generation and transformer-training experiments delayed the first official candidate until roughly nineteen minutes remained.

41:23

Established a 16% checkpoint

The first two valid packaged candidates each answered 4 of 25 public calibration questions exactly.

Final 12m

Promoted a regression

Contract changes reduced measured calibration to 3 of 25. Despite the unresolved conflict with its local 25/25 result, the agent promoted and finalized that branch; it scored 4 of 100 hidden.

04 · Assessment

Good, bad, and unresolved

What worked

A valid and auditable run

The session stayed pinned, tools remained usable, nine candidate actions completed, and the selected artifact passed size and grading checks without recorded broker errors.

What hurt

Proxy confidence overrode real evidence

The decisive mistake was trusting a custom local evaluator after the actual packaged artifact had visibly regressed. V4 Flash encountered a similar mismatch but kept debugging until both paths agreed.

What remains unknown

Model variance and provider routing

This is one seed, and the allowed OpenRouter pool did not retain per-request provider attribution. Additional seeds are needed before treating 4% as the model's stable capability.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-00914%60:48$3.3810.09 MB9Valid
1.4MInput tokens
103.8KOutput tokens
25.4MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 9 candidate submission actions across the published run set. Selected artifacts averaged 10.09 MB compressed.

03

Time use

The runs averaged 60:48 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 285 agent messages · 9 candidate actions · 1 published run.

06 · Economics

What the result cost

$3.38 per displayed run at current configured rates.

The run cost an estimated $3.38, or 1.18 score points per dollar. V4 Flash produced nearly five times the score for less than half the cost because it converted more of the hour into verified inference hardening.

At current configured API-equivalent rates, the displayed run cost is $3.38 and value is 1.18 score points per dollar.

#17
Value positionValue rank #17 of 23 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-07-30.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare DeepSeek V4 Pro 0813 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard