Run postmortem · Published 2026-08-14

DeepSeek V4 Flash 0731xhigh reasoning

Nineteen candidates, no recorded training run, and a 19% solver built almost entirely in inference code.

19/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
19%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#5
Current cost$1.58
Points / $12.04
Average time57:38
01 · Verdict

A strong solver engineering run

DeepSeek V4 Flash spent nearly the full hour expanding and hardening a broad deterministic inference system rather than training a large fallback. The 19% hidden score is the strongest generated one-seed result below Sol and GPT-5.5 xhigh, and its $1.58 cost gives it excellent current value.

One valid autonomous run produced a provisional 19% hidden exact-match score.

#5
Score position19% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What DeepSeek V4 Flash 0731 built

The artifact treated each prefix as a structured reasoning problem and solved it directly when possible.

01

Broad deterministic coverage

The solver handled copying, arithmetic, sequences, mappings, records, lists, text transformations, and passage extraction.

02

Context-window hardening

Later patches deliberately searched recent and full context to avoid missing repeated records or mappings.

03

Package-test-repair loop

Thirteen packaging actions and nineteen submissions kept the solver under continuous artifact-level validation.

03 · Trajectory

How the run unfolded

Opening

Chose solver breadth over training

No training run was retained; effort went into direct inference and packaging from the beginning.

Late middle

Fixed full-context failures

The agent identified repeated-record and mapping misses, broadened the search window, and reached perfect calibration.

57:38

Finalized after sustained hardening

The resulting 5.35 MB artifact scored 19% hidden.

04 · Assessment

Good, bad, and unresolved

What worked

High score from disciplined inference engineering

DeepSeek converted almost the entire hour into broader solver coverage and a near-frontier one-seed result.

What hurt

Calibration saturation encouraged patching

Once the public set reached 100%, further special-case additions could not be ranked for hidden generalization.

What remains unknown

Repeatability

A second and third seed are needed to determine whether 19% is stable or a favorable solver path.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-007819%57:38$1.585.35 MB19Valid
2.9MInput tokens
96.8KOutput tokens
40.9MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 19 candidate submission actions across the published run set. Selected artifacts averaged 5.35 MB compressed.

03

Time use

The runs averaged 57:38 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 232 agent messages · 19 candidate actions · 1 published run.

06 · Economics

What the result cost

$1.58 per displayed run at current configured rates.

DeepSeek's $1.58 run delivered 12.04 points per dollar, placing it among the best current values without relying on later repricing.

At current configured API-equivalent rates, the displayed run cost is $1.58 and value is 12.04 score points per dollar.

#4
Value positionValue rank #4 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-08-01.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare DeepSeek V4 Flash 0731 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard