Run postmortem · Published 2026-08-14

Qwen 3.7 Plusxhigh reasoning

The run looked promising—then failed the probability contract.

0/ 100
hidden exact

One-hour track · xhigh reasoning · invalid

ARI score scaleexact-match accuracy
0%
0255075100
StatusInvalid
Valid seeds0 / 3
Score rankUnranked
Current cost$0.19
Points / $
Average time11:40
01 · Verdict

Not a zero score: an invalid submission

Qwen 3.7 Plus built a task-aware SmartInfer layer and reached 96% public calibration in under twelve minutes. The promoted artifact was not accepted as a valid model result because one probability distribution summed to 0.999686275 instead of one; the agent also exited without finalizing a replacement.

The submitted artifact reached evaluation, but the result was invalid and produced no valid seed.

Unranked
Score positionValidation failed, so this attempt is not placed on the score table.

No valid seed · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Qwen 3.7 Plus built

The underlying idea was a peaked deterministic solver wrapped around a neural fallback.

01

Smart contextual extraction

The inference layer attempted to recover direct answers from the current prefix and convert them into peaked byte distributions.

02

Neural fallback

When the direct path did not apply, the package retained the standard trained-model inference route.

03

Promoted, never finalized

Three candidates were submitted, but the run ended with only a promoted checkpoint available for official grading.

03 · Trajectory

How the run unfolded

Opening

Built a fast task-aware candidate

The agent reached 96% public calibration after less than twelve minutes.

Submission

Left a promoted checkpoint

The agent process failed before a corrected artifact was explicitly finalized.

Grading

Failed strict distribution validation

The grader rejected the artifact's probability normalization, so no accuracy score exists.

04 · Assessment

Good, bad, and unresolved

What worked

Strong visible progress

The run understood the task shape quickly and nearly solved the public calibration set.

What hurt

Numerical validity was not hardened

A tiny normalization error invalidated the entire artifact despite otherwise plausible predictions.

What remains unknown

The hidden score it might have earned

Because the artifact was invalid, assigning it 0% or comparing it to valid zero-score models would be misleading.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-0044Invalid11:40$0.193.14 MB3Invalid
154.2KInput tokens
22.6KOutput tokens
1.8MCache-read tokens
01

No valid accuracy estimate

The displayed 0% is a failure-state convention, not a valid zero-score seed. The attempt is documented but excluded from score and value ranks.

02

Candidate behavior

The agent made 3 candidate submission actions across the published run set. Selected artifacts averaged 3.14 MB compressed.

03

Time use

The runs averaged 11:40 of wall time. Runtime alone cannot show whether more time would have improved the result.

Retained totals: 58 agent messages · 3 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.19 per displayed run at current configured rates.

The invalid attempt cost only nineteen cents, but it produced no benchmark seed. Its value and score ranks should therefore remain unranked.

The attempt cost $0.19 at current configured rates, but an invalid result has no score-per-dollar rank.

Value positionValue rank unavailable. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Qwen 3.7 Plus with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard