Built a fast task-aware candidate
The agent reached 96% public calibration after less than twelve minutes.
Run postmortem · Published 2026-08-14
The run looked promising—then failed the probability contract.
One-hour track · xhigh reasoning · invalid
Qwen 3.7 Plus built a task-aware SmartInfer layer and reached 96% public calibration in under twelve minutes. The promoted artifact was not accepted as a valid model result because one probability distribution summed to 0.999686275 instead of one; the agent also exited without finalizing a replacement.
The submitted artifact reached evaluation, but the result was invalid and produced no valid seed.
No valid seed · Hidden exact-match accuracy, not public calibration accuracy.
The underlying idea was a peaked deterministic solver wrapped around a neural fallback.
The inference layer attempted to recover direct answers from the current prefix and convert them into peaked byte distributions.
When the direct path did not apply, the package retained the standard trained-model inference route.
Three candidates were submitted, but the run ended with only a promoted checkpoint available for official grading.
The agent reached 96% public calibration after less than twelve minutes.
The agent process failed before a corrected artifact was explicitly finalized.
The grader rejected the artifact's probability normalization, so no accuracy score exists.
The run understood the task shape quickly and nearly solved the public calibration set.
A tiny normalization error invalidated the entire artifact despite otherwise plausible predictions.
Because the artifact was invalid, assigning it 0% or comparing it to valid zero-score models would be misleading.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0044 | Invalid | 11:40 | $0.19 | 3.14 MB | 3 | Invalid |
The displayed 0% is a failure-state convention, not a valid zero-score seed. The attempt is documented but excluded from score and value ranks.
The agent made 3 candidate submission actions across the published run set. Selected artifacts averaged 3.14 MB compressed.
The runs averaged 11:40 of wall time. Runtime alone cannot show whether more time would have improved the result.
Retained totals: 58 agent messages · 3 candidate actions · 1 published run.
$0.19 per displayed run at current configured rates.
The invalid attempt cost only nineteen cents, but it produced no benchmark seed. Its value and score ranks should therefore remain unranked.
The attempt cost $0.19 at current configured rates, but an invalid result has no score-per-dollar rank.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Qwen 3.7 Plus with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard