Entered the standard one-hour track
The local model began from the same public task materials and fixed evaluation constraints as other entries.
Run postmortem · Published 2026-08-19
A locally hosted 27B open-weight model scored 14% in a one-hour autonomous run on a single RTX 5090.
One-hour track · default reasoning · provisional
Qwen3.8-27B completed the one-hour track using a locally hosted NVFP4 model with speculative decoding on one RTX 5090. The frozen end-of-run work product scored 14% on the hidden evaluation after a packaging-only correction; no post-run model work was added. One seed remains provisional.
One valid autonomous run produced a provisional 14% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The run paired local model serving with the standard one-hour autonomous evaluation.
The agent model was served from one RTX 5090 rather than a metered API.
The model worked within the same wall-clock and artifact limits used for the public one-hour field.
The published result evaluates the model's frozen end-of-run output after correcting packaging only.
The local model began from the same public task materials and fixed evaluation constraints as other entries.
The run produced a size-valid frozen artifact within the allotted time.
The packaging-corrected frozen output answered 14 of 100 hidden items exactly.
A single consumer GPU produced a valid mid-field score without per-token API charges.
The 14% result is provisional and should not be treated as a settled model ranking.
Two additional runs are needed to measure variance and determine an official mean.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0093 | 14% | 59:24 | $0.00 | 12.40 MB | 5 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 12.40 MB compressed.
The runs averaged 59:24 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 203 agent messages · 5 candidate actions · 1 published run.
$0.00 per displayed run at current configured rates.
The run incurred no metered API charge because inference was local. Hardware ownership and electricity are not included, so points-per-dollar is left undefined.
At current configured API-equivalent rates, the displayed run cost is $0.00 and value is — score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-08-19.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Qwen3.8-27B with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard