Run postmortem · Published 2026-08-19

Qwen3.8-27Bdefault reasoning

A locally hosted 27B open-weight model scored 14% in a one-hour autonomous run on a single RTX 5090.

14/ 100
hidden exact

One-hour track · default reasoning · provisional

ARI score scaleexact-match accuracy
14%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#12
Current cost$0.00
Points / $
Average time59:24
01 · Verdict

A credible single-GPU result from a compact open model

Qwen3.8-27B completed the one-hour track using a locally hosted NVFP4 model with speculative decoding on one RTX 5090. The frozen end-of-run work product scored 14% on the hidden evaluation after a packaging-only correction; no post-run model work was added. One seed remains provisional.

One valid autonomous run produced a provisional 14% hidden exact-match score.

#12
Score position14% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Qwen3.8-27B built

The run paired local model serving with the standard one-hour autonomous evaluation.

01

Single-GPU local inference

The agent model was served from one RTX 5090 rather than a metered API.

02

Fixed one-hour budget

The model worked within the same wall-clock and artifact limits used for the public one-hour field.

03

Frozen-output grading

The published result evaluates the model's frozen end-of-run output after correcting packaging only.

03 · Trajectory

How the run unfolded

Start

Entered the standard one-hour track

The local model began from the same public task materials and fixed evaluation constraints as other entries.

One hour

Completed a full autonomous attempt

The run produced a size-valid frozen artifact within the allotted time.

Final grade

Reached 14% hidden accuracy

The packaging-corrected frozen output answered 14 of 100 hidden items exactly.

04 · Assessment

Good, bad, and unresolved

What worked

Strong result for a local 27B model

A single consumer GPU produced a valid mid-field score without per-token API charges.

What hurt

One run is not a stable estimate

The 14% result is provisional and should not be treated as a settled model ranking.

What remains unknown

Repeatability across seeds

Two additional runs are needed to measure variance and determine an official mean.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-009314%59:24$0.0012.40 MB5Valid
6.4MInput tokens
252.4KOutput tokens
0Cache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 12.40 MB compressed.

03

Time use

The runs averaged 59:24 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 203 agent messages · 5 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.00 per displayed run at current configured rates.

The run incurred no metered API charge because inference was local. Hardware ownership and electricity are not included, so points-per-dollar is left undefined.

At current configured API-equivalent rates, the displayed run cost is $0.00 and value is — score points per dollar.

Value positionValue rank unavailable. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-08-19.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Qwen3.8-27B with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard