Run postmortem · Published 2026-08-20

Qwen3.8-27B Aggressive FastMTP Q3_K_Pxhigh reasoning

A Q3_K_P HauhauCS derivative with FastMTP reached 19% on one RTX 5090—five points above the earlier local Qwen run.

19/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
19%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#5
Current cost$0.00
Points / $
Average time55:49
01 · Verdict

A materially stronger, distinct single-GPU Qwen result

This is not a rerun of the previously published 14% Qwen3.8-27B entry. It used the HauhauCS Aggressive GGUF derivative at Q3_K_P with a FastMTP sidecar at depth 3, xhigh agent effort, and a 200K-class context configuration on one RTX 5090. The selected artifact passed validation and answered 19 of 100 hidden questions exactly. With one seed, the result remains provisional.

One valid autonomous run produced a provisional 19% hidden exact-match score.

#5
Score position19% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Qwen3.8-27B Aggressive FastMTP Q3_K_P built

The run combined a quantized local agent model with speculative multi-token prediction while preserving the standard one-hour ARI Bench artifact workflow.

01

Distinct HauhauCS checkpoint

The agent used the Aggressive Qwen3.8-27B GGUF derivative rather than the NVFP4 checkpoint behind the earlier 14% row.

02

FastMTP sidecar at depth 3

Speculative multi-token prediction accelerated generation while the main model retained final-token verification.

03

Single RTX 5090 deployment

Inference remained local and unmetered, with xhigh effort and a large-context configuration tuned to fit one consumer GPU.

03 · Trajectory

How the run unfolded

Start

Entered the one-hour track at xhigh

The model began from the same public materials, hidden-exam boundary, and artifact constraints used across the leaderboard.

During run

Maintained an iterative candidate loop

Five candidate actions were retained, and the selected package reached 25 of 25 on public calibration.

55:49

Finalized a valid 19% artifact

The final 11.94 MB compressed package passed grading and answered 19 of 100 hidden items exactly.

04 · Assessment

Good, bad, and unresolved

What worked

Near-frontier one-seed local result

The run matched the 19% score of a strong API-served entry while using one locally owned RTX 5090 and no metered API calls.

What hurt

Throughput weakened at the longest active context

Generation usually exceeded 100 tokens per second at shorter and moderate context, but dropped below that target as the active context approached roughly 131K tokens.

What remains unknown

Repeatability and component attribution

Two more seeds are needed, and one aggregate score cannot isolate how much of the gain came from the checkpoint, quantization, sidecar, context, or run variance.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-010019%55:49$0.0011.94 MB5Valid
298.8KInput tokens
187.8KOutput tokens
15MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 11.94 MB compressed.

03

Time use

The runs averaged 55:49 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 156 agent messages · 5 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.00 per displayed run at current configured rates.

The run incurred no metered API charge because inference was local. Hardware ownership and electricity are excluded, so points-per-dollar is not reported.

At current configured API-equivalent rates, the displayed run cost is $0.00 and value is — score points per dollar.

Value positionValue rank unavailable. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-08-20.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Qwen3.8-27B Aggressive FastMTP Q3_K_P with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard