Entered the one-hour track at xhigh
The model began from the same public materials, hidden-exam boundary, and artifact constraints used across the leaderboard.
Run postmortem · Published 2026-08-20
A Q3_K_P HauhauCS derivative with FastMTP reached 19% on one RTX 5090—five points above the earlier local Qwen run.
One-hour track · xhigh reasoning · provisional
This is not a rerun of the previously published 14% Qwen3.8-27B entry. It used the HauhauCS Aggressive GGUF derivative at Q3_K_P with a FastMTP sidecar at depth 3, xhigh agent effort, and a 200K-class context configuration on one RTX 5090. The selected artifact passed validation and answered 19 of 100 hidden questions exactly. With one seed, the result remains provisional.
One valid autonomous run produced a provisional 19% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The run combined a quantized local agent model with speculative multi-token prediction while preserving the standard one-hour ARI Bench artifact workflow.
The agent used the Aggressive Qwen3.8-27B GGUF derivative rather than the NVFP4 checkpoint behind the earlier 14% row.
Speculative multi-token prediction accelerated generation while the main model retained final-token verification.
Inference remained local and unmetered, with xhigh effort and a large-context configuration tuned to fit one consumer GPU.
The model began from the same public materials, hidden-exam boundary, and artifact constraints used across the leaderboard.
Five candidate actions were retained, and the selected package reached 25 of 25 on public calibration.
The final 11.94 MB compressed package passed grading and answered 19 of 100 hidden items exactly.
The run matched the 19% score of a strong API-served entry while using one locally owned RTX 5090 and no metered API calls.
Generation usually exceeded 100 tokens per second at shorter and moderate context, but dropped below that target as the active context approached roughly 131K tokens.
Two more seeds are needed, and one aggregate score cannot isolate how much of the gain came from the checkpoint, quantization, sidecar, context, or run variance.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0100 | 19% | 55:49 | $0.00 | 11.94 MB | 5 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 11.94 MB compressed.
The runs averaged 55:49 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 156 agent messages · 5 candidate actions · 1 published run.
$0.00 per displayed run at current configured rates.
The run incurred no metered API charge because inference was local. Hardware ownership and electricity are excluded, so points-per-dollar is not reported.
At current configured API-equivalent rates, the displayed run cost is $0.00 and value is — score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-08-20.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Qwen3.8-27B Aggressive FastMTP Q3_K_P with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard