Secured an early valid candidate
The first candidate established a gradeable fallback while the agent analyzed every public calibration pattern.
Run postmortem · Published 2026-09-02
It solved every visible calibration item in under twelve minutes, then the hidden exam held it to 14%.
One-hour track · high reasoning · provisional
Gemini 3.8 Flash completed a clean maximum-supported-effort run with five candidate submissions and a valid 14% hidden exact-match result. Its final package was independently re-imported, checkpoint-loaded, and hash-matched to the source Gemini finalized. The result is operationally sound, but this single seed did not beat Gemini 3.7 Flash's earlier 17% run and remains provisional.
One valid autonomous run produced a provisional 14% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The run combined a compact trained byte model with a broad deterministic inference layer, then stopped as soon as that hybrid saturated public calibration.
After rejecting an oversized expert, the agent trained an 8-layer, 192-wide, 3.59M-parameter transformer with Muon and AdamW; the final compressed bundle was 13,344,510 bytes.
It generated tens of thousands of examples spanning arithmetic, sequences, sorting, transformations, mappings, records, and passage-oriented completions.
The final inference path parsed recognizable tasks directly and used the trained neural checkpoint as a fallback for prefixes outside those rules.
The first candidate established a gradeable fallback while the agent analyzed every public calibration pattern.
The run rejected a 39.79 MB checkpoint and repaired an incompatible checkpoint/model pairing before training the smaller production model.
The selected hybrid reached 100% public calibration and 14% hidden accuracy, then ended roughly forty-eight minutes before the one-hour limit.
The graded infer.py, model.py, model.pt, and config.json matched the intended source hashes; packaged multi-prefix and autoregressive checks also passed.
A perfect 25/25 visible score concealed an 86-point hidden gap, yet the agent treated that calibration result as sufficient evidence to finish early.
Two additional seeds are required, and this run cannot show whether spending the remaining forty-eight minutes would broaden hidden coverage or merely add more public-perfect variants.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0109 | 14% | 11:33 | $1.50 | 13.34 MB | 5 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 13.34 MB compressed.
The runs averaged 11:33 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 157 agent messages · 5 candidate actions · 1 published run.
$1.50 per displayed run at current configured rates.
The archived token buckets imply $1.50 at the September 2 OpenRouter rates, or 9.32 hidden-score points per dollar. That is inexpensive in absolute terms, but weaker value than Gemini 3.7 Flash's earlier one-seed result.
At current configured API-equivalent rates, the displayed run cost is $1.50 and value is 9.32 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-09-02.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Gemini 3.8 Flash with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard