Run postmortem · Published 2026-09-02

Gemini 3.8 Flashhigh reasoning

It solved every visible calibration item in under twelve minutes, then the hidden exam held it to 14%.

14/ 100
hidden exact

One-hour track · high reasoning · provisional

ARI score scaleexact-match accuracy
14%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#14
Current cost$1.50
Points / $9.32
Average time11:33
01 · Verdict

Fast and valid, but not a generational gain

Gemini 3.8 Flash completed a clean maximum-supported-effort run with five candidate submissions and a valid 14% hidden exact-match result. Its final package was independently re-imported, checkpoint-loaded, and hash-matched to the source Gemini finalized. The result is operationally sound, but this single seed did not beat Gemini 3.7 Flash's earlier 17% run and remains provisional.

One valid autonomous run produced a provisional 14% hidden exact-match score.

#14
Score position14% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Gemini 3.8 Flash built

The run combined a compact trained byte model with a broad deterministic inference layer, then stopped as soon as that hybrid saturated public calibration.

01

Compact RoPE byte model

After rejecting an oversized expert, the agent trained an 8-layer, 192-wide, 3.59M-parameter transformer with Muon and AdamW; the final compressed bundle was 13,344,510 bytes.

02

Synthetic task curriculum

It generated tens of thousands of examples spanning arithmetic, sequences, sorting, transformations, mappings, records, and passage-oriented completions.

03

Deterministic solver overlay

The final inference path parsed recognizable tasks directly and used the trained neural checkpoint as a fallback for prefixes outside those rules.

03 · Trajectory

How the run unfolded

Opening

Secured an early valid candidate

The first candidate established a gradeable fallback while the agent analyzed every public calibration pattern.

Middle

Recovered from two model-path problems

The run rejected a 39.79 MB checkpoint and repaired an incompatible checkpoint/model pairing before training the smaller production model.

11:33

Finalized with most of the hour unused

The selected hybrid reached 100% public calibration and 14% hidden accuracy, then ended roughly forty-eight minutes before the one-hour limit.

04 · Assessment

Good, bad, and unresolved

What worked

Exact artifact provenance is verified

The graded infer.py, model.py, model.pt, and config.json matched the intended source hashes; packaged multi-prefix and autoregressive checks also passed.

What hurt

Public saturation triggered a premature stop

A perfect 25/25 visible score concealed an 86-point hidden gap, yet the agent treated that calibration result as sufficient evidence to finish early.

What remains unknown

Repeatability and unused-time upside

Two additional seeds are required, and this run cannot show whether spending the remaining forty-eight minutes would broaden hidden coverage or merely add more public-perfect variants.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-010914%11:33$1.5013.34 MB5Valid
854.6KInput tokens
57.7KOutput tokens
8.6MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 13.34 MB compressed.

03

Time use

The runs averaged 11:33 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 157 agent messages · 5 candidate actions · 1 published run.

06 · Economics

What the result cost

$1.50 per displayed run at current configured rates.

The archived token buckets imply $1.50 at the September 2 OpenRouter rates, or 9.32 hidden-score points per dollar. That is inexpensive in absolute terms, but weaker value than Gemini 3.7 Flash's earlier one-seed result.

At current configured API-equivalent rates, the displayed run cost is $1.50 and value is 9.32 score points per dollar.

#7
Value positionValue rank #7 of 28 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-02.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Gemini 3.8 Flash with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard