Tested several model sizes
Gemini moved quickly through small and large configurations rather than deeply tuning one.
Run postmortem · Published 2026-08-14
Six candidates in nine minutes: speed won, breadth did not.
One-hour track · xhigh reasoning · provisional
Gemini 3.5 Flash explored several neural configurations, wrapped the largest model in a compact solver, and finalized after only nine minutes. The 11% hidden score is respectable for such a short run, but the unfinished public calibration and unused wall time make it look more like a promising checkpoint than a completed search.
One valid autonomous run produced a provisional 11% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The run prioritized quick model comparison over prolonged solver hardening.
Five training attempts ranged from tiny baselines to a four-layer, 256-wide model near 12 MB.
A short inference layer handled common transformations and structured completions before neural fallback.
The largest augmented model became the final candidate after only six submissions.
Gemini moved quickly through small and large configurations rather than deeply tuning one.
The final public score still had visible errors, but the agent treated the package as sufficient.
The hidden grader returned 11%, leaving substantial unexplored time.
Eleven hidden matches from a nine-minute run is one of the library's strongest demonstrations of fast execution.
Unlike many peers, Gemini still had a visible improvement signal and enough time to act on it.
A longer run might have widened the solver—or revealed that the fast 11% checkpoint was already near this approach's ceiling.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0046 | 11% | 9:07 | $2.09 | 12.24 MB | 6 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 6 candidate submission actions across the published run set. Selected artifacts averaged 12.24 MB compressed.
The runs averaged 9:07 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 60 agent messages · 6 candidate actions · 1 published run.
$2.09 per displayed run at current configured rates.
The run's dollar cost was moderate, but its real unspent resource was time. Most of the fixed one-hour budget remained available when it finalized.
At current configured API-equivalent rates, the displayed run cost is $2.09 and value is 5.25 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Gemini 3.5 Flash with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard