Built the direct solver first
The run focused immediately on task-aware inference rather than lengthy model training.
Run postmortem · Published 2026-08-14
The shortest generated run produced a 14% score with a 3.2 MB artifact.
One-hour track · xhigh reasoning · provisional
Gemini 3.6 Flash assembled a small deterministic and n-gram hybrid, reached perfect public calibration, and finalized in six and a half minutes. Its 14% hidden score is strong relative to both time and artifact size, but it leaves the largest unused exploration budget of any successful generated page.
One valid autonomous run produced a provisional 14% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The final package emphasized direct inference and lightweight statistical fallback over a large trained network.
The solver covered arithmetic, structured records, mappings, sequences, and contextual extraction.
When a direct answer was unavailable, the system used local byte statistics rather than a large neural checkpoint.
The selected package was only 3.20 MB and required seven candidate submissions.
The run focused immediately on task-aware inference rather than lengthy model training.
Every visible item passed, removing the public set as a useful selector.
The hidden result landed at 14%, well above several full-hour attempts.
The run converted very little time and a small package into a competitive upper-midfield score.
The agent did not use the hour to diversify tests after public calibration saturated.
With only one seed and six minutes of work, the result says little about what a patient 3.6 Flash run could do.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0072 | 14% | 6:33 | $1.18 | 3.20 MB | 7 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 7 candidate submission actions across the published run set. Selected artifacts averaged 3.20 MB compressed.
The runs averaged 6:33 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 57 agent messages · 7 candidate actions · 1 published run.
$1.18 per displayed run at current configured rates.
Gemini 3.6 Flash is one of the better value results even without repricing. Its strongest advantage was operational: it achieved 14% with minimal wall time and a small artifact.
At current configured API-equivalent rates, the displayed run cost is $1.18 and value is 11.9 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Gemini 3.6 Flash with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard