Run postmortem · Published 2026-08-14

Gemini 3.5 Flashxhigh reasoning

Six candidates in nine minutes: speed won, breadth did not.

11/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
11%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#13
Current cost$2.09
Points / $5.25
Average time9:07
01 · Verdict

A fast experiment that stopped at the first credible answer

Gemini 3.5 Flash explored several neural configurations, wrapped the largest model in a compact solver, and finalized after only nine minutes. The 11% hidden score is respectable for such a short run, but the unfinished public calibration and unused wall time make it look more like a promising checkpoint than a completed search.

One valid autonomous run produced a provisional 11% hidden exact-match score.

#13
Score position11% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Gemini 3.5 Flash built

The run prioritized quick model comparison over prolonged solver hardening.

01

Rapid architecture sweep

Five training attempts ranged from tiny baselines to a four-layer, 256-wide model near 12 MB.

02

Compact task solver

A short inference layer handled common transformations and structured completions before neural fallback.

03

Early promotion

The largest augmented model became the final candidate after only six submissions.

03 · Trajectory

How the run unfolded

Opening

Tested several model sizes

Gemini moved quickly through small and large configurations rather than deeply tuning one.

~09:00

Promoted at 92% calibration

The final public score still had visible errors, but the agent treated the package as sufficient.

09:07

Ended with most of the hour unused

The hidden grader returned 11%, leaving substantial unexplored time.

04 · Assessment

Good, bad, and unresolved

What worked

Excellent speed-to-score

Eleven hidden matches from a nine-minute run is one of the library's strongest demonstrations of fast execution.

What hurt

The run stopped before calibration was solved

Unlike many peers, Gemini still had a visible improvement signal and enough time to act on it.

What remains unknown

The cost of the early exit

A longer run might have widened the solver—or revealed that the fast 11% checkpoint was already near this approach's ceiling.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-004611%9:07$2.0912.24 MB6Valid
939.3KInput tokens
29.4KOutput tokens
2.8MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 6 candidate submission actions across the published run set. Selected artifacts averaged 12.24 MB compressed.

03

Time use

The runs averaged 9:07 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 60 agent messages · 6 candidate actions · 1 published run.

06 · Economics

What the result cost

$2.09 per displayed run at current configured rates.

The run's dollar cost was moderate, but its real unspent resource was time. Most of the fixed one-hour budget remained available when it finalized.

At current configured API-equivalent rates, the displayed run cost is $2.09 and value is 5.25 score points per dollar.

#7
Value positionValue rank #7 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Gemini 3.5 Flash with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard