Run analysis 003 · August 13, 2026

Gemini 3.7 Flash moved fast.The hidden exam moved differently.

Google’s newest Flash reached 25/25 public calibration in six minutes, built a 5M-parameter hybrid system, and finished at 17/100—at roughly ninety cents through the discounted Vertex/global route.

17/ 100
hidden exact

One-hour track · maximum available “high” reasoning · one seed · provisional

public calibration25 / 25finalized early
StatusValid
Actual charge$0.8994
Time used16:25
Final artifact9.25 MB
Candidate actions9
Run0086

Gemini 3.7 Flash produced a valid, inexpensive improvement—and stopped after optimizing a public signal that could no longer distinguish better ideas.

It answered 17 of 100 hidden questions exactly, improving on Gemini 3.6 Flash’s observed 14% and Gemini 3.5 Flash’s 11%. The run was operationally sound: one pinned agent session, no broker or tool-protocol errors, six valid promoted candidates, and a self-contained final artifact within the size limit.

The result is provisional. It is one seed at “high,” the strongest reasoning setting Gemini 3.7 Flash exposes through OpenRouter. Earlier Gemini Flash runs on the leaderboard used “xhigh,” so the generation comparison is useful but not perfectly controlled.

The exact Vertex/global endpoint was 50% off. It turned 11.7 million tokens of agent work into a sub-dollar run.

The measured OpenRouter account-balance change was $0.8994, excluding the preflight. Reconstructing the run from rounded archived token buckets gives $0.9140; the small difference is rounding, so both figures are retained rather than forced to match.

$0.90

Actual charge. The composition below uses the $0.914 rounded-token reconstruction.

$0.401310.7M cache-read tokens · 43.9%
$0.3419911.8K fresh input · 37.4%
$0.170891.1K output · 18.7%
Standard-rate equivalent~$1.828
50% off
Discounted estimate~$0.914

On August 13, 2026, OpenRouter listed google-vertex/global at $0.375/M input, $1.875/M output, and $0.0375/M cache reads for this model. The route was pinned to that endpoint only, with provider fallback disabled. Sources: OpenRouter endpoint catalog, provider-routing documentation, and prompt-caching documentation.

Gemini treated the task as a mixture of recognizable symbolic problems and uncertain byte prediction. Its final 665-line inference program routed inputs through explicit solvers before falling back to a small trained transformer.

01

Pattern-specific solvers

Deterministic paths handled copying, text transforms, arithmetic, sorting, sequences, schedules, local mappings, structured records, and table-like inputs.

02

Few-shot induction

Local examples and suffix patterns were used to infer transformations that did not fit a single hard-coded category.

03

Synthetic multitask training

The agent generated task-shaped training data and mixed it with the ordinary corpus to teach a compact byte-level fallback.

04

Neural fallback

The final 5M-parameter transformer used six layers, 256-dimensional embeddings, eight attention heads, and a 512-byte training block.

Gemini moved unusually quickly from a safety checkpoint to a larger hybrid system. The session recovered from malformed early packages and a training-device error, but it finalized with most of the one-hour budget unused.

Established a safety checkpoint

The first promoted candidate was valid but scored 0/25 on public calibration.

Two packages failed validation

Early inference revisions returned malformed output shapes. Gemini corrected the packaging path instead of losing the run.

Jumped to 22/25

The first effective hybrid candidate captured most of the visible task patterns.

Saturated public calibration

The next candidate reached 25/25. Every subsequent valid promotion remained perfect on that small visible set.

Expanded the neural fallback

After correcting a CPU/CUDA generator mismatch, Gemini completed 3,000-step and 6,000-step training passes and grew the artifact to 9.25 MB.

Finalized early

The model ended with roughly 43½ minutes still available. The hidden grader then returned 17/100.

The good

Fast recovery, clean delivery

Gemini created an early valid checkpoint, recovered from three concrete failures, and delivered a compact hybrid artifact without tool or protocol problems.

The bad

Its tests stopped being informative

Public calibration and 58 self-generated checks were perfect, but hidden accuracy was 17%. The solver was effective on familiar structures and brittle outside them.

The unknown

Would the remaining time help?

More adversarial test generation might have exposed blind spots. But without a fresh signal, more iterations could also have deepened the same overfitting.

The most plausible explanation for the gap is not context loss or broken tool calling. The session remained continuous, tools worked, and the model executed a coherent plan. The weaker point was decision quality after visible validation saturated: synthetic tests resembled the system it had already designed, while the hidden distribution demanded broader transfer.

Across the three observed Flash runs, score rose while cost fell. Gemini 3.7 Flash delivered the best score and by far the best value of the series: about 18.9 hidden score points per charged dollar. Each model still has only one seed.

ModelScoreCostPoints / $
Gemini 3.5 Flash11%$2.095.25
Gemini 3.6 Flash14%$1.1811.90
Gemini 3.7 Flash17%$0.9018.90

Gemini 3.5 and 3.6 were run at xhigh; 3.7 exposes high as its maximum OpenRouter reasoning setting. These are one-seed observations, not repeatable model averages.

See Gemini 3.7 Flash in context—including score, measured run cost, value per dollar, seed count, and results from other frontier models.

View the leaderboard