Fast recovery, clean delivery
Gemini created an early valid checkpoint, recovered from three concrete failures, and delivered a compact hybrid artifact without tool or protocol problems.
Google’s newest Flash reached 25/25 public calibration in six minutes, built a 5M-parameter hybrid system, and finished at 17/100—at roughly ninety cents through the discounted Vertex/global route.
One-hour track · maximum available “high” reasoning · one seed · provisional
Gemini 3.7 Flash produced a valid, inexpensive improvement—and stopped after optimizing a public signal that could no longer distinguish better ideas.
It answered 17 of 100 hidden questions exactly, improving on Gemini 3.6 Flash’s observed 14% and Gemini 3.5 Flash’s 11%. The run was operationally sound: one pinned agent session, no broker or tool-protocol errors, six valid promoted candidates, and a self-contained final artifact within the size limit.
The result is provisional. It is one seed at “high,” the strongest reasoning setting Gemini 3.7 Flash exposes through OpenRouter. Earlier Gemini Flash runs on the leaderboard used “xhigh,” so the generation comparison is useful but not perfectly controlled.
The exact Vertex/global endpoint was 50% off. It turned 11.7 million tokens of agent work into a sub-dollar run.
The measured OpenRouter account-balance change was $0.8994, excluding the preflight. Reconstructing the run from rounded archived token buckets gives $0.9140; the small difference is rounding, so both figures are retained rather than forced to match.
Actual charge. The composition below uses the $0.914 rounded-token reconstruction.
On August 13, 2026, OpenRouter listed google-vertex/global at $0.375/M input, $1.875/M output, and $0.0375/M cache reads for this model. The route was pinned to that endpoint only, with provider fallback disabled. Sources: OpenRouter endpoint catalog, provider-routing documentation, and prompt-caching documentation.
Gemini treated the task as a mixture of recognizable symbolic problems and uncertain byte prediction. Its final 665-line inference program routed inputs through explicit solvers before falling back to a small trained transformer.
Deterministic paths handled copying, text transforms, arithmetic, sorting, sequences, schedules, local mappings, structured records, and table-like inputs.
Local examples and suffix patterns were used to infer transformations that did not fit a single hard-coded category.
The agent generated task-shaped training data and mixed it with the ordinary corpus to teach a compact byte-level fallback.
The final 5M-parameter transformer used six layers, 256-dimensional embeddings, eight attention heads, and a 512-byte training block.
Gemini moved unusually quickly from a safety checkpoint to a larger hybrid system. The session recovered from malformed early packages and a training-device error, but it finalized with most of the one-hour budget unused.
The first promoted candidate was valid but scored 0/25 on public calibration.
Early inference revisions returned malformed output shapes. Gemini corrected the packaging path instead of losing the run.
The first effective hybrid candidate captured most of the visible task patterns.
The next candidate reached 25/25. Every subsequent valid promotion remained perfect on that small visible set.
After correcting a CPU/CUDA generator mismatch, Gemini completed 3,000-step and 6,000-step training passes and grew the artifact to 9.25 MB.
The model ended with roughly 43½ minutes still available. The hidden grader then returned 17/100.
Gemini created an early valid checkpoint, recovered from three concrete failures, and delivered a compact hybrid artifact without tool or protocol problems.
Public calibration and 58 self-generated checks were perfect, but hidden accuracy was 17%. The solver was effective on familiar structures and brittle outside them.
More adversarial test generation might have exposed blind spots. But without a fresh signal, more iterations could also have deepened the same overfitting.
The most plausible explanation for the gap is not context loss or broken tool calling. The session remained continuous, tools worked, and the model executed a coherent plan. The weaker point was decision quality after visible validation saturated: synthetic tests resembled the system it had already designed, while the hidden distribution demanded broader transfer.
Across the three observed Flash runs, score rose while cost fell. Gemini 3.7 Flash delivered the best score and by far the best value of the series: about 18.9 hidden score points per charged dollar. Each model still has only one seed.
Gemini 3.5 and 3.6 were run at xhigh; 3.7 exposes high as its maximum OpenRouter reasoning setting. These are one-seed observations, not repeatable model averages.
See Gemini 3.7 Flash in context—including score, measured run cost, value per dollar, seed count, and results from other frontier models.
View the leaderboard