Established a trained baseline
The run retained a baseline checkpoint while building and testing the direct inference helpers.
Run postmortem · Published 2026-09-04
38 correct answers, a verified 5.4 MB artifact, and a new highest observed score on the one-hour track.
One-hour track · xhigh reasoning · provisional
GPT-6 Astra scored 38/100 on the same private one-hour hard exam, 12 percentage points above the previous highest displayed score of 26%. It finalized its selected artifact after 54 minutes. The exact submitted code and weights matched the selected source by hash, passed checkpoint and autoregressive checks, and received a valid original Docker grade. No packaging correction or regrade was needed. This is one provisional seed; it does not establish a repeatable 38% average.
One valid autonomous run produced a provisional 38% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The final artifact combines a compact trained byte model with deterministic handlers that work from the visible prefix. The retained code embeds four solver modules inside a self-contained inference file.
A 3,487,232-parameter model supplies byte probabilities when the direct solvers do not find a supported answer. The final checkpoint and matching configuration are included in the package.
Symbolic, structured-data, textual, and induction helpers propose candidate completions. The inference layer combines compatible candidates and falls back to the trained model when necessary.
Five candidate submission actions were recorded. Astra selected the augmented final candidate, and its inference code and weights survived finalization byte-for-byte.
The run retained a baseline checkpoint while building and testing the direct inference helpers.
The workspace records separate solver tests, an augmented training path, and explicit packaging and inference checks before final selection.
The selected artifact reached 25/25 on public calibration and 38/100 on the private exam. The run ended voluntarily at 54:02 with a 5,404,660-byte package, below the 16 MB cap.
All substantive source/package hashes matched. The exact package imported, loaded its checkpoint, produced valid multi-prefix distributions, and passed two independent autoregressive checks.
Perfect public calibration still left 62 of the 100 private questions without an exact match. The available evidence does not isolate the contribution of each solver family.
Two more valid seeds are required for an official estimate. This run used OpenCode 1.18.28, upgraded from the previous 1.17.11 controller, and ChatGPT subscription OAuth. The exam, one-hour limit, GPU, and 16 MB cap were unchanged; the CLI version and subscription backend are comparability limits.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0113 | 38% | 54:02 | — | 5.40 MB | 5 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 5.40 MB compressed.
The runs averaged 54:02 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 140 agent messages · 5 candidate actions · 1 published run.
$18.65 per displayed run at current configured rates.
This run used ChatGPT subscription quota. The displayed $18.65 is an approximate API-equivalent estimate, not an API charge: 540K uncached input tokens at $10 per million, 149K output at $50 per million, and 5.8M cached input at $1 per million. These are rounded aggregate OpenCode statistics. The September 4 official OpenAI rate card also applies premiums above 272K input tokens per request; a complete per-request breakdown was not retained, so this base-rate estimate cannot verify those premiums. Pricing source: https://developers.openai.com/api/docs/models/gpt-6-astra. The interrupted launch run-0112 produced no scored artifact and is excluded.
At current configured API-equivalent rates, the displayed run cost is $18.65 and value is 2.04 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-09-04.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare GPT-6 Astra with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard