Run postmortem · Published 2026-09-04

GPT-6 Astraxhigh reasoning

38 correct answers, a verified 5.4 MB artifact, and a new highest observed score on the one-hour track.

38/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
38%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#1
Current cost$18.65
Points / $2.04
Average time54:02
01 · Verdict

A new high score, with repeatability still to establish

GPT-6 Astra scored 38/100 on the same private one-hour hard exam, 12 percentage points above the previous highest displayed score of 26%. It finalized its selected artifact after 54 minutes. The exact submitted code and weights matched the selected source by hash, passed checkpoint and autoregressive checks, and received a valid original Docker grade. No packaging correction or regrade was needed. This is one provisional seed; it does not establish a repeatable 38% average.

One valid autonomous run produced a provisional 38% hidden exact-match score.

#1
Score position38% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GPT-6 Astra built

The final artifact combines a compact trained byte model with deterministic handlers that work from the visible prefix. The retained code embeds four solver modules inside a self-contained inference file.

01

Trained byte-model fallback

A 3,487,232-parameter model supplies byte probabilities when the direct solvers do not find a supported answer. The final checkpoint and matching configuration are included in the package.

02

Four complementary solver families

Symbolic, structured-data, textual, and induction helpers propose candidate completions. The inference layer combines compatible candidates and falls back to the trained model when necessary.

03

Verified artifact selection

Five candidate submission actions were recorded. Astra selected the augmented final candidate, and its inference code and weights survived finalization byte-for-byte.

03 · Trajectory

How the run unfolded

Opening

Established a trained baseline

The run retained a baseline checkpoint while building and testing the direct inference helpers.

Iteration

Expanded the hybrid and checked packaging

The workspace records separate solver tests, an augmented training path, and explicit packaging and inference checks before final selection.

Finalization

Finished within the hour

The selected artifact reached 25/25 on public calibration and 38/100 on the private exam. The run ended voluntarily at 54:02 with a 5,404,660-byte package, below the 16 MB cap.

04 · Assessment

Good, bad, and unresolved

What worked

The submitted artifact earned the result

All substantive source/package hashes matched. The exact package imported, loaded its checkpoint, produced valid multi-prefix distributions, and passed two independent autoregressive checks.

What hurt

Public calibration did not cover the hidden gap

Perfect public calibration still left 62 of the 100 private questions without an exact match. The available evidence does not isolate the contribution of each solver family.

What remains unknown

One seed and an updated CLI

Two more valid seeds are required for an official estimate. This run used OpenCode 1.18.28, upgraded from the previous 1.17.11 controller, and ChatGPT subscription OAuth. The exam, one-hour limit, GPU, and 16 MB cap were unchanged; the CLI version and subscription backend are comparability limits.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-011338%54:025.40 MB5Valid
540KInput tokens
149KOutput tokens
5.8MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 5.40 MB compressed.

03

Time use

The runs averaged 54:02 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 140 agent messages · 5 candidate actions · 1 published run.

06 · Economics

What the result cost

$18.65 per displayed run at current configured rates.

This run used ChatGPT subscription quota. The displayed $18.65 is an approximate API-equivalent estimate, not an API charge: 540K uncached input tokens at $10 per million, 149K output at $50 per million, and 5.8M cached input at $1 per million. These are rounded aggregate OpenCode statistics. The September 4 official OpenAI rate card also applies premiums above 272K input tokens per request; a complete per-request breakdown was not retained, so this base-rate estimate cannot verify those premiums. Pricing source: https://developers.openai.com/api/docs/models/gpt-6-astra. The interrupted launch run-0112 produced no scored artifact and is excluded.

At current configured API-equivalent rates, the displayed run cost is $18.65 and value is 2.04 score points per dollar.

#21
Value positionValue rank #21 of 30 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-04.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GPT-6 Astra with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard