Run postmortem · Published 2026-09-17

Union Alphadefault reasoning

32% hidden accuracy from a completed run, with a projected cost of $4.08.

32/ 100
hidden exact

One-hour track · default reasoning · provisional

ARI score scaleexact-match accuracy
32%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#2
Projected cost$4.08
Points / $7.85
Average time62:56
01 · Verdict

A completed run reached 32% on the hidden exam

Union Alpha finalized a valid candidate and answered 32 of 100 hidden questions correctly. Public calibration was 25/25. This row uses the completed run-0123, not an average of all four attempts: the earlier attempts ended with errors. The result remains provisional. ARI Bench evaluates the small system the agent builds, not the controller model answering the hidden questions directly.

One valid autonomous run produced a provisional 32% hidden exact-match score.

#2
Score position32% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Union Alpha built

The agent combined a trained byte-level model with deterministic task solvers in a compact inference package.

01

A learned fallback

The agent completed baseline training and a later synthetic-data training job. The final compressed artifact is 6,356,851 bytes, below the 16 MB cap.

02

Task-specific solver code

The submitted hybrid combines learned predictions with symbolic and record-processing logic. The agent revised and tested the inference code before selecting its final candidate.

03

An explicitly finalized checkpoint

Four candidate submissions were recorded, followed by an explicit finalization. The hidden grader used that finalized snapshot, not unfinished workspace edits.

03 · Trajectory

How the run unfolded

Initial candidate

Valid package, no public matches

The first candidate scored 0/25 on public calibration.

Hybrid candidate

Public calibration reached 25/25

After additional training and solver development, the agent submitted a valid hybrid that passed all 25 public examples.

Final selection

The tested checkpoint was finalized

Later submissions retained 25/25 public calibration. A last experimental addition encountered a build error, and the agent finalized its previously validated checkpoint.

Hidden grade

32 of 100 exact matches

The final artifact passed validation and scored 32%. The controller exited successfully; the archive reports no broker errors.

04 · Assessment

Good, bad, and unresolved

What worked

Recovered and delivered a valid artifact

The agent continued through intermittent provider errors, improved its baseline, and explicitly finalized a valid submission.

What hurt

Public success did not generalize fully

A perfect 25/25 public score became 32/100 on the hidden exam. Nine provider stream-error events were captured during this run, and a final build experiment failed.

What remains unknown

One selected run and estimated economics

One completed run does not establish a stable model ranking. Post-preview token and cache pricing are not confirmed, so the plotted cost is a scenario estimate.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-012332%62:56$0.006.36 MB5Valid
992.8KInput tokens
154KOutput tokens
6.7MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 6.36 MB compressed.

03

Time use

The runs averaged 62:56 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 246 agent messages · 5 candidate actions · 1 published run.

06 · Economics

What the result cost

$4.08 per run at the unconfirmed projected rates.

The graphs use $4.0774 (displayed as $4.08), applying the third-party projection published at union-alpha.com: $0.50 per million input tokens and $1.50 per million output tokens. No projected cache rate is published, so 6.7 million cache-read tokens are charged at the full input price. Together with 992,800 input and 154,000 output tokens, the calculation is $0.4964 + $3.35 + $0.231 = $4.0774. Usage totals are rounded. These are unconfirmed projected rates; the actual free-preview cost was $0.00.

At the unconfirmed projected rates, the displayed run cost is $4.08 and value is 7.85 score points per dollar. The archived cost at run time was $0.00; the difference is a projection, not a changed score.

Value positionValue rank unavailable. Higher score points per dollar indicate better value.

Projected API-equivalent cost / run · pricing snapshot 2026-09-17.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Union Alpha with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard