Valid package, no public matches
The first candidate scored 0/25 on public calibration.
Run postmortem · Updated 2026-09-18
Pareto by Unbiased: 32% hidden accuracy in the Union Alpha preview run, now $5.31 at published token rates.
One-hour track · default reasoning · provisional
Formerly Union Alpha · Pareto model documentation.
The preserved Union Alpha preview run is now listed as Pareto by Unbiased. It finalized a valid candidate and answered 32 of 100 hidden questions correctly, with 25/25 public calibration. The model identity and current-price equivalent were updated after the reveal; no new benchmark was run. This row uses the completed run-0123, not an average of all four attempts. The earlier attempts ended with errors. The result remains provisional. ARI Bench evaluates the small system the agent builds, not the controller model answering the hidden questions directly.
One valid autonomous run produced a provisional 32% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The agent combined a trained byte-level model with deterministic task solvers in a compact inference package.
The agent completed baseline training and a later synthetic-data training job. The final compressed artifact is 6,356,851 bytes, below the 16 MB cap.
The submitted hybrid combines learned predictions with symbolic and record-processing logic. The agent revised and tested the inference code before selecting its final candidate.
Four candidate submissions were recorded, followed by an explicit finalization. The hidden grader used that finalized snapshot, not unfinished workspace edits.
The first candidate scored 0/25 on public calibration.
After additional training and solver development, the agent submitted a valid hybrid that passed all 25 public examples.
Later submissions retained 25/25 public calibration. A last experimental addition encountered a build error, and the agent finalized its previously validated checkpoint.
The final artifact passed validation and scored 32%. The controller exited successfully; the archive reports no broker errors.
The agent continued through intermittent provider errors, improved its baseline, and explicitly finalized a valid submission.
A perfect 25/25 public score became 32/100 on the hidden exam. Nine provider stream-error events were captured during this run, and a final build experiment failed.
The preserved result was obtained under the Union Alpha preview name. The revealed Pareto endpoint has not been benchmarked again here, and one completed run does not establish a stable model ranking.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0123 | 32% | 62:56 | $0.00 | 6.36 MB | 5 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 5 candidate submission actions across the published run set. Selected artifacts averaged 6.36 MB compressed.
The runs averaged 62:56 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 246 agent messages · 5 candidate actions · 1 published run.
$5.31 per displayed run at current configured rates.
OpenRouter publishes Pareto rates of $2.50 per million input tokens, $7.50 per million output tokens and $0.25 per million cache-read tokens. Applied to 992,800 input, 154,000 output and 6.7 million cache-read tokens, the current-price equivalent is $2.482 + $1.155 + $1.675 = $5.312, displayed as $5.31. Token totals are rounded. This replaces the earlier $4.08 projection; the actual preview cost remains $0.00. The plotted cost reprices the historical run and does not represent a new paid run.
At current configured API-equivalent rates, the displayed run cost is $5.31 and value is 6.02 score points per dollar. The archived cost at run time was $0.00; the difference is repricing, not a changed score.
Current configured API-equivalent cost / run · pricing snapshot 2026-09-18.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Pareto with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard