V3 Pi-harness postmortem · Published 2026-08-21

Ox Alpha · Pi 0.84.2max reasoning

Pi improved public calibration from 44% to 48%, but the finalized artifact earned only two hidden exact matches.

2/ 100
hidden exact

V3 · Pi harness · max reasoning · provisional

ARI score scaleexact-match accuracy
2%
0255075100
StatusProvisional
Valid seeds1 / 3
V3 rank#1
Current cost$0.00
Points / $
Run time56:03
01 · Verdict

The Pi run was valid; its learned strategy transferred poorly

Ox Alpha completed one pinned Pi 0.84.2 session at max reasoning with no broker errors and a clean return code. It finalized a valid 12.41 MB candidate that scored 48% on the 25-question public calibration set, then 2% on the 100-question hidden exam.

Hashes for infer.py, model.py, model.pt, and config.json match across the selected candidate, promoted copy, finalized copy, and final model. The grader explicitly scored the finalized directory.

#1
V3 score position2% on the Pi-only V3 track. One valid run is not an official estimate.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Ox Alpha built through Pi

The run built a pure byte-level neural predictor from scratch, using task-shaped synthetic data and a 1,024-byte inference window.

01

Environment and smoke validation

After correcting an initial command-path mistake, the agent verified Python, PyTorch, CUDA, and the RTX 5090 before committing to training.

02

Cap-aware architecture reduction

The first 54.61 MB package was rejected. The agent reduced the model to eight layers, eight attention heads, and width 256; the replacement compressed to 12.41 MB.

03

Synthetic curriculum refinement

The agent generated task-shaped training examples, promoted a 44% public candidate, then refined the same architecture to 48% before finalization.

03 · Trajectory

How the run unfolded

Opening

Probed the rig and built a smoke model

Pi established the training path, created evaluation utilities, and generated a synthetic completion corpus.

Middle

Recovered from the size-cap rejection

A valid 12.41 MB replacement scored 44% publicly and was promoted as a timeout-safe fallback.

56:03

Finalized 48% public, scored 2% hidden

The last refinement added one public exact match, but the broader hidden exam exposed the same large generalization gap.

04 · Assessment

Good, bad, and unresolved

What worked

Pi completed a disciplined, recoverable run

The pinned session kept a valid fallback, recovered from an oversized package, promoted the improved artifact, explicitly finalized it, and ended without transport or broker failures.

What hurt

The final predictor remained almost entirely neural

The submitted inference path used the trained transformer and cache acceleration, but no broad symbolic extraction or arithmetic layer. Public task-shaped learning did not generalize to the hidden distribution.

What remains unknown

Whether Pi materially changes model quality

The main-track OpenCode run scored 1% and this Pi V3 run scored 2%, but each is a single stochastic seed. That difference is not enough to claim a harness advantage.

05 · Evidence

The retained V3 run record

The published evidence uses aggregate run metadata, public calibration, and audited artifact hashes. It does not expose hidden questions, hidden answers, or private raw transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactActionsGrade
run-01042%56:03$0.0012.41 MB5Valid
134.0KInput tokens
42.2KOutput tokens
1.79MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The history records one rejected oversized package, a valid 44% promotion, a 48% calibration candidate, its promoted copy, and explicit finalization.

03

Provenance

The final core-file hashes match at every promotion and finalization boundary; the 12,408,097-byte compressed artifact passed the 16 MB cap.

Retained totals: 54 agent messages · 5 candidate-history actions · 1 published V3 run.

06 · Economics

What the result cost

$0.00 per displayed run at current configured rates.

Ox Alpha was priced at zero through OpenRouter for this evaluation. Hardware and electricity are excluded, and a zero-cost row does not receive a finite points-per-dollar ratio.

Value positionValue rank unavailable because the displayed API cost is zero.

Current configured API-equivalent cost / run · pricing snapshot 2026-08-21.

07 · Context

Pi-only V3 track

V3 is the Pi-harness benchmark track. The OpenCode 1.18.21 control remains on the main one-hour chart and is not included in this V3 ranking.

For the separately published OpenCode run, see the main-track postmortem.

Compare the Pi-only V3 result with the main one-hour leaderboard while keeping the harness tracks distinct.

View the leaderboard