Run postmortem · Published 2026-08-28

Hy4 Previewmax reasoning

It reached 96% public calibration, finalized a 92% package, and still scored only 4% hidden.

4/ 100
hidden exact

One-hour track · max reasoning · provisional

ARI score scaleexact-match accuracy
4%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#20
Current cost$0.46
Points / $8.79
Average time66:48
01 · Verdict

The strongest public signal in the run was nearly useless

Hy4 Preview produced a valid 4% hidden exact-match result in its first published max-effort run. The agent built a hybrid byte predictor, raised official public calibration from 52% to 96%, then replaced that package with a fine-tuned 92% candidate near the deadline. Independent grading reproduced 4% from the exact finalized artifact, so the 88-point public-to-hidden gap reflects severe selection and generalization failure rather than the Ox Alpha package-path defect. One seed remains provisional.

One valid autonomous run produced a provisional 4% hidden exact-match score.

#20
Score position4% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Hy4 Preview built

The final artifact combined a nine-layer byte transformer with a large deterministic inference layer and repeated package-level calibration checks.

01

Multi-stage byte-model training

The run trained and fine-tuned progressively larger byte transformers, ending with a nine-layer, 256-wide neural fallback.

02

Programmatic completion solvers

Inference logic handled copying, arithmetic, records, sequences, schedules, mappings, sorting, and text transformations before blending in the neural model.

03

Atomic candidate packaging

An oversized first build was rejected safely, while later candidates were packaged, scored, promoted, and independently checked under the 16 MB cap.

03 · Trajectory

How the run unfolded

Opening 9m

The first package exceeded the cap

A 24.09 MB candidate failed closed instead of becoming selectable, forcing the run toward a smaller model.

Middle

Public calibration climbed to 96%

The compact hybrid moved from 52% to 92% and then 96% on the official public set, while synthetic holdouts reinforced the same narrow direction.

Deadline

Finalized a lower-scoring fine-tune

Hy4 replaced the preserved 96% package with a 92% fine-tuned candidate, finalized it, reached the controller deadline, and received a valid 4% hidden grade.

04 · Assessment

Good, bad, and unresolved

What worked

The final artifact is fully auditable

The selected package passed size and inference checks, matched its recorded code and weights, and reproduced 4% in an independent Docker re-grade.

What hurt

Calibration collapsed out of sample

Twenty-three of twenty-five public questions were correct, but only four of one hundred hidden questions matched exactly.

What remains unknown

Replication and model identity

Two more valid seeds are needed, and the equal 4% score does not establish that Hy4 Preview, GLM-5.3 Flash, and the former Ox Alpha endpoint are identical models.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-01054%66:48$0.4613.57 MB10Valid
107.6KInput tokens
38.6KOutput tokens
6.4MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 10 candidate submission actions across the published run set. Selected artifacts averaged 13.57 MB compressed.

03

Time use

The runs averaged 66:48 of wall time. Runtime alone cannot show whether more time would have improved the result.

Retained totals: 140 agent messages · 10 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.46 per displayed run at current configured rates.

The official run cost an estimated $0.4551 at its launch pricing, excluding preflight smoke requests. That is inexpensive in absolute terms, but the 4% hidden result and one-seed status make the displayed value ratio provisional.

At current configured API-equivalent rates, the displayed run cost is $0.46 and value is 8.79 score points per dollar.

#7
Value positionValue rank #7 of 26 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-08-28.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Hy4 Preview with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard