Run postmortem · Published 2026-08-14

Inklingmax reasoning

One candidate consumed the hour and never produced a positive calibration signal.

0/ 100
hidden exact

One-hour track · max reasoning · provisional

ARI score scaleexact-match accuracy
0%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#22
Current cost$2.45
Points / $0
Average time59:08
01 · Verdict

The run never established a working feedback loop

Inkling used almost the entire hour to train a small transformer and add heuristic inference, but submitted only one candidate. That artifact scored 0% on both the retained public calibration and the hidden exam, making this a valid failure rather than an invalid run.

One valid autonomous run completed, but it produced no exact matches on the hidden exam.

#22
Score position0% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Inkling built

The final system combined a small trained model with a compact set of direct rules.

01

Small neural baseline

The run trained a three-layer, 128-wide byte model with a conventional next-byte objective.

02

Limited heuristic layer

It added arithmetic, record parsing, mapping, and bracket-oriented extraction around the neural fallback.

03

Single-shot submission

Only one candidate reached the runner, eliminating the normal promote-test-repair loop.

03 · Trajectory

How the run unfolded

Opening

Built a conventional baseline

Most of the early work went into establishing the small neural model.

Middle

Added a narrow solver

The agent supplemented the model with several task-aware rules but did not demonstrate them through iterative submissions.

59:08

Submitted once and scored zero

The final artifact was valid, but neither public nor hidden grading found an exact match.

04 · Assessment

Good, bad, and unresolved

What worked

The result is diagnostically clean

The artifact passed validation, so the zero can be attributed to the submitted system rather than a grading-format failure.

What hurt

No iterative checkpoint loop

One submission gave the agent no opportunity to use public feedback to repair the artifact.

What remains unknown

Whether the model or workflow dominated

A single seed with a single candidate is too little evidence to characterize Inkling's repeatable capability.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-00750%59:08$2.452.40 MB1Valid
2.1MInput tokens
19.5KOutput tokens
1.6MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 1 candidate submission actions across the published run set. Selected artifacts averaged 2.40 MB compressed.

03

Time use

The runs averaged 59:08 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 101 agent messages · 1 candidate actions · 1 published run.

06 · Economics

What the result cost

$2.45 per displayed run at current configured rates.

Inkling spent $2.45 and almost the full hour without producing a score. The result is economically weak, but its larger lesson is the cost of not iterating through the runner.

At current configured API-equivalent rates, the displayed run cost is $2.45 and value is 0 score points per dollar.

#22
Value positionValue rank #22 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Inkling with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard