Run postmortem · Published 2026-08-14

Kimi K3max reasoning

Three seeds stayed within three points—even after one nearly lost its solver at the package boundary.

15.67/ 100
hidden exact

One-hour track · max reasoning · official

ARI score scaleexact-match accuracy
15.67%
0255075100
StatusOfficial
Valid seeds3 / 3
Score rank#10
Current cost$1.06
Points / $14.78
Average time58:22
01 · Verdict

Stable, economical, and operationally instructive

Kimi K3 produced one of the library's cleanest official results: 14%, 16%, and 17% across three max-effort seeds for about one dollar per run. One seed discovered that its helper solver was not visible inside the packaged artifact, corrected the boundary, and still recovered to 16%.

3 valid autonomous runs produced an official 15.67% mean hidden exact-match score.

#10
Score position15.67% on the common 0–100 scale. Tied valid scores share a rank.

3-seed protocol met · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Kimi K3 built

All three runs relied on compact deterministic solvers, with only limited neural training.

01

Self-contained contextual solver

The artifacts handled arithmetic, records, mappings, lists, sequences, text transformations, and extraction directly from the prefix.

02

Minimal neural dependence

Only one seed recorded a small training run; another final package was under half a megabyte.

03

Packaging-boundary repair

The first seed noticed that an external solver module was absent from the candidate bundle and moved the logic into the submitted inference path.

03 · Trajectory

How the run unfolded

Seed 1

Recovered from a packaging bug

A byte-identical evaluation exposed that the helper solver was not executing; the self-contained replacement finished at 16%.

Seed 2

Used the full hour

A slightly larger hybrid reached perfect calibration and scored 14%.

Seed 3

Smallest package, best score

The 0.47 MB solver finished at 17%, reinforcing that artifact size was not the main performance driver.

04 · Assessment

Good, bad, and unresolved

What worked

Strong three-seed consistency

A three-point spread and an official mean make Kimi K3 much easier to trust than similarly scored one-off runs.

What hurt

Public calibration still saturated

Two seeds reached 100% visible accuracy without a signal for choosing broader hidden generalization.

What remains unknown

How much max effort mattered

The row is stable at max, but there is no lower-effort Kimi K3 control to price the benefit of that setting.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-006916%56:51$0.955.34 MB4Valid
run-007014%60:10$1.093.66 MB4Valid
run-007117%58:06$1.140.47 MB5Valid
247.2KInput tokens
54.6KOutput tokens
5.4MCache-read tokens
01

Seeded evidence

The estimate is based on 3 valid runs spanning 14% to 17%, a 3-point spread.

02

Candidate behavior

The agent made 13 candidate submission actions across the published run set. Selected artifacts averaged 3.16 MB compressed.

03

Time use

The runs averaged 58:22 of wall time. 2 runs explicitly finalized before the one-hour limit.

Retained totals: 127 agent messages · 13 candidate actions · 3 published runs.

06 · Economics

What the result cost

$1.06 per displayed run at current configured rates.

Kimi K3 combines official replication with roughly one-dollar runs, making it one of the strongest evidence-per-dollar results rather than merely a cheap single seed.

At current configured API-equivalent rates, the displayed run cost is $1.06 and value is 14.78 score points per dollar.

#3
Value positionValue rank #3 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Kimi K3 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard