Recovered from a packaging bug
A byte-identical evaluation exposed that the helper solver was not executing; the self-contained replacement finished at 16%.
Run postmortem · Published 2026-08-14
Three seeds stayed within three points—even after one nearly lost its solver at the package boundary.
One-hour track · max reasoning · official
Kimi K3 produced one of the library's cleanest official results: 14%, 16%, and 17% across three max-effort seeds for about one dollar per run. One seed discovered that its helper solver was not visible inside the packaged artifact, corrected the boundary, and still recovered to 16%.
3 valid autonomous runs produced an official 15.67% mean hidden exact-match score.
3-seed protocol met · Hidden exact-match accuracy, not public calibration accuracy.
All three runs relied on compact deterministic solvers, with only limited neural training.
The artifacts handled arithmetic, records, mappings, lists, sequences, text transformations, and extraction directly from the prefix.
Only one seed recorded a small training run; another final package was under half a megabyte.
The first seed noticed that an external solver module was absent from the candidate bundle and moved the logic into the submitted inference path.
A byte-identical evaluation exposed that the helper solver was not executing; the self-contained replacement finished at 16%.
A slightly larger hybrid reached perfect calibration and scored 14%.
The 0.47 MB solver finished at 17%, reinforcing that artifact size was not the main performance driver.
A three-point spread and an official mean make Kimi K3 much easier to trust than similarly scored one-off runs.
Two seeds reached 100% visible accuracy without a signal for choosing broader hidden generalization.
The row is stable at max, but there is no lower-effort Kimi K3 control to price the benefit of that setting.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0069 | 16% | 56:51 | $0.95 | 5.34 MB | 4 | Valid |
| run-0070 | 14% | 60:10 | $1.09 | 3.66 MB | 4 | Valid |
| run-0071 | 17% | 58:06 | $1.14 | 0.47 MB | 5 | Valid |
The estimate is based on 3 valid runs spanning 14% to 17%, a 3-point spread.
The agent made 13 candidate submission actions across the published run set. Selected artifacts averaged 3.16 MB compressed.
The runs averaged 58:22 of wall time. 2 runs explicitly finalized before the one-hour limit.
Retained totals: 127 agent messages · 13 candidate actions · 3 published runs.
$1.06 per displayed run at current configured rates.
Kimi K3 combines official replication with roughly one-dollar runs, making it one of the strongest evidence-per-dollar results rather than merely a cheap single seed.
At current configured API-equivalent rates, the displayed run cost is $1.06 and value is 14.78 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Kimi K3 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard