Chose direct inference
The run treated the benchmark as a family of contextual transformations rather than a pure language-model training task.
Run postmortem · Published 2026-08-14
Twenty-nine submissions expanded a solver that never became broad enough.
One-hour track · xhigh reasoning · provisional
Kimi K2.7 Code built a compact rule-driven inference system and submitted candidates twenty-nine times. The solver could handle several visible task families, but it never moved beyond 52% public calibration and produced only three hidden exact matches.
One valid autonomous run produced a provisional 3% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The final artifact prioritized explicit solvers over a large trained model.
It implemented records, mappings, arithmetic, sorting, sequences, time operations, text transforms, and odd-one-out logic.
The selected artifact was only 1.84 MB, leaving size headroom for more coverage.
Twenty-nine candidates show that the run repeatedly patched and tested the solver rather than waiting for a final build.
The run treated the benchmark as a family of contextual transformations rather than a pure language-model training task.
The agent kept adding solver branches and resubmitting, but public accuracy stalled near half.
The hidden exam returned 3%, confirming that the visible parser remained brittle.
Direct contextual solving was the same broad mechanism used by the benchmark's stronger systems.
High submission volume without a strong public score suggests narrow fixes rather than robust abstractions.
The archive cannot show whether fewer, more general mechanisms would have transferred better than many special cases.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0045 | 3% | 52:13 | $3.77 | 1.84 MB | 29 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 29 candidate submission actions across the published run set. Selected artifacts averaged 1.84 MB compressed.
The runs averaged 52:13 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 192 agent messages · 29 candidate actions · 1 published run.
$3.77 per displayed run at current configured rates.
Kimi K2.7 Code was neither a cheap experiment nor a competitive scorer. Most of the spend funded repeated solver iteration that did not generalize.
At current configured API-equivalent rates, the displayed run cost is $3.77 and value is 0.8 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Kimi K2.7 Code with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard