Model run analyses
Open the run.
Keep the exam closed.
Evidence-based accounts of how frontier models used their time, what they built, what it cost, and what the final score can—and cannot—support.
Grok 4.6: the $9.80 run, explained
A fast hybrid build, the strongest observed single-run score, and 18 million cached-context tokens.
26%provisional
Claude Opus 5 on ARI Bench
An ambitious neural-symbolic system, perfect public calibration, and a valid hidden-exam result.
23%provisional
These pages disclose aggregate evidence and high-level agent behavior. Hidden questions, answers, private grading records, credentials, and raw private transcripts remain unpublished.