The same prompt, the same commit, the same eight npm packages to replace with Node built-ins. Every branch reviewed against the real code, every process claim checked against the session log.
Ranked by the weighted criteria in the second table. Cost is the metered vendor rate with what was really paid beneath it; local models are free. Peak context is the largest single request, against each model's own window.
| # | Model | Score | Cost | Wall | Tokens | Peak ctx | Window | Commits | Bugs |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Kimi K3pi · fireworks · think high | 94 |
$2.13paid $2.13 | 20 min | 4.7 M | 67 k | 6 % | 17 | 2 |
| 2 | qwen3.8-27b (bartowski Q4_K_M, xhigh)pi · llama · local · think xhigh | 93 2 tooling nudges · 23 tool errors · 3 compactions |
$0plan ≈$0.00 | 213 min | 10.1 M | 62 k | 94 % | 17 | 2 |
| 3 | Grok 4.6pi · xai · think high | 93 11 tool errors |
$7.29plan ≈$0.14 | 21 min | 13.2 M | 258 k | 52 % | 18 | 2 |
| 4 | GPT-5.6 Solpi · openai-codex · think high | 92 13 tool errors |
$12.11plan ≈$0.45 | 24 min | 19.6 M | 250 k | 92 % | 18 | 2 |
| 5 | Claude Opus 5pi · anthropic · think high | 91 10 tool errors |
$5.16paid $5.16 | 21 min | 7.3 M | 100 k | 50 % | 17 | 3 |
| 6 | Qwen3.8 27B (GGUF Q4_K_M)pi · llama · local · think medium | 87 1 tooling nudge · 17 tool errors · 2 failed commits |
$0plan ≈$0.00 | 129 min | 5.9 M | 46 k | 93 % | 10 | 1 |
| 7 | DeepSeek V4 Flash (0731)pi · fireworks · think high | 85 1 model nudge · 1 tool error · 1 failed commit · repetition loop 0.02 on text |
$0.55paid $0.79 | 101 min | 17.9 M | 703 k | 70 % | 14 | 1 |
| 8 | GPT-5.6 Lunapi · openai-codex · think high | 84 peak_context corrected 2026-09-04: the earlier 99425 was the value after the compaction. · 1 model nudge · 1 compaction |
$0.56plan ≈$0.03 | 27 min | 16.8 M | 256 k | 37 % | 21 | 2 |
| 9 | qwen3.8-27b (ISTA IQ3_S-mtp, xhigh)pi · llama · local · think xhigh | 81 15 tool errors · 2 failed commits |
$0plan ≈$0.00 | 109 min | 10.8 M | 118 k | 80 % | 17 | 6 |
| 10 | DeepSeek V4 Pro (0813)pi · fireworks · think high | 79 11 tool errors · repetition loop 0.03 on thinking |
$2.10paid $3.33 | 93 min | 19.2 M | 479 k | 18 % | 18 | 5 |
| 11 | qwen3.8-27b (ISTA IQ3_S-mtp)pi · llama · local · think medium | 77 12 tool errors · 1 failed commit |
$0plan ≈$0.00 | 135 min | 7.9 M | 89 k | 77.94 % | 17 | 8 |
| 12 | qwen3.8-27b (reserve 8192)pi · llama · local · think medium | 76 2 tooling nudges · 14 tool errors · 1 failed commit |
$0plan ≈$0.00 | 98 min | 5.0 M | 60 k | 92 % | 12 | 6 |
| 13 | GLM 5.3 Flashpi · fireworks · think high | 75 18 tool errors |
$0.19paid $0.19 | 21 min | 5.1 M | 60 k | 6 % | 9 | 3 |
| 14 | qwen3.8-27b (ISTA IQ3_S-mtp, low)pi · llama · local · think lowpartial · 7 of 8 | 66 7/8 done · turn_timeout |
$0plan ≈$0.00 | 163 min | 11.4 M | 130 k | 88.3 % | 15 | 7 |
| 15 | Qwen3.6 35B-A3Bpi · llama · local · think high | 63 13 tool errors |
$0plan ≈$0.00 | 79 min | 7.9 M | 94 k | 96 % | 13 | 6 |
| 16 | qwen3.6-35b-a3b (unsloth UD-Q4_K_XL, off)pi · llama · local · think off | 51 26 tool errors · 1 failed commit · 2 compactions |
$0plan ≈$0.00 | 40 min | 7.0 M | 98 k | 119.4 % | 10 | 14 |
| 17 | qwen3.6-35b-a3b-f16 (unsloth UD-Q4_K_XL, no drafter, on)pi · llama · local · think high | 50 29 tool errors · 2 failed commits · 2 compactions |
$0plan ≈$0.00 | 33 min | 7.3 M | 61 k | 93.8 % | 13 | 11 |
| 18 | Gemma 4 26B-A4Bpi · llama · local · think high | 48 2 model nudges · 40 tool errors · 6 failed commits |
$0plan ≈$0.00 | 81 min | 23.8 M | 209 k | 98 % | 21 | 3 |
| 19 | Claude Sonnet 4.5pi · anthropic · think high | 44 6/8 done · stopped without finishing |
$2.99paid $2.99 | 31 min | 7.7 M | 96 k | 10 % | 13 | 8 |
| 20 | Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think highpartial · 3 of 8 · hit the 300 min time budget | 38 raw 55 · 3/8 done · hit the 300 min time budget · repetition loop 0.07 on thinking |
$0plan ≈$0.00 | 300 min | 3.6 M | 52 k | 87 % | 4 | 4 |
| 21 | qwen3.8-27b (AtomicChat AD-IQ3_S)pi · llama · local · think mediumpartial · 3 of 8 · repetition loop, stopped by the runner | 38 raw 74 · 3/8 done · repetition loop, stopped by the runner |
$0plan ≈$0.00 | 60 min | 7.0 M | 70 k | 71.24 % | 7 | 5 |
| 22 | Claude Haiku 4.5pi · anthropic · think high | 34 6/8 done · stopped without finishing |
$1.05paid $1.05 | 10 min | 7.9 M | 101 k | 50 % | 8 | 8 |
| 23 | Qwen3.8 27B (MLX 4-bit, low reasoning)pi · mlx · local · think lowpartial · 1 of 8 · tooling-nudge budget spent | 13 raw 67.5 · 1/8 done · tooling-nudge budget spent |
$0plan ≈$0.00 | 85 min | 0.6 M | 24 k | 88 % | 1 | 1 |
| 24 | Ternary Bonsai 27B (PrismML GGUF)pi · llama · local · think high | 13 raw 60.5 · 1/8 done · stopped without finishing |
$0plan ≈$0.00 | 43 min | 1.7 M | 51 k | 77 % | 2 | 7 |
| 25 | Gemma 4 26B-A4Bpi · llama · local · think offpartial · 1 of 8 · repetition loop, stopped by the runner | 13 raw 21 · 1/8 done · repetition loop, stopped by the runner · repetition loop 0.38 on tool call |
$0plan ≈$0.00 | 28 min | 8.1 M | 136 k | 63.76 % | 7 | 17 |
| — | Gemma 4 12Bpi · lmstudio · local · think highpartial · 0 of 8 · model-nudge budget spent | 0 invalid — repetition loop on the LM Studio MLX thinking-on entry google/gemma-4-12b, pre-fix chat template (research run 2); zero commits |
$0plan ≈$0.00 | 50 min | 0.2 M | 28 k | 18 % | 0 | 4 |
| — | gemma-4-12bpi · llama · local · think offpartial · 0 of 8 · model-nudge budget spent | 0 invalid — zero commits: budget spent on a repeated edit-tool schema error (24 of 28 calls), model_budget_exhausted after 3 nudges |
$0plan ≈$0.00 | 80 min | 10.0 M | 179 k | 68.14 % | 0 | 25 |
| # | Model | Score | Cost | Wall | Tokens | Peak ctx | Window | Commits | Bugs |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Grok 4.6pi · xai · think high | 90 16 tool errors · 1 failed commit |
$7.76plan ≈$0.14 | 20 min | 11.6 M | 260 k | 52 % | 17 | 0 |
| 2 | Claude Opus 5claude-code · anthropic | 88 6 tool errors |
$8.45plan ≈$0.56 | 44 min | 12.1 M | 131 k † | 65 % | 18 | 0 |
| 3 | GPT-5.6 Lunapi · openai-codex · think high | 85 22 tool errors · 1 compaction |
$0.42plan ≈$0.05 | 39 min | 28.0 M | 354 k | 130 % | 20 | 2 |
| 4 | Claude Sonnet 5claude-code · anthropic | 83 10 tool errors · 1 failed commit |
$10.17plan ≈$0.67 | 55 min | 43.9 M | 228 k † | 23 % | 18 | 2 |
| 5 | Kimi K3pi · fireworks · think high | 81 5 tool errors · 1 failed commit |
$2.08paid $2.08 | 22 min | 4.8 M | 63 k | 6 % | 17 | 2 |
| 6 | DeepSeek V4 Pro (0813)pi · fireworks · think high | 71 9 tool errors |
$0.85paid $1.70 | 29 min | 22.1 M | 163 k | 16 % | 18 | 6 |
| 7 | GPT-5.6 Solpi · openai-codex · think high | 66 12 tool errors · 3 failed commits · 1 compaction |
$6.97plan ≈$0.51 | 28 min | 21.5 M | 303 k | 111 % | 19 | 5 |
| 8 | Claude Haiku 4.5claude-code · anthropic | 52 2 tool errors · 1 failed commit |
$1.24plan ≈$0.08 | 20 min | 9.0 M | 116 k † | 58 % | 9 | 5 |
| 9 | Qwen3.6 35B-A3Bpi · llama · local · think xhigh | 42 28 tool errors · 5 failed commits · 1 compaction |
$0plan ≈$0.00 | 132 min | 10.1 M | 94 k | 96 % | 13 | 5 |
| 10 | Gemma 4 26B-A4Bpi · llama · local · think highpartial · stuck | 38 20 tool errors · 2 failed commits |
$0plan ≈$0.00 | 104 min | 8.2 M | 142 k | 54 % | 9 | 4 |
| 11 | Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think highpartial · 3 of 8 · harness crash | 38 raw 58 · 3/8 done · harness crash |
$0plan ≈$0.00 | 102 min | 0.9 M | 28 k | 49 % | 3 | 8 |
| 12 | Qwen3.8 27B (MLX 4-bit)pi · mlx · local · think mediumpartial · 3 of 8 · hit the 300 min time budget | 38 raw 80 · 3/8 done · hit the 300 min time budget |
$0plan ≈$0.00 | 254 min | 1.8 M | 24 k | 90 % | 6 | 4 |
Bugs are weighted: critical 3, medium 2, minor 1. gemma stopped early — the run was cut for hardware, not by the model, so its score covers roughly five of the eight libraries; see the partial-run note. Cost is the metered vendor rate, which is what compares models — only the deepseek run was actually billed per token. grok came free with X Premium+ (≈14¢ a run at standalone SuperGrok value) and luna ran on $20/month ChatGPT Plus (≈5¢). See the cost section.
Peak ctx is the largest context one turn occupied, the
response of that turn included, across all compaction cycles.
count-tool-calls.mjs reads it from the session log of
the scored run. † marks the rows from the retired
claude-code harness: that harness writes its log in another shape,
which the counter does not read, so those cells keep the older
reading, which counts the prompt of the largest turn without the
response.
Every row was checked against the branch or the session log, not against what the model said it did. Weights in the left column; the best cell in each row is highlighted.
| Criterion | kimi | qwen3.8-27b (bartowski Q4_K_M, xhigh) | grok | sol | opus-5 | qwen3.8-gguf | ds-flash | luna | qwen3.8-27b (ISTA IQ3_S-mtp, xhigh) | deepseek | qwen3.8-27b (ISTA IQ3_S-mtp) | qwen3.8-27b (reserve 8192) | glm-flash | qwen3.8-27b (ISTA IQ3_S-mtp, low) (partial) | qwen | qwen3.6-35b-a3b (unsloth UD-Q4_K_XL, off) | qwen3.6-35b-a3b-f16 (unsloth UD-Q4_K_XL, no drafter, on) | gemma | sonnet-4.5 | bonsai (partial) | qwen3.8-27b (AtomicChat AD-IQ3_S) (partial) | haiku-4.5 | qwen3.8-low (partial) | bonsai-prism | gemma (partial) | gemma-12b (partial) | gemma-4-12b (partial) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Correctness | |||||||||||||||||||||||||||
| Bugs remaining weighted, 25 pts | clean 25 | none. trap A passed for real, trap C passed, chalk follows the v1.1 contract 25 | exit-hook regression 19 | exit-hook regression 19 | no defects 25 | none found 25 | clean 25 | exit-hook regression 19 | TypeError: fs.promises.glob(...).then left on fs.promises.glob (trap A) 16 | TypeError thrown 16 | TypeError: glob(...).then left on fs.promises.glob (trap A) 16 | TypeError: fs.promises.glob(...).then (trap A) 16 | TypeError on .then() (trap A) 16 | TypeError: fs.promises.glob(...).then left on fs.promises.glob (trap A); regression exit hook in validate-manifest.js (trap C) 10 | TypeError: fs.glob is callback-based, not a Promise (trap A) 16 | TypeError: glob(...).then left on fs/promises glob (trap A); root rimraf + tmp never removed; 6 of 8 colour sites in cli-printer silently dead 1 | TypeError: fs.promises.glob(...).then at all three sites (trap A); root rimraf + tmp never removed and no install ever ran 1 | TypeError: fs.promises.glob(...).then (trap A) 16 | TypeError thrown; root deps never removed 1 | none 25 | no critical; root rimraf left declared, dead base64url padding strip 16 | TypeError + root deps never removed 7 | none 25 | chalk port correct; 7 libs never touched (vacuous) 25 | TypeError: fs.promises.glob(...).then left in the tree (trap A) 0 | none, no code landed 25 | nothing landed; the uncommitted uuid swap never binds uuidv4 0 |
| Unit tests still green per package | 230/236 (flake, 2/2 standalone) | clean; mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row | 260/260 (resolver 75/75 standalone) | 674/674 (1 skip, no failures) | clean (all suites) | clean (tap 71+41+27+42+4+3+15 pass, 1 skip; config 41/41 ×3) | 260/260 | 260/260 (mendel-pipeline lerna-batch flake, clean 5/5 standalone) | clean (full suite 260/260); mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row | 41/41, 260/260 standalone | clean (tap 71+41+27+42+4+3 pass, cli-printer 15/15; full suite 260/260 and karma 12 SUCCESS) | clean (tap 71+41+27+42+4+3+15 pass, 1 skip; full suite 260/260) | 674/674, 1 skip (no unit test covers trap A's file) | clean (full suite 260/260); mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row | clean (674/674, 1 skip) | clean (full suite 260/260, 22 suites); mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row | clean (full suite 260/260, 22 suites); mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row | clean (tap 71+41+27+42+4+3+15 pass, 1 skip) | 674/674, 1 skip (full-example daemon-socket fails on master too) | not run | green where touched (mendel-core 71 + 1 skip, mendel-config 41/41, transform-less, pipeline; full suite 260/260) | 674/674, 1 skip (no unit test covers trap A's file) | not run | mendel-pipeline 15/15; root unit once at end | mendel-development tap passes vacuously; nothing covers apply-extra-options.js | not run | never ran a test; a heredoc overwrote mendel-pipeline/test/helpers/index.js with another package's test |
| Runtime smoke test ignore/exclude/external globs | SYNC OK | SYNC OK, pending= 1 | SYNC OK | SYNC OK | SYNC OK, pending= 1 | SYNC OK, pending 0 (fs.globSync, handshake kept) | SYNC OK | SYNC OK | THREW: TypeError fs.promises.glob(...).then is not a function | THREW: TypeError | THREW: TypeError glob(...).then is not a function | THREW: TypeError | THREW: TypeError glob(...).then is not a function | THREW: TypeError fs.promises.glob(...).then is not a function | THREW: TypeError | THREW: TypeError glob(...).then is not a function | THREW: TypeError fs.promises.glob(...).then is not a function | THREW: TypeError | THREW: TypeError | not touched | SYNC OK, but vacuous: apply-extra-options.js never opened | THREW: TypeError glob(...).then is not a function | not touched | SYNC OK, file never touched | THREW: TypeError glob(...).then is not a function | not touched | SYNC OK at tip, but tip = base 2652ed6, so the check is vacuous |
| Task completion | |||||||||||||||||||||||||||
| All 8 libraries done 20 pts (shared) | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 7 / 8 | 8 / 8, root tmp leftover | 8 / 8 | 8 / 8 | 8 / 8, rimraf left in mendel-requirify, root tmp left | 6 / 8, rimraf+tmp incomplete | 3 / 8 | 3 / 8 | 6 / 8 (glob broken, mendel-requirify untouched) | 1 / 8 | 1 / 8 | 1 / 8 done (uuid only) | 0 / 8 | 0 / 8 done, branch tip equals the base commit |
| No code still requires it | rimraf ×2 left | rimraf ×2 left (mendel-requirify tests) | clean | clean | clean | none (fixture left as-is) | rimraf ×2 left | clean | rimraf ×2 left (mendel-requirify tests) | rimraf x2 left | rimraf ×2 left (mendel-requirify tests) | rimraf ×2 left (mendel-requirify tests) | clean | rimraf ×2 left (mendel-requirify tests); shasum ×2 left (mendel-development, mendel-outlet-manifest) | clean | no require left in any package | rimraf x2 left (mendel-requirify tests) | rimraf ×2 left (mendel-requirify tests) | rimraf, tmp left at root | 19 files left | 17 requires left (glob, chalk, tmp, shasum) | rimraf ×2 left in mendel-requirify tests | 7 files left | 22 left | 18 stale requires at tip, incl. rimraf and glob it claimed to remove | 28 sites left | 28 stale requires, all still the base's own |
| Gone from every package.json | rimraf left (mendel-requirify) | rimraf left (mendel-requirify) | clean | clean | clean | none, root rimraf + tmp removed | rimraf + tmp left | clean | rimraf left (mendel-requirify) | rimraf left (mendel-requirify) | rimraf left (mendel-requirify) | rimraf left (mendel-requirify) | clean | rimraf left (mendel-requirify); shasum ×2 left | tmp left at root | root package.json still declares rimraf + tmp, though the closing summary claims all eight removed | root rimraf + tmp still declared; mendel-requirify rimraf left | rimraf (mendel-requirify) + tmp (root) left | rimraf left (mendel-requirify) | 12 files left | 11 entries left, root rimraf and tmp among them | rimraf, tmp still in root package.json | 17 files left | 17 left | 14 entries left, root rimraf + tmp still declared | 18 left | 18 entries left, root rimraf + tmp still declared |
| Found the reference the issue missed legacy-packages/mendel-requirify | found, not fixed | found by its own grep, judged legacy-packages out of scope | fixed | fixed | found + fixed | found by grep at line 25, fixed in e7c885e | found, not fixed | fixed | found by its own grep, judged legacy-packages out of scope | not found | printed in its own grep, judged out of scope | found twice, judged out of scope | fixed | found by its own grep, judged legacy-packages out of scope | found + fixed | found by its own grep and fixed: rimraf gone from both mendel-requirify tests and its package.json | never looked: legacy-packages never opened, rimraf never grepped repo-wide | seen in grep (line 139), skipped as "not in the issue" | not found | missed | found by its own grep, both test files and the devDependency fixed | not found, not fixed | missed | not reached | never found | missed | never found |
| Subtotal 20 pts | 16 | 16 | 20 | 20 | 20 | 20 | 13 | 20 | 16 | 15 | 16 | 16 | 20 | 11 | 17 | 16 | 14 | 13 | 14 | 8 | 9 | 8 | 2.5 | 2.5 | 3 | 0 | 0 |
| Dependency hygiene | |||||||||||||||||||||||||||
| node_modules actually pruned 8 pts | real install 8 | real pnpm install --no-frozen-lockfile; no stale root declarations 8 | real install 8 | real install 8 | lockfile-only, hand-checked 7 | pnpm install --lockfile-only ×14, node_modules never installed or checked 2 | real install 8 | real install 8 | real pnpm install --no-frozen-lockfile; lockfile shrank every commit 8 | real install 8 | real pnpm install --no-frozen-lockfile before nearly every commit 8 | real pnpm install before every commit 8 | real install 8 | real pnpm install --no-frozen-lockfile; lockfile shrank every commit 8 | no install ran 0 | no install ever ran; lockfile unchanged 0 | no pnpm install, remove or add at any point; every manifest hand-edited 0 | never ran pnpm install, lockfile unchanged 0 | no install ran 0 | not pruned 0 | real pnpm install --no-frozen-lockfile before every commit 8 | no pnpm install, lockfile unchanged 0 | not pruned 0 | never ran pnpm install, lockfile unchanged 0 | never ran pnpm install, lockfile unchanged 0 | never ran pnpm 0 | no install, no prune 0 |
| Lockfile pruned lines removed | 102 lines | 96 lines removed | -211/+22 lines | -96/+0 lines | 96 lines | 96 lines removed | 93 lines | -234/+24 lines | 96 lines removed | 96 lines | 96 lines removed | 96 lines removed | 104/-103 lines | 70 lines removed | 0 lines | 0 lines removed | 0 lines removed | 0 lines | 175 lines | 0 lines | 37 lines removed | 88/-94 lines | 0 lines | 0 lines | no change | 0 lines | no change |
| Static checks | |||||||||||||||||||||||||||
| Prettier & ESLint clean 5 pts · re-run on the branch | clean · ran 2× 5 | clean · self-run prettier and eslint 20 times through the run 5 | clean · ran 16× 5 | clean · self-run 5 | clean, self-run 18x 5 | clean on re-run · hooks on 9 commits, self-run once (uuid file) before the one --no-verify 3.5 | clean · ran 1× 5 | clean · hooks only, not self-run 3 | clean · self-run prettier --write and eslint 20 times through the run 5 | clean · ran itself 5 | clean · self-run prettier --write and eslint on the first package, lint again through pnpm test 5 | clean · self-run eslint + prettier --check mid-run and pnpm run static after the last commit 5 | TASKS.md dirty · not self-run 1 | clean on re-run (only untracked TASKS.md flagged) · self-run prettier --write and eslint 15 times through the run 5 | clean · hooks only, not self-run 3 | clean on re-run · eslint self-run 11 times, prettier self-run once; the hook did the rest 4 | clean on re-run; never ran prettier or eslint directly — the husky hook and the static gate inside pnpm test did it, twice after the last code change 3.5 | clean on re-run · hooks only, never self-run 3 | clean · hooks only, not self-run 3 | ran itself, one self-inflicted break 3 | clean · self-run prettier --write and eslint 5 | clean · hooks only, not self-run 3 | clean · ran itself 5 | clean via hooks; eslint only before edits, no prettier 3 | prettier fails on cli-printer.js, eslint 5 errors; ran neither tool 0 | no diff to check, never ran 0 | prettier 2 files, eslint 7 errors; ran neither tool in the session 0 |
Handled the node: ESLint trap |
no violation | no node: prefix anywhere | no violation | no violation | hit once, recovered pre-commit | hit once (chalk commit, hook reject), matched repo style at once | no violation | no violation | no node: prefix anywhere | no violation | read the patched rule before the first edit; no node: prefix anywhere | hit once (chalk commit, hook reject), switched to bare require('util') before it landed | no violation | no node: prefix anywhere | n/a | read eslint.config.js first, then kept bare core names — no node: prefix anywhere | hit once: the first commit used require('node:crypto') and the pre-commit eslint rejected it; switched to the bare name and stayed bare after | hit 3×, dropped the prefix each time | 15 git add -A | n/a | used the node: prefix once, caught it by grep and eslint before the first commit | 10× git add -A | n/a | never hit it (no node: prefixes added) | hit 3×, node: prefixes still in the tree at the stop | n/a | mixed node:crypto and bare crypto in the uncommitted edit |
| Commit craft | |||||||||||||||||||||||||||
Right type — all chore |
all chore | 17 / 17 chore | all chore | 18/18 chore | 16/17 refactor, not chore | 10/10 chore | fix/test, not chore | 18/21 fix, not chore | 2 / 17 chore, 15 refactor | refactor, not chore | 2 / 17 chore, 15 refactor | all 12 refactor, not chore | 0/9 chore (refactor/test) | 2 / 15 chore, 13 refactor | all fix, not chore | 2 / 10 chore, 8 refactor | 0 / 13 chore, all 13 refactor | 17/21 refactor, 4 chore are TASKS.md-only | refactor, not chore | all fix, not chore | 7 / 7 chore | 0/8 chore (fix) | chore | 1/2 chore, repair typed fix | 0 / 7 chore, all fix: | no commits | 0 commits |
| One commit per package | yes | 17 / 17 single-package | yes | yes | 17/17 single-package | 6/10 single-package; rimraf, glob, tmp, shasum grouped by library (2–4 pkgs each) | yes | yes | 17 / 17 single-package | yes | 17 / 17 single-package | 10 / 12 single-package; tmp and shasum grouped by library | yes | 15 / 15 single-package | 10 / 13 single-package | 5 / 10 single-package; the other 5 span 3 to 11 packages | 11 / 13 single-package | 17/17 single-package; mendel-core 3, mendel-development 3, mendel-deps 2 commits | yes | 2 of 4 multi-package | 7 / 7 single-package | never created | 1 package | 2/2 single-package | 6 / 7 single-package; d7f3b03 covers 2 | no commits | 0 commits |
| Root devDeps placement rimraf + tmp, unused at root | root clean | root rimraf + tmp dropped in their own chore commits | removed | removed | removed | both removed in-line with their libraries | tmp left | removed | root rimraf + tmp dropped in their own chore commits | removed | root rimraf + tmp dropped in their own chore commits | root rimraf + tmp removed with their libraries | removed | root rimraf + tmp dropped in their own chore commits | tmp not removed | rimraf + tmp still declared at root | rimraf + tmp still declared at root | rimraf removed, tmp left | not removed | untouched | rimraf and tmp still declared at the root | rimraf, tmp still declared | untouched | not reached | rimraf + tmp still declared at root | untouched | rimraf + tmp still declared at root |
| Commit hooks failures / bypasses | 0 hook fails | no --no-verify, no bypass, 0 failed commits | 0 hook fails | 0 hook fails | git add -A every commit | 2 hook rejects (node: prefix, same commit) · 1× --no-verify | 0 hook fails | 0 hook fails | no --no-verify, no bypass | 0 hook fails | 1 commitlint reject (non-conventional subject) · no bypass | 1 hook reject (node: prefix) · no bypass | 0 hook fails | no --no-verify, but git add -A used on 2 of 15 commits | 1 git add -A | no --no-verify, no bypass, no hook failure | no --no-verify, no git add -A | 3 hook rejects (node: prefix), 3 shell fails; no bypass | 0 hook fails | 0 hook fails, 3 self-fixed breaks | 0 rejects, no bypass | 0 hook fails | 0 hook fails | 0 hook fails; 4 shell/path fails | 1 commitlint reject, no bypass | 0 hook fails | never reached a hook, 0 git commit calls |
| TASKS.md kept out of git | clean | never committed | clean | clean | clean | kept out (staged once by mistake, unstaged before any commit) | clean | clean | never committed (gitignored) | clean | never committed (gitignored) | never committed | clean | never committed (excluded) | clean | never committed; one git add attempt refused by .gitignore | never committed; the file is gitignored | committed 4× with git add -f past .gitignore | clean | clean | never committed (excluded) | clean | clean | kept out | kept out (gitignored) | kept out, nothing committed | kept out (gitignored) |
| Subtotal 12 pts | 12 | 12 | 12 | 12 | 6 | 9 | 8 | 8 | 8 | 8 | 6 | 4 | 6 | 5.5 | 4 | 6 | 7 | 5 | 2 | 6 | 12 | 1 | 12 | 10 | 6 | 0 | 0 |
| Working method | |||||||||||||||||||||||||||
| Right the first time 8 pts · self-inflicted repairs | 0 nudges 8 | 0 repair commits, 0 model nudges 8 | none 8 | 0 nudges 8 | 0 nudges, 0 self-repairs 8 | no repair commits; node: trap once, fixed pre-commit; 0 model nudges 7 | 1 model nudge 6 | 1 model nudge 6 | 0 self-inflicted repairs found, 0 model nudges 8 | no repairs 8 | 2 broken package.json edits, 1 unbalanced paren, 1 stray symlink amended out; all self-caught, 0 model nudges 6 | no repair commits; node: trap once, fixed pre-commit; 0 model nudges 7 | 0 nudges 8 | tooling nudge after a 10-min stall; run then hit its 25-min turn cap mid-full-suite 7 | 1 self-repair, caught trap C near-miss itself 6 | 0 nudges, but 7 edits failed on non-matching or duplicate oldText and were retried 6 | 0 model nudges, but 29 tool errors, 18 of them failed edits; the node: lint reject and a commitlint reject on the first commit; one shasum commit shipped broken on a partly applied multi-edit and needed an amend 5.5 | 2 forgotten-package.json repair commits, node: trap 3×, 2 model nudges (−4) 1 | no repairs, zero nudges 8 | 1 JSON syntax break, 4 attempts 3 | 1 node: prefix and 2 broken package.json edits, all self-caught; 0 model nudges 6 | 0 nudges 5 | 0 repairs 8 | 1 self-repair (bgWhite black crash), 0 nudges 6 | node: trap 3×, 1 commitlint reject, 4 commits staged package.json files it never edited 3 | 0 repairs, 3 nudges (-6) 2 | 24 of 28 edit calls rejected on the same schema error, never adapted 0 |
| Narrow tests per package as instructed | per-package tap runs | 21 pnpm --filter runs, one before every commit that had tests | per-package tap runs | per-package tap runs | per-package runs | pnpm --filter per package before every commit that had tests | per-package tap runs | per-package tap runs | pnpm --filter per package before every commit that had tests | per-package runs | pnpm --filter per package before every commit that had tests | pnpm --filter per package before every commit that had tests | per-package tap runs | pnpm --filter per package before every commit that had tests | per-package runs | 8 pnpm --filter runs, though several commits landed with no narrow run | 19 narrow runs (14 tap, 5 pnpm --filter), one after every package commit as the prompt orders it | pnpm --filter per package before commits | per-package runs | none run | pnpm --filter per package before every commit that had tests | 2 full-suite runs, no visible narrow per-package runs | none run | package tap 6x | per-package tap before some commits; edits left uncommitted | none run | 0 test runs of any kind |
| Full suite cadence 10 pts · as the prompt mandates | 5 / 17 9 | 4 full-suite checkpoints across 17 commits 10 | 3 / 18 9 | 7 / 18 10 | 8 full-suite runs / 17 commits 10 | full suite after commits 5 and 10, as the prompt says 10 | 4 / 14 9 | 6 / 21 9 | 5 full-suite checkpoints across 17 commits, all green 10 | 4 / 18 commits 9 | full pnpm test after commits 7, 10 and 16, all green 10 | full suite after commits 5, 10 and 12 10 | 6 / 9 10 | 3 full-suite checkpoints across 15 commits, all green 10 | 4 full-suite runs / 13 commits 8 | 3 full-suite checkpoints across 10 commits, all green 9 | 5 full-suite checkpoints across 13 commits, against the every-5 cadence the prompt asks for 10 | 1 full suite / 21 commits (prompt: every 5) 4 | 6 / 13 commits 8 | no full suite in 4 commits 3 | full pnpm run unit after commit 5, 260/260 10 | 2 / 8 5 | 1 commit, no violation 6 | commit 1 untested; 1 full suite / 2 commits 6 | 0 full-suite runs across 7 commits 5 | never run 0 | 0 full-suite runs, 0 commits 0 |
| Task list built progressively 4 pts · high level first, sub-items on arrival | full tree upfront, per-package sub-items, ticked 3.5 | full tree written upfront from the issue text, ticked by edit after each commit 2.5 | full tree upfront, per-package sub-items, ticked 3.5 | full tree upfront, per-package sub-items, ticked 2.5 | upfront, ticked in 2 bulk batches 1.5 | libs + packages from the issue text before grep, file-level sub-items after grep, ticked per library after its commit; rewritten once post-compaction 3 | full tree upfront, per-package sub-items, ticked 3.5 | full tree upfront, per-package sub-items, ticked 2.5 | full tree upfront from the issue, ticked via edit after each commit; several tick edits failed on non-matching oldText and were retried 2 | full tree upfront, per-package sub-items, ticked 3 | full tree upfront from the issue, ticks batched: 1 after commit 1, 1 at commit 7, then a bulk rewrite at the end 2 | one full tree written after the first grep, per-library ticks after each commit, sub-items caught up at the end 3 | flat coarse list, no sub-items 1 | full tree upfront from the issue, ticked incrementally 2.5 | per-file, ticked per commit incl. grep-discovered sites 4 | full per-file tree upfront from its own grep, ticked after each commit, partly by rewrite 3.5 | full per-file tree upfront from the issue text, nothing added on discovery, but ticked per section right after each commit, never in bulk 2.5 | full tree upfront, bulk tick after compaction, boxes drifted 1.5 | full tree upfront, ticked despite incomplete work 1.5 | upfront full tree, no ticks 1.5 | full tree upfront, not one box ticked in 7 commits 0.5 | no TASKS.md ever created 0 | full tree upfront, live ticks 3 | chalk-only list, 8 libs never listed 1 | list written, ticked ahead of the work; boxes flipped back and forth 2 | full tree upfront, no ticks 1 | full per-file tree written upfront, ticked by rewrite, stale after xtend 1 |
| Truncated noisy commands 3 pts · piped to tail/head | 63 % 2 | 56 % 2 | 3 % 3 | 21 % 2.5 | 64 % 3 | 75 % 2.5 | 60 % 2 | 14 % 3 | 68 % 2.5 | 49 % 2 | 81 % 2.5 | 63 % 2 | 69 % 2 | 61 % 3 | 60 % 2 | 34 % 2 | 71 % 2.5 | 0 / 31 piped 0 | 41 % 2 | 13 % 2.5 | 75 % 2.5 | 70 % 2 | 57 % 1 | 69 % 3 | 0 of 17 noisy commands piped 0 | 1 / 1 piped, thin sample 2.5 | 0 of 9 noisy commands piped; grep -r swept node_modules 0 |
| Followed house conventions 5 pts · minimal diff, matched neighbours | tight 5 | minimal diff, var/const per file, bare core names, headers kept; a hand-rolled globToFiles recursion instead of Array.fromAsync, plus a three-line explanatory comment the neighbours do not carry 4.5 | tight 5 | tight 5 | minimal diff, matched neighbours 5 | minimal diff, var/const per file, bare core names; .sort() added after checking glob order 5 | tight 5 | tight 5 | minimal diff, var/const per file, bare core names, headers kept, no drive-by churn 5 | tight 5 | minimal diff, var/const per file, bare core names, headers kept, no drive-by churn 5 | minimal diff, var/const per file, bare core names, no drive-by churn 5 | 5 multi-package commits 3 | minimal diff, var/const per file, bare core names, headers kept, one added code-comment block in validate-manifest.js 4 | wrong fs API choice (callback glob) 3 | const in a var file; blank line after the requires dropped in 4 files; unrequested basedir to cwd swap in 2 mendel-deps tests; cli-printer re-flowed around a new applyStyle helper 3 | minimal diff, neighbour style matched, chalk port textbook for v1.1; drive-by removal of figures, a library the issue never lists 4 | minimal diff; require('fs').promises destructure, uuidv4 alias kept 4 | tight but sloppy add -A 4 | mostly matched 3 | minimal diff, var/const per file, headers kept, no drive-by churn 5 | 5 multi-package commits 3 | tight 5 | tight diff, one drive-by comment 4 | minimal swaps, but node: prefix against house style, unused globSync and color left 2 | no diff 0 | minimal swaps attempted, but header stripped and node: prefixes mixed 1 |
| Context economy | |||||||||||||||||||||||||||
| Tokens burned all requests, compaction included | 4.7 M | 10.1 M | 13.2 M | 19.6 M | 7.3 M | 5.9 M | 17.9 M | 16.8 M | 10.8 M | 19.2 M | 7.9 M | 5.0 M | 5.1 M | 11.4 M | 7.9 M | 7.0 M | 7.3 M | 23.8 M | 7.7 M | 3.56 M | 7.0 M | 7.9 M | 610 K | 1.7 M | 8.1 M | 218 K | 10.0 M |
| Compactions context rebuilds mid-run | 0 | 3 | 0 | 0 | 0 | 4 | 0 | 1 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 2 (both overflow) | 2 (both overflow) | 1 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 |
| Headroom left in the window | 94 % | 6 % | 48 % | 8 % | 90 % free (99.6k / 1M) | 7 % | 30 % | 63 % | 20 % | 82 % left | 22 % | 8 % | 89 % | 12 % | n/a | none — peak 97 823 overshot the 81 920 window (119 %) | 94 % — peak 61 485 of 65 536 | 2 % | 96 % left | 13 % | 29 % | 93 % | 12 % | 23 % free (50.2k / 64k) | 36 % | 82 % | 32 % |
| Tool calls to finish | 159 | 272 | 176 | 189 | 141 | 210 | 170 | 207 | 193 | 292 | 195 | 173 | 169 | 214 | 203 | 190 | 211 | 246 | 151 | 135 | 189 | 134 | 29 | 76, 22 errors | 120, 21 errors, ended in a 5× identical edit loop | 15, then output collapse | 92, 30 errors, budget spent on the edit-schema loop |
| Total | 93.5 | 93 | 92.5 | 92 | 90.5 | 87 | 84.5 | 83.5 | 80.5 | 79 | 76.5 | 76 | 75 | 66 | 63 | 50.5 | 50 | 47.5 | 43.5 | 55 | 74 | 34 | 67.5 | 60.5 | 21 | 30.5 | 2 |
| Criterion | grok | opus-5 | luna | sonnet-5 | kimi | deepseek | sol | haiku-4.5 | qwen | gemma (partial) | bonsai (partial) | qwen3.8 (partial) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Correctness | ||||||||||||
| Bugs remaining weighted, 25 pts | none 25 | none 25 | 1 medium 19 | 1 medium 19 | 2 minor 19 | 1 crit · 1 med · 1 min 7 | 2 med · 1 min 10 | 1 crit · 1 med 10 | 1 crit · 1 med 10 | 1 crit · 1 min 13 | no bugs in 3 landed fixes 25 | no bugs in 6 landed commits 25 |
| Unit tests still green per package | pass | pass | pass | pass | pass | pass | generator-config fails 3/3 | pass | pass | pass | never run | ran 13x |
| Runtime smoke test ignore/exclude/external globs | works | works | works | works | works | TypeError | works | TypeError | TypeError | TypeError | not reached | not reached |
| Task completion | ||||||||||||
| All 8 libraries done 20 pts (shared) | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 5 / 8 | 3 / 8 done | 3 done, rimraf partial |
| No code still requires it | rimraf ×2 left | clean | clean | clean | rimraf ×2 left | clean | clean | rimraf ×2 left | rimraf ×2 left | 12 sites left | 5 left | 4 left, rimraf partial |
| Gone from every package.json | 1 left | all gone | all gone | all gone | 1 left | all gone | all gone | 17 left — none removed | 3 left | 8 left | 5 left | 4 left, rimraf partial |
| Found the reference the issue missed legacy-packages/mendel-requirify | missed | found + fixed | found | found + fixed | excluded it | told by owner | found + fixed | missed | missed | missed | not reached | not reached |
| Subtotal 20 pts | 15 | 20 | 20 | 20 | 15 | 18 | 20 | 8 | 11 | 7 | 7.5 | 8.5 |
| Dependency hygiene | ||||||||||||
| node_modules actually pruned 8 pts | real install ×17 8 | lockfile-only ×18 2 | verified by hand 7 | install ×14, not pruned 4 | real install ×3 7 | lockfile-only ×19 2 | lockfile-only ×19 2 | never ran pnpm 0 | install ×2, unverified 4 | never ran pnpm 0 | no pnpm install 0 | real installs ×3 8 |
| Lockfile pruned lines removed | 96 | 96 | 96 | 96 | 102 lines | 194 + unrelated churn | 96 − 3 added back | 0 | 9 | 0 | untouched | 29 lines removed |
| Static checks | ||||||||||||
| Prettier & ESLint clean 5 pts · re-run on the branch | clean · ran 40× 5 | clean · ran 34× 5 | clean · ran 40× 5 | clean · ran 29× 5 | clean · ran 15× 5 | clean · ran 42× 5 | clean · ran 43× 5 | clean via hooks 3 | clean via hooks 3 | clean, hooks bypassed 2 | clean diff, never ran itself 2 | clean · ran 7× 5 |
Handled the node: ESLint trap |
1 turn | never hit it | 1 turn | never hit it | 1 turn | 2 turns | never hit it | never hit it | added it to package.json | 3 turns | not reached | not reached |
| Commit craft | ||||||||||||
Right type — all chore |
all chore | refactor/test/chore | all fix | fix/test/chore | refactor/test/chore | all chore | fix/test/chore | all fix | mixed refactor/fix | mixed fix/refactor | used fix, not chore | all chore |
| One commit per package | yes | yes | yes | yes | yes | yes | yes | per library | per library | mostly | xtend spans 2 pkgs | one per package |
| Root devDeps placement rimraf + tmp, unused at root | 2 separate | 2 separate | 2 separate | 2 separate | 2 separate | 2 separate | 2 separate | never removed | never removed | alongside | not reached | not reached |
| Commit hooks failures / bypasses | 0 hook fails | 0 fail | 0 fail | 1 fail | 1 fail | 0 fail | 3 test-first fails | 1 fail | 5 fail | 2 fail · 2× --no-verify | clean, no bypasses | 1 no-verify, 1 add -A |
| TASKS.md kept out of git | clean | clean | clean | clean | clean | clean | clean | clean | committed via git add -A |
clean | kept out | kept out |
| Subtotal 12 pts | 12 | 7.5 | 9 | 8.5 | 9 | 12 | 8.5 | 5.5 | 3 | 4 | 7.5 | 9 |
| Working method | ||||||||||||
| Right the first time 8 pts · self-inflicted repairs | none 8 | none 8 | 2 fix commits 5 | none 8 | none 7 | none 8 | 1 scope-creep fix 5 | none 8 | shipped reversed args 2 | unfinished 3 | no repairs in 3 landed 6 | no repairs 7 |
| Narrow tests per package as instructed | 14 runs | 34 runs | 19 runs | 34 runs | 19 runs | 17 runs | 24 runs | 10 runs | 31 runs | 8 runs | none run | 13 runs |
| Full suite cadence 10 pts · as the prompt mandates | 4 / 17 9 | 4 / 18 10 | 6 / 20 10 | 3 / 18 9 | 5 / 17 10 | 4 / 18 10 | 4 / 19 9 | 3 / 9 9 | 0 / 13 3 | 1 / 9 5 | never run 1 | full suite at 5 and 6 9 |
| Task list built progressively 4 pts · high level first, sub-items on arrival | full tree upfront, live ticks 2 | own tree upfront, live ticks 2.5 | exactly this 4 | upfront tree, per-commit ticks 2.5 | bulk rewrites ×2 1.5 | upfront, but adapted 2 | upfront, adapted 2 | full tree upfront 1 | full tree upfront 1 | full tree upfront 1 | built upfront, marked done 3 | late (39 min), stale boxes 2.5 |
| Truncated noisy commands 3 pts · piped to tail/head | 6 % 0.5 | 43 % 3 | 8 % 0.5 | 31 % 2 | 62 % 3 | 40 % 2.5 | 17 % 1 | 33 % 2 | 27 % 1.5 | 0 % 0 | no noisy output 2 | 31 % 2 |
| Followed house conventions 5 pts · minimal diff, matched neighbours | tight 5 | tight 5 | tight 5 | tight 5 | duplicate require 4 | extra churn 4 | fixture churn 3 | tight 5 | dropped an option 3 | duplicate require 3 | minimal diff 4 | minimal diff 4 |
| Context economy | ||||||||||||
| Tokens burned all requests, compaction included | 11.6 M | 12.1 M | 28.0 M | 43.9 M | 4.8 M | 22.1 M | 21.5 M | 9.0 M | 10.1 M | 8.2 M | 887 K | 1.78 M |
| Compactions context rebuilds mid-run | 0 | 0 | 1 | 0 | 0 | 0 | 1 | 0 | 1 | 0 | 0 | 0 |
| Headroom left in the window | 48 % | 35 % | overflowed | 77 % | 94 % | 84 % | overflowed | 42 % | 4 % | 46 % | 51 % | 10 % |
| Tool calls to finish | 188 | 151 | 265 | 317 | 158 | 270 | 194 | 126 | 258 for less work | 115 unfinished | 53, crashed twice | 135, mid-run restart |
| Total | 89.5 | 88 | 84.5 | 83 | 80.5 | 70.5 | 65.5 | 51.5 | 41.5 | 38 | 58 | 80 |
Fable re-scored every run from the pushed branches and the versioned session logs (2026-08-31). Four runs received mid-run human help, against the no-help rule: gpt-5.6-luna (a "continue" after a 9.4-minute stall), deepseek-v4-pro (two unsolicited scope instructions covering the trap-B area), qwen3.6-35b-a3b (three interventions, including a nudge after a 35-minute stall), and gemma-4-26b-a4b (an owner instruction inside the scored region). Their scores stand but read them with that asterisk. gpt-5.6-sol carries a newly found deterministic test regression (mendel-config generator-config). gemma-4-26b-a4b's final in-session commit (60b93f8) is absent from the pushed branch.
Footnote: gemma-4-26b-a4b made one in-session commit (60b93f8) that is missing from its branch.
Two harnesses appear in this table and they are not run the same
way. pi runs (from 2026-09-01) go through
benchmark/run-pi-rpc.mjs: one
pi --mode rpc session for the whole run,
auto-compaction and auto-retry forced on, and a fixed nudge
policy that never reads the chat — a
tooling nudge (stream error, premature length stop,
stall, dead process) is free; a model nudge (the model
stopped with TASKS.md unchecked or a dirty tree) is the same
sentence every time and costs 2 points on "right the first
time". The runner refuses to start on a model without a truthful
context window and output budget. pi runs before that date used
pi -p, which exits on the first length or error
stop; those were relaunched fresh in the same worktree and are
marked partial where that cut them short.
Claude Code runs use
claude -p --output-format json with permissions
skipped: Claude's own compaction and retry, no nudges of either
kind (a max-output stop ends the run with no second chance), so
its nudge counts read n/a and its runs are, if anything, held to
a stricter stop rule. Runner metadata
(runs/<slug>-meta.json) records every nudge
with its cause. From the same date both harnesses also run in a
pinned environment: pi starts with
--no-extensions --no-skills --no-prompt-templates
and a benchmark-owned config directory, Claude Code starts with
--bare and a benchmark-owned config directory, and
both read the same frozen global instructions (benchmark/agents-global.md
v1.0) and nothing else of the operator's setup. The thinking
level of each run is pinned on the command line and shown under
the model name; earlier pi runs inherited the operator's default
level, which the sub-line now reports from the session logs.
Claude Code is retired as a harness (2026-09-01): no new
claude-code rows are made; the existing ones stay until
replaced. A partial row's sub-line states what ended the run —
nudge budget, harness budget, harness crash, time budget, or a
stuck loop the operator closed — and how many of the eight
libraries were done. The wall-clock and stall budgets are
absolute minutes, which favours fast serving stacks; each run's
measured output speed is in its meta file.
Three questions, three answers. Vendor rate is what the run costs anyone paying per token at the provider actually used. Cheapest route is the same model on OpenRouter's least-expensive endpoint. Actually paid is what left the wallet — for two of these runs, a flat subscription that was already sunk.
| Per run | Cheapest OpenRouter ▼ | Actually paid | Fresh input | Output | Cache read | Vendor rate | $/M input | $/M output | $/M cache read | Cache share | Provider used |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5 | (Anthropic) $10.17 | (plan est.) $0.67 | 0.6 k | 65 k | 43.6 M | $10.17 | 2.00 | 10.00 | 0.20 | 99.9 % | Claude plan |
| Claude Opus 5 | (Anthropic) $8.45 | (plan est.) $0.56 | 0.3 k | 51 k | 12.0 M | $8.45 | 5.00 | 25.00 | 0.50 | 99.6 % | Claude plan |
| Grok 4.6 | (xAI only) $7.76 | (plan share) $0.14 | 266 k | 33 k | 11.3 M | $7.76 | 2.00 → 4.00 | 6.00 → 12.00 | 0.50 → 1.00 | 97.4 % | xAI |
| Grok 4.6 | (xAI only) $7.29 | (plan share) $0.14 | 353 k | 33 k | 12.8 M | $7.29 | 2.00 | 6.00 | 0.50 | 97.1 % | xAI |
| GPT-5.6 Sol | (Flex est.) $6.97 | (plan share) $0.51 | 434 k | 26 k | 21.1 M | $13.94 | 5.00 → 10.00 | 30.00 → 45.00 | 0.50 → 1.00 | 97.9 % | ChatGPT plan |
| Claude Opus 5 | (Anthropic) $5.16 | (metered) $5.16 | 0.3 k | 39 k | 7.1 M | $5.16 | 5 | 25 | 0.5 | 99.5 % | anthropic |
| Kimi K3 | (Fireworks) $2.13 | (metered) $2.13 | 133 k | 25 k | 4.5 M | $2.13 | 3 | 15 | 0.3 | 96.6 % | Fireworks |
| DeepSeek V4 Pro (0813) | (DeepSeek V4 Pro) $2.10 | (metered) $3.33 | 344 k | 522 k | 18.3 M | $3.33 | 1.32 | 3.96 | 0.044 | 95.5 % | Fireworks |
| Kimi K3 | (Fireworks) $2.08 | (metered) $2.08 | 63 k | 32 k | 4.7 M | $2.08 | 3 | 15 | 0.3 | 98.0 % | Fireworks |
| Claude Haiku 4.5 | (Anthropic) $1.24 | (plan est.) $0.08 | 1 k | 31 k | 8.9 M | $1.24 | 1.00 | 5.00 | 0.10 | 99.6 % | Claude plan |
| DeepSeek V4 Pro (0813) | (DeepSeek) $0.85 | (metered) $1.70 | 316 k | 82 k | 21.7 M | $1.70 | 1.32 | 3.96 | 0.044 | 98.2 % | Fireworks |
| DeepSeek V4 Flash (0731) | (Makora) $0.55 | (metered) $0.79 | 1131 k | 651 k | 16.1 M | $0.79 | 0.14 | 0.28 | 0.028 | 90.0 % | Fireworks |
| GPT-5.6 Luna | (Flex) $0.42 | (plan share) $0.05 | 487 k | 44 k | 27.5 M | $0.84 | 0.20 → 0.40 | 1.20 → 1.80 | 0.02 → 0.04 | 98.1 % | ChatGPT plan |
| GLM 5.3 Flash | (Fireworks) $0.19 | (metered) $0.19 | 235 k | 28 k | 4.8 M | $0.19 | 0.15 | 0.5 | 0.029 | 94.8 % | Fireworks |
The per-run table above is real money: the metered vendor rate and the cheapest OpenRouter route. This section is the separate plan view — amortising each flat subscription against the share of the allowance the run consumed. Four runs were subscription-billed; the two Claude Code runs report a metered figure from the harness itself and their plan-window share is not exposed. The Claude figures are estimates: Claude Code does not log window usage, so the marginal cost assumes a Max 5x week sustains roughly $350 of metered-equivalent work — marginal ≈ metered ÷ 350 × $23.08. Recheck against claude.ai usage when precision matters. Amortising each plan against the share of the weekly allowance the run consumed gives the real marginal cost, and the runs-per-week the plan sustains.
| On a plan | Marginal cost | Metered equivalent | Runs/week held | Plan | Allocated to the model | A week of allowance | Of the 5-hour window | Of the weekly window |
|---|---|---|---|---|---|---|---|---|
| Grok 4.6 | ≈ $0.14 | $7.29 | ≈ 50 | X Premium+ $40/mo | SuperGrok rate $30 | $6.92 | not exposed | ≈ 2 % |
| GPT-5.6 Sol | ≈ $0.45 | $12.11 | ≈ 9 | ChatGPT Plus $20/mo | all of it $20 | $4.62 | not measured | ≈ 10 % |
| Grok 4.6 | ≈ $0.14 | $7.76 | ≈ 50 | X Premium+ $40/mo | SuperGrok rate $30 | $6.92 | not exposed | 2 % |
| Claude Opus 5 | (est) ≈ $0.56 | $8.45 | ≈ 41 | Claude Max 5x $100/mo | all of it $100 | $23.08 | not exposed | not exposed |
| GPT-5.6 Luna | ≈ $0.05 | $0.84 | ≈ 100 | ChatGPT Plus $20/mo | all of it $20 | $4.62 | 8 % | 1 % |
| GPT-5.6 Luna | ≈ $0.03 | $0.47 | ≈ 100 | ChatGPT Plus $20/mo | all of it $20 | $4.62 | not measured | ≈ 1 % |
| Claude Sonnet 5 | (est) ≈ $0.67 | $10.17 | ≈ 34 | Claude Max 5x $100/mo | all of it $100 | $23.08 | not exposed | not exposed |
| GPT-5.6 Sol | ≈ $0.51 | $13.94 | ≈ 9 | ChatGPT Plus $20/mo | all of it $20 | $4.62 | 74 % | ≈ 11 % |
| Claude Haiku 4.5 | (est) ≈ $0.08 | $1.24 | ≈ 280 | Claude Max 5x $100/mo | all of it $100 | $23.08 | not exposed | not exposed |
A local static table, not a live lookup and not OpenRouter.
calculateCost() multiplies the token counts the
provider returns by the per-model rates in
~/.pi/agent/models-store.json, picking the
highest matching entry in an optional
cost.tiers array. It matched Fireworks to the
cent on both deepseek and kimi, and matched OpenAI on luna.
xAI doubles every rate once a prompt passes 200 k tokens,
and 11 of grok's 113 requests did. pi's
grok-4.6 entry carries no
tiers array, so it billed all 113 at the low
rate. The true figure is $7.76 — pi understates it by
21.5 %. Its gpt-5.6-luna entry does have tiers
and priced correctly, so this is a data gap, not a logic
bug.
Between 97 and 98 % of every run's tokens are cache reads, so the cache-read rate alone sets the ranking. grok sent the fewest tokens of the metered three and still cost the most, because xAI charges 25 % of input for a cache hit where Fireworks charges 3.3 % and OpenAI 10 %.
It resells deepseek-v4-pro-0813 from 14
providers and kimi-k3 from 16, and Fireworks
sits mid-pack in both. DeepSeek's own endpoint halves the
deepseek run to $0.85; Makora cuts kimi 15 % to
$1.77. Pin the provider — default routing balances on
uptime and throughput, and bouncing between providers also
breaks the prompt cache. Rates are pass-through, but card
top-ups carry a 5.5 % fee. For grok it changes nothing: xAI
passes straight through and the only other route is Bedrock
at $2.20/$6.60/$0.55.
Grok came in free with X Premium+ at $40/month, so its price has to be allocated. The defensible split is the standalone rate: SuperGrok sells the same access for $30, leaving ~$10 for the X features. That puts a run at 14¢. An even split would say 9¢ and charging the whole bundle would say 18¢ — the range is narrow and none of it changes a conclusion, but 14¢ is the honest figure, and it is three times what the same run costs on ChatGPT Plus.
OpenAI exposes it:
GET https://chatgpt.com/backend-api/wham/usage
with the Codex OAuth bearer token returns
plan_type plus primary_window (5
h) and secondary_window (7 d), each with
used_percent and reset_at. The
Codex CLI records the same block in every rollout log. xAI
publishes no equivalent route — /v1/usage and
/v1/subscription both 404; the weekly pool is
only visible in grok.com's Usage tab.
fs.promises.glob() returns an async iterator,
not a promise. Three models kept the old
.then(files => …) chain and shipped code
that throws on the first bundle using an
ignore, exclude or
external glob. Two wrapped it in
Array.fromAsync. No test in the repo covers
that file, so all five suites stayed green.
legacy-packages/mendel-requirify still required
rimraf in two test files. Only deepseek and
luna grepped widely enough to find it; both handled the
callback form correctly. grok, qwen and gemma trusted the
issue's list. kimi is the awkward case — it ran a repo-wide
sweep but passed -g '!legacy-packages/**', so
it looked straight past it.
It claims tmp auto-removes its directories on
exit. It does not — tmp 0.2 needs
setGracefulCleanup(). deepseek and luna
dutifully added an exit hook to
validate-manifest.js, which now deletes the
debug manifest one line after printing its path. grok, qwen
and kimi left it alone and matched the old behaviour exactly
— kimi is the only one that wrote down why.
Stopped at five of eight libraries when the GPU was needed
elsewhere — not the model's failure. But its delivered
portion still carried the critical glob bug, never ran
pnpm once, and reached for
--no-verify twice, so the score is not
depressed only by the missing scope.
4.8 M tokens for the whole job — less than half the
next-lowest run, and a third of deepseek's — with a peak
request of 63 k against a 1 M window. It finished in 22
minutes on 158 tool calls, and piped 62 % of its shell
commands through tail or head, the
highest rate in the field. Same eight libraries, a quarter
of the context.
pnpm install && git add <explicit files>
&& git commit, every single time. It is the only run where node_modules
on disk actually matched the branch, and the only one that
finished with nothing to fix.
The only model that wrote the eight top-level items and nothing else, then added each library's sub-items just before starting it. Everyone else planned the whole tree upfront from the issue text — which is also why they inherited its mistakes.
glob shim, the
styleText colour switch and the
temp-dir scoping are all correct.
validate-manifest.js deletes its
own debug artifact.
The added process.once('exit') removes
the temp dir immediately after printing
<path> written, so the file the
message points at is always gone. It also registers
a new exit listener per call.
mendel-mocha-runner/index.js gained
const { globSync } = require('fs') next
to the const fs = require('fs') already
on the line above.
supports-color peer
re-resolution across eslint, debug and axios, folded
into the dependency commits.
fs.promises.glob(...).then is not a
function
at three call sites in
apply-extra-options.js. Any bundle
config using an ignore,
exclude or external glob
throws. Reproduced directly.
validate-manifest.js debug-artifact
deletion as luna.
babel-plugin-polyfill-corejs2 and
re-resolved a babel version, folded into the
dependency commits — 98 of its 195 changed lines
have nothing to do with the issue.
fs.promises.glob bug at
three call sites.
chalk.level = options.enableColor !== false ? 3
: 0
was deleted outright, so enableColor is
now dead config, and without
validateStream: false the printer emits
no colour at all whenever stdout is not a TTY.
Verified: the other three still force colour, qwen
does not.
fs.promises.glob bug at
three call sites.
mendel-mocha-runner/index.js gained
const { globSync } = require('fs') next
to the const fs = require('fs') already
on the line above.
Nothing found. Array.fromAsync glob shim,
correct styleText colour switch with
validateStream: false,
validate-manifest.js left alone, and it is
the only model that both found and fixed the
mendel-requirify references the issue
missed. Process dings live in the matrix:
git add -A on all eighteen commits,
--lockfile-only installs only,
refactor/test commit types.
styleText wrapper honours
enableColor: false, but leaves stream
validation on, so the analytics printer emits no
colour when stdout is not a TTY — master forced
level 3 regardless. Same class of regression as qwen
and haiku, milder shape.
Otherwise clean: trap A passed via
fs.promises.glob handled correctly,
validate-manifest.js left alone, and it
found and fixed the
mendel-requirify references. Ran fourteen
real pnpm installs yet the dropped packages
still resolve from the worktree root — pruning never
verified.
fs.promises.glob bug at
three call sites in
apply-extra-options.js. Reproduced:
TypeError: fs.glob(...).then is not a
function.
styleText runs with stream validation
on, so enableColor no longer forces
colour off a TTY.
Not a bug but the completion gap of the field: haiku
replaced every require() site, then never
touched a single package.json, never ran
pnpm, and left the lockfile untouched. All seventeen
dependency declarations are still there.
shasum(x) with
createHash('sha1').update(x) without
guarding non-string input.
shasum stable-stringifies;
createHash throws. Both call sites
always pass strings today, so it is latent, not
live.
require('crypto') in the end, and all
six hit the
implicit-dependencies ESLint rule on
node:crypto first.
chalk item
found the two call sites the issue did not list —
cache/client.js and
pipeline.js.