Controlled bake-off · irae/mendel · issue #13

Dependency purge as AI model benchmark

The same prompt, the same commit, the same eight npm packages to replace with Node built-ins. Every branch reviewed against the real code, every process claim checked against the session log.

Base tag benchmark-blind-base (per-row base_commit) Branches <model>-issue-13 Libraries 8 Runs 25–26 Aug 2026 Models 6 Harness pi agent

Guided report →  · 

Scoreboard

Ranked by the weighted criteria in the second table. Cost is the metered vendor rate with what was really paid beneath it; local models are free. Peak context is the largest single request, against each model's own window.

     

Prompt v1.1

# Model Score Cost Wall Tokens Peak ctx Window Commits Bugs
1 Kimi K3pi · fireworks · think high
94
$2.13paid $2.13 20 min 4.7 M 67 k 6 % 17 2
2 qwen3.8-27b (bartowski Q4_K_M, xhigh)pi · llama · local · think xhigh
93
2 tooling nudges · 23 tool errors · 3 compactions
$0plan ≈$0.00 213 min 10.1 M 62 k 94 % 17 2
3 Grok 4.6pi · xai · think high
93
11 tool errors
$7.29plan ≈$0.14 21 min 13.2 M 258 k 52 % 18 2
4 GPT-5.6 Solpi · openai-codex · think high
92
13 tool errors
$12.11plan ≈$0.45 24 min 19.6 M 250 k 92 % 18 2
5 Claude Opus 5pi · anthropic · think high
91
10 tool errors
$5.16paid $5.16 21 min 7.3 M 100 k 50 % 17 3
6 Qwen3.8 27B (GGUF Q4_K_M)pi · llama · local · think medium
87
1 tooling nudge · 17 tool errors · 2 failed commits
$0plan ≈$0.00 129 min 5.9 M 46 k 93 % 10 1
7 DeepSeek V4 Flash (0731)pi · fireworks · think high
85
1 model nudge · 1 tool error · 1 failed commit · repetition loop 0.02 on text
$0.55paid $0.79 101 min 17.9 M 703 k 70 % 14 1
8 GPT-5.6 Lunapi · openai-codex · think high
84
peak_context corrected 2026-09-04: the earlier 99425 was the value after the compaction. · 1 model nudge · 1 compaction
$0.56plan ≈$0.03 27 min 16.8 M 256 k 37 % 21 2
9 qwen3.8-27b (ISTA IQ3_S-mtp, xhigh)pi · llama · local · think xhigh
81
15 tool errors · 2 failed commits
$0plan ≈$0.00 109 min 10.8 M 118 k 80 % 17 6
10 DeepSeek V4 Pro (0813)pi · fireworks · think high
79
11 tool errors · repetition loop 0.03 on thinking
$2.10paid $3.33 93 min 19.2 M 479 k 18 % 18 5
11 qwen3.8-27b (ISTA IQ3_S-mtp)pi · llama · local · think medium
77
12 tool errors · 1 failed commit
$0plan ≈$0.00 135 min 7.9 M 89 k 77.94 % 17 8
12 qwen3.8-27b (reserve 8192)pi · llama · local · think medium
76
2 tooling nudges · 14 tool errors · 1 failed commit
$0plan ≈$0.00 98 min 5.0 M 60 k 92 % 12 6
13 GLM 5.3 Flashpi · fireworks · think high
75
18 tool errors
$0.19paid $0.19 21 min 5.1 M 60 k 6 % 9 3
14 qwen3.8-27b (ISTA IQ3_S-mtp, low)pi · llama · local · think lowpartial · 7 of 8
66
7/8 done · turn_timeout
$0plan ≈$0.00 163 min 11.4 M 130 k 88.3 % 15 7
15 Qwen3.6 35B-A3Bpi · llama · local · think high
63
13 tool errors
$0plan ≈$0.00 79 min 7.9 M 94 k 96 % 13 6
16 qwen3.6-35b-a3b (unsloth UD-Q4_K_XL, off)pi · llama · local · think off
51
26 tool errors · 1 failed commit · 2 compactions
$0plan ≈$0.00 40 min 7.0 M 98 k 119.4 % 10 14
17 qwen3.6-35b-a3b-f16 (unsloth UD-Q4_K_XL, no drafter, on)pi · llama · local · think high
50
29 tool errors · 2 failed commits · 2 compactions
$0plan ≈$0.00 33 min 7.3 M 61 k 93.8 % 13 11
18 Gemma 4 26B-A4Bpi · llama · local · think high
48
2 model nudges · 40 tool errors · 6 failed commits
$0plan ≈$0.00 81 min 23.8 M 209 k 98 % 21 3
19 Claude Sonnet 4.5pi · anthropic · think high
44
6/8 done · stopped without finishing
$2.99paid $2.99 31 min 7.7 M 96 k 10 % 13 8
20 Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think highpartial · 3 of 8 · hit the 300 min time budget
38
raw 55 · 3/8 done · hit the 300 min time budget · repetition loop 0.07 on thinking
$0plan ≈$0.00 300 min 3.6 M 52 k 87 % 4 4
21 qwen3.8-27b (AtomicChat AD-IQ3_S)pi · llama · local · think mediumpartial · 3 of 8 · repetition loop, stopped by the runner
38
raw 74 · 3/8 done · repetition loop, stopped by the runner
$0plan ≈$0.00 60 min 7.0 M 70 k 71.24 % 7 5
22 Claude Haiku 4.5pi · anthropic · think high
34
6/8 done · stopped without finishing
$1.05paid $1.05 10 min 7.9 M 101 k 50 % 8 8
23 Qwen3.8 27B (MLX 4-bit, low reasoning)pi · mlx · local · think lowpartial · 1 of 8 · tooling-nudge budget spent
13
raw 67.5 · 1/8 done · tooling-nudge budget spent
$0plan ≈$0.00 85 min 0.6 M 24 k 88 % 1 1
24 Ternary Bonsai 27B (PrismML GGUF)pi · llama · local · think high
13
raw 60.5 · 1/8 done · stopped without finishing
$0plan ≈$0.00 43 min 1.7 M 51 k 77 % 2 7
25 Gemma 4 26B-A4Bpi · llama · local · think offpartial · 1 of 8 · repetition loop, stopped by the runner
13
raw 21 · 1/8 done · repetition loop, stopped by the runner · repetition loop 0.38 on tool call
$0plan ≈$0.00 28 min 8.1 M 136 k 63.76 % 7 17
Gemma 4 12Bpi · lmstudio · local · think highpartial · 0 of 8 · model-nudge budget spent
0
invalid — repetition loop on the LM Studio MLX thinking-on entry google/gemma-4-12b, pre-fix chat template (research run 2); zero commits
$0plan ≈$0.00 50 min 0.2 M 28 k 18 % 0 4
gemma-4-12bpi · llama · local · think offpartial · 0 of 8 · model-nudge budget spent
0
invalid — zero commits: budget spent on a repeated edit-tool schema error (24 of 28 calls), model_budget_exhausted after 3 nudges
$0plan ≈$0.00 80 min 10.0 M 179 k 68.14 % 0 25

Prompt v1.0

# Model Score Cost Wall Tokens Peak ctx Window Commits Bugs
1 Grok 4.6pi · xai · think high
90
16 tool errors · 1 failed commit
$7.76plan ≈$0.14 20 min 11.6 M 260 k 52 % 17 0
2 Claude Opus 5claude-code · anthropic
88
6 tool errors
$8.45plan ≈$0.56 44 min 12.1 M 131 k † 65 % 18 0
3 GPT-5.6 Lunapi · openai-codex · think high
85
22 tool errors · 1 compaction
$0.42plan ≈$0.05 39 min 28.0 M 354 k 130 % 20 2
4 Claude Sonnet 5claude-code · anthropic
83
10 tool errors · 1 failed commit
$10.17plan ≈$0.67 55 min 43.9 M 228 k † 23 % 18 2
5 Kimi K3pi · fireworks · think high
81
5 tool errors · 1 failed commit
$2.08paid $2.08 22 min 4.8 M 63 k 6 % 17 2
6 DeepSeek V4 Pro (0813)pi · fireworks · think high
71
9 tool errors
$0.85paid $1.70 29 min 22.1 M 163 k 16 % 18 6
7 GPT-5.6 Solpi · openai-codex · think high
66
12 tool errors · 3 failed commits · 1 compaction
$6.97plan ≈$0.51 28 min 21.5 M 303 k 111 % 19 5
8 Claude Haiku 4.5claude-code · anthropic
52
2 tool errors · 1 failed commit
$1.24plan ≈$0.08 20 min 9.0 M 116 k † 58 % 9 5
9 Qwen3.6 35B-A3Bpi · llama · local · think xhigh
42
28 tool errors · 5 failed commits · 1 compaction
$0plan ≈$0.00 132 min 10.1 M 94 k 96 % 13 5
10 Gemma 4 26B-A4Bpi · llama · local · think highpartial · stuck
38
20 tool errors · 2 failed commits
$0plan ≈$0.00 104 min 8.2 M 142 k 54 % 9 4
11 Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think highpartial · 3 of 8 · harness crash
38
raw 58 · 3/8 done · harness crash
$0plan ≈$0.00 102 min 0.9 M 28 k 49 % 3 8
12 Qwen3.8 27B (MLX 4-bit)pi · mlx · local · think mediumpartial · 3 of 8 · hit the 300 min time budget
38
raw 80 · 3/8 done · hit the 300 min time budget
$0plan ≈$0.00 254 min 1.8 M 24 k 90 % 6 4

Bugs are weighted: critical 3, medium 2, minor 1. gemma stopped early — the run was cut for hardware, not by the model, so its score covers roughly five of the eight libraries; see the partial-run note. Cost is the metered vendor rate, which is what compares models — only the deepseek run was actually billed per token. grok came free with X Premium+ (≈14¢ a run at standalone SuperGrok value) and luna ran on $20/month ChatGPT Plus (≈5¢). See the cost section.

Peak ctx is the largest context one turn occupied, the response of that turn included, across all compaction cycles. count-tool-calls.mjs reads it from the session log of the scored run. † marks the rows from the retired claude-code harness: that harness writes its log in another shape, which the counter does not read, so those cells keep the older reading, which counts the prompt of the largest turn without the response.

Criteria

Every row was checked against the branch or the session log, not against what the model said it did. Weights in the left column; the best cell in each row is highlighted.

Prompt v1.1

Criterion kimi qwen3.8-27b (bartowski Q4_K_M, xhigh) grok sol opus-5 qwen3.8-gguf ds-flash luna qwen3.8-27b (ISTA IQ3_S-mtp, xhigh) deepseek qwen3.8-27b (ISTA IQ3_S-mtp) qwen3.8-27b (reserve 8192) glm-flash qwen3.8-27b (ISTA IQ3_S-mtp, low) (partial) qwen qwen3.6-35b-a3b (unsloth UD-Q4_K_XL, off) qwen3.6-35b-a3b-f16 (unsloth UD-Q4_K_XL, no drafter, on) gemma sonnet-4.5 bonsai (partial) qwen3.8-27b (AtomicChat AD-IQ3_S) (partial) haiku-4.5 qwen3.8-low (partial) bonsai-prism gemma (partial) gemma-12b (partial) gemma-4-12b (partial)
Correctness
Bugs remaining weighted, 25 pts clean 25 none. trap A passed for real, trap C passed, chalk follows the v1.1 contract 25 exit-hook regression 19 exit-hook regression 19 no defects 25 none found 25 clean 25 exit-hook regression 19 TypeError: fs.promises.glob(...).then left on fs.promises.glob (trap A) 16 TypeError thrown 16 TypeError: glob(...).then left on fs.promises.glob (trap A) 16 TypeError: fs.promises.glob(...).then (trap A) 16 TypeError on .then() (trap A) 16 TypeError: fs.promises.glob(...).then left on fs.promises.glob (trap A); regression exit hook in validate-manifest.js (trap C) 10 TypeError: fs.glob is callback-based, not a Promise (trap A) 16 TypeError: glob(...).then left on fs/promises glob (trap A); root rimraf + tmp never removed; 6 of 8 colour sites in cli-printer silently dead 1 TypeError: fs.promises.glob(...).then at all three sites (trap A); root rimraf + tmp never removed and no install ever ran 1 TypeError: fs.promises.glob(...).then (trap A) 16 TypeError thrown; root deps never removed 1 none 25 no critical; root rimraf left declared, dead base64url padding strip 16 TypeError + root deps never removed 7 none 25 chalk port correct; 7 libs never touched (vacuous) 25 TypeError: fs.promises.glob(...).then left in the tree (trap A) 0 none, no code landed 25 nothing landed; the uncommitted uuid swap never binds uuidv4 0
Unit tests still green per package 230/236 (flake, 2/2 standalone) clean; mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row 260/260 (resolver 75/75 standalone) 674/674 (1 skip, no failures) clean (all suites) clean (tap 71+41+27+42+4+3+15 pass, 1 skip; config 41/41 ×3) 260/260 260/260 (mendel-pipeline lerna-batch flake, clean 5/5 standalone) clean (full suite 260/260); mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row 41/41, 260/260 standalone clean (tap 71+41+27+42+4+3 pass, cli-printer 15/15; full suite 260/260 and karma 12 SUCCESS) clean (tap 71+41+27+42+4+3+15 pass, 1 skip; full suite 260/260) 674/674, 1 skip (no unit test covers trap A's file) clean (full suite 260/260); mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row clean (674/674, 1 skip) clean (full suite 260/260, 22 suites); mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row clean (full suite 260/260, 22 suites); mendel-full-example:test fails on the pre-existing broken bin symlink, same as every other row clean (tap 71+41+27+42+4+3+15 pass, 1 skip) 674/674, 1 skip (full-example daemon-socket fails on master too) not run green where touched (mendel-core 71 + 1 skip, mendel-config 41/41, transform-less, pipeline; full suite 260/260) 674/674, 1 skip (no unit test covers trap A's file) not run mendel-pipeline 15/15; root unit once at end mendel-development tap passes vacuously; nothing covers apply-extra-options.js not run never ran a test; a heredoc overwrote mendel-pipeline/test/helpers/index.js with another package's test
Runtime smoke test ignore/exclude/external globs SYNC OK SYNC OK, pending= 1 SYNC OK SYNC OK SYNC OK, pending= 1 SYNC OK, pending 0 (fs.globSync, handshake kept) SYNC OK SYNC OK THREW: TypeError fs.promises.glob(...).then is not a function THREW: TypeError THREW: TypeError glob(...).then is not a function THREW: TypeError THREW: TypeError glob(...).then is not a function THREW: TypeError fs.promises.glob(...).then is not a function THREW: TypeError THREW: TypeError glob(...).then is not a function THREW: TypeError fs.promises.glob(...).then is not a function THREW: TypeError THREW: TypeError not touched SYNC OK, but vacuous: apply-extra-options.js never opened THREW: TypeError glob(...).then is not a function not touched SYNC OK, file never touched THREW: TypeError glob(...).then is not a function not touched SYNC OK at tip, but tip = base 2652ed6, so the check is vacuous
Task completion
All 8 libraries done 20 pts (shared) 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 7 / 8 8 / 8, root tmp leftover 8 / 8 8 / 8 8 / 8, rimraf left in mendel-requirify, root tmp left 6 / 8, rimraf+tmp incomplete 3 / 8 3 / 8 6 / 8 (glob broken, mendel-requirify untouched) 1 / 8 1 / 8 1 / 8 done (uuid only) 0 / 8 0 / 8 done, branch tip equals the base commit
No code still requires it rimraf ×2 left rimraf ×2 left (mendel-requirify tests) clean clean clean none (fixture left as-is) rimraf ×2 left clean rimraf ×2 left (mendel-requirify tests) rimraf x2 left rimraf ×2 left (mendel-requirify tests) rimraf ×2 left (mendel-requirify tests) clean rimraf ×2 left (mendel-requirify tests); shasum ×2 left (mendel-development, mendel-outlet-manifest) clean no require left in any package rimraf x2 left (mendel-requirify tests) rimraf ×2 left (mendel-requirify tests) rimraf, tmp left at root 19 files left 17 requires left (glob, chalk, tmp, shasum) rimraf ×2 left in mendel-requirify tests 7 files left 22 left 18 stale requires at tip, incl. rimraf and glob it claimed to remove 28 sites left 28 stale requires, all still the base's own
Gone from every package.json rimraf left (mendel-requirify) rimraf left (mendel-requirify) clean clean clean none, root rimraf + tmp removed rimraf + tmp left clean rimraf left (mendel-requirify) rimraf left (mendel-requirify) rimraf left (mendel-requirify) rimraf left (mendel-requirify) clean rimraf left (mendel-requirify); shasum ×2 left tmp left at root root package.json still declares rimraf + tmp, though the closing summary claims all eight removed root rimraf + tmp still declared; mendel-requirify rimraf left rimraf (mendel-requirify) + tmp (root) left rimraf left (mendel-requirify) 12 files left 11 entries left, root rimraf and tmp among them rimraf, tmp still in root package.json 17 files left 17 left 14 entries left, root rimraf + tmp still declared 18 left 18 entries left, root rimraf + tmp still declared
Found the reference the issue missed legacy-packages/mendel-requirify found, not fixed found by its own grep, judged legacy-packages out of scope fixed fixed found + fixed found by grep at line 25, fixed in e7c885e found, not fixed fixed found by its own grep, judged legacy-packages out of scope not found printed in its own grep, judged out of scope found twice, judged out of scope fixed found by its own grep, judged legacy-packages out of scope found + fixed found by its own grep and fixed: rimraf gone from both mendel-requirify tests and its package.json never looked: legacy-packages never opened, rimraf never grepped repo-wide seen in grep (line 139), skipped as "not in the issue" not found missed found by its own grep, both test files and the devDependency fixed not found, not fixed missed not reached never found missed never found
Subtotal 20 pts 16 16 20 20 20 20 13 20 16 15 16 16 20 11 17 16 14 13 14 8 9 8 2.5 2.5 3 0 0
Dependency hygiene
node_modules actually pruned 8 pts real install 8 real pnpm install --no-frozen-lockfile; no stale root declarations 8 real install 8 real install 8 lockfile-only, hand-checked 7 pnpm install --lockfile-only ×14, node_modules never installed or checked 2 real install 8 real install 8 real pnpm install --no-frozen-lockfile; lockfile shrank every commit 8 real install 8 real pnpm install --no-frozen-lockfile before nearly every commit 8 real pnpm install before every commit 8 real install 8 real pnpm install --no-frozen-lockfile; lockfile shrank every commit 8 no install ran 0 no install ever ran; lockfile unchanged 0 no pnpm install, remove or add at any point; every manifest hand-edited 0 never ran pnpm install, lockfile unchanged 0 no install ran 0 not pruned 0 real pnpm install --no-frozen-lockfile before every commit 8 no pnpm install, lockfile unchanged 0 not pruned 0 never ran pnpm install, lockfile unchanged 0 never ran pnpm install, lockfile unchanged 0 never ran pnpm 0 no install, no prune 0
Lockfile pruned lines removed 102 lines 96 lines removed -211/+22 lines -96/+0 lines 96 lines 96 lines removed 93 lines -234/+24 lines 96 lines removed 96 lines 96 lines removed 96 lines removed 104/-103 lines 70 lines removed 0 lines 0 lines removed 0 lines removed 0 lines 175 lines 0 lines 37 lines removed 88/-94 lines 0 lines 0 lines no change 0 lines no change
Static checks
Prettier & ESLint clean 5 pts · re-run on the branch clean · ran 2× 5 clean · self-run prettier and eslint 20 times through the run 5 clean · ran 16× 5 clean · self-run 5 clean, self-run 18x 5 clean on re-run · hooks on 9 commits, self-run once (uuid file) before the one --no-verify 3.5 clean · ran 1× 5 clean · hooks only, not self-run 3 clean · self-run prettier --write and eslint 20 times through the run 5 clean · ran itself 5 clean · self-run prettier --write and eslint on the first package, lint again through pnpm test 5 clean · self-run eslint + prettier --check mid-run and pnpm run static after the last commit 5 TASKS.md dirty · not self-run 1 clean on re-run (only untracked TASKS.md flagged) · self-run prettier --write and eslint 15 times through the run 5 clean · hooks only, not self-run 3 clean on re-run · eslint self-run 11 times, prettier self-run once; the hook did the rest 4 clean on re-run; never ran prettier or eslint directly — the husky hook and the static gate inside pnpm test did it, twice after the last code change 3.5 clean on re-run · hooks only, never self-run 3 clean · hooks only, not self-run 3 ran itself, one self-inflicted break 3 clean · self-run prettier --write and eslint 5 clean · hooks only, not self-run 3 clean · ran itself 5 clean via hooks; eslint only before edits, no prettier 3 prettier fails on cli-printer.js, eslint 5 errors; ran neither tool 0 no diff to check, never ran 0 prettier 2 files, eslint 7 errors; ran neither tool in the session 0
Handled the node: ESLint trap no violation no node: prefix anywhere no violation no violation hit once, recovered pre-commit hit once (chalk commit, hook reject), matched repo style at once no violation no violation no node: prefix anywhere no violation read the patched rule before the first edit; no node: prefix anywhere hit once (chalk commit, hook reject), switched to bare require('util') before it landed no violation no node: prefix anywhere n/a read eslint.config.js first, then kept bare core names — no node: prefix anywhere hit once: the first commit used require('node:crypto') and the pre-commit eslint rejected it; switched to the bare name and stayed bare after hit 3×, dropped the prefix each time 15 git add -A n/a used the node: prefix once, caught it by grep and eslint before the first commit 10× git add -A n/a never hit it (no node: prefixes added) hit 3×, node: prefixes still in the tree at the stop n/a mixed node:crypto and bare crypto in the uncommitted edit
Commit craft
Right type — all chore all chore 17 / 17 chore all chore 18/18 chore 16/17 refactor, not chore 10/10 chore fix/test, not chore 18/21 fix, not chore 2 / 17 chore, 15 refactor refactor, not chore 2 / 17 chore, 15 refactor all 12 refactor, not chore 0/9 chore (refactor/test) 2 / 15 chore, 13 refactor all fix, not chore 2 / 10 chore, 8 refactor 0 / 13 chore, all 13 refactor 17/21 refactor, 4 chore are TASKS.md-only refactor, not chore all fix, not chore 7 / 7 chore 0/8 chore (fix) chore 1/2 chore, repair typed fix 0 / 7 chore, all fix: no commits 0 commits
One commit per package yes 17 / 17 single-package yes yes 17/17 single-package 6/10 single-package; rimraf, glob, tmp, shasum grouped by library (2–4 pkgs each) yes yes 17 / 17 single-package yes 17 / 17 single-package 10 / 12 single-package; tmp and shasum grouped by library yes 15 / 15 single-package 10 / 13 single-package 5 / 10 single-package; the other 5 span 3 to 11 packages 11 / 13 single-package 17/17 single-package; mendel-core 3, mendel-development 3, mendel-deps 2 commits yes 2 of 4 multi-package 7 / 7 single-package never created 1 package 2/2 single-package 6 / 7 single-package; d7f3b03 covers 2 no commits 0 commits
Root devDeps placement rimraf + tmp, unused at root root clean root rimraf + tmp dropped in their own chore commits removed removed removed both removed in-line with their libraries tmp left removed root rimraf + tmp dropped in their own chore commits removed root rimraf + tmp dropped in their own chore commits root rimraf + tmp removed with their libraries removed root rimraf + tmp dropped in their own chore commits tmp not removed rimraf + tmp still declared at root rimraf + tmp still declared at root rimraf removed, tmp left not removed untouched rimraf and tmp still declared at the root rimraf, tmp still declared untouched not reached rimraf + tmp still declared at root untouched rimraf + tmp still declared at root
Commit hooks failures / bypasses 0 hook fails no --no-verify, no bypass, 0 failed commits 0 hook fails 0 hook fails git add -A every commit 2 hook rejects (node: prefix, same commit) · 1× --no-verify 0 hook fails 0 hook fails no --no-verify, no bypass 0 hook fails 1 commitlint reject (non-conventional subject) · no bypass 1 hook reject (node: prefix) · no bypass 0 hook fails no --no-verify, but git add -A used on 2 of 15 commits 1 git add -A no --no-verify, no bypass, no hook failure no --no-verify, no git add -A 3 hook rejects (node: prefix), 3 shell fails; no bypass 0 hook fails 0 hook fails, 3 self-fixed breaks 0 rejects, no bypass 0 hook fails 0 hook fails 0 hook fails; 4 shell/path fails 1 commitlint reject, no bypass 0 hook fails never reached a hook, 0 git commit calls
TASKS.md kept out of git clean never committed clean clean clean kept out (staged once by mistake, unstaged before any commit) clean clean never committed (gitignored) clean never committed (gitignored) never committed clean never committed (excluded) clean never committed; one git add attempt refused by .gitignore never committed; the file is gitignored committed 4× with git add -f past .gitignore clean clean never committed (excluded) clean clean kept out kept out (gitignored) kept out, nothing committed kept out (gitignored)
Subtotal 12 pts 12 12 12 12 6 9 8 8 8 8 6 4 6 5.5 4 6 7 5 2 6 12 1 12 10 6 0 0
Working method
Right the first time 8 pts · self-inflicted repairs 0 nudges 8 0 repair commits, 0 model nudges 8 none 8 0 nudges 8 0 nudges, 0 self-repairs 8 no repair commits; node: trap once, fixed pre-commit; 0 model nudges 7 1 model nudge 6 1 model nudge 6 0 self-inflicted repairs found, 0 model nudges 8 no repairs 8 2 broken package.json edits, 1 unbalanced paren, 1 stray symlink amended out; all self-caught, 0 model nudges 6 no repair commits; node: trap once, fixed pre-commit; 0 model nudges 7 0 nudges 8 tooling nudge after a 10-min stall; run then hit its 25-min turn cap mid-full-suite 7 1 self-repair, caught trap C near-miss itself 6 0 nudges, but 7 edits failed on non-matching or duplicate oldText and were retried 6 0 model nudges, but 29 tool errors, 18 of them failed edits; the node: lint reject and a commitlint reject on the first commit; one shasum commit shipped broken on a partly applied multi-edit and needed an amend 5.5 2 forgotten-package.json repair commits, node: trap 3×, 2 model nudges (−4) 1 no repairs, zero nudges 8 1 JSON syntax break, 4 attempts 3 1 node: prefix and 2 broken package.json edits, all self-caught; 0 model nudges 6 0 nudges 5 0 repairs 8 1 self-repair (bgWhite black crash), 0 nudges 6 node: trap 3×, 1 commitlint reject, 4 commits staged package.json files it never edited 3 0 repairs, 3 nudges (-6) 2 24 of 28 edit calls rejected on the same schema error, never adapted 0
Narrow tests per package as instructed per-package tap runs 21 pnpm --filter runs, one before every commit that had tests per-package tap runs per-package tap runs per-package runs pnpm --filter per package before every commit that had tests per-package tap runs per-package tap runs pnpm --filter per package before every commit that had tests per-package runs pnpm --filter per package before every commit that had tests pnpm --filter per package before every commit that had tests per-package tap runs pnpm --filter per package before every commit that had tests per-package runs 8 pnpm --filter runs, though several commits landed with no narrow run 19 narrow runs (14 tap, 5 pnpm --filter), one after every package commit as the prompt orders it pnpm --filter per package before commits per-package runs none run pnpm --filter per package before every commit that had tests 2 full-suite runs, no visible narrow per-package runs none run package tap 6x per-package tap before some commits; edits left uncommitted none run 0 test runs of any kind
Full suite cadence 10 pts · as the prompt mandates 5 / 17 9 4 full-suite checkpoints across 17 commits 10 3 / 18 9 7 / 18 10 8 full-suite runs / 17 commits 10 full suite after commits 5 and 10, as the prompt says 10 4 / 14 9 6 / 21 9 5 full-suite checkpoints across 17 commits, all green 10 4 / 18 commits 9 full pnpm test after commits 7, 10 and 16, all green 10 full suite after commits 5, 10 and 12 10 6 / 9 10 3 full-suite checkpoints across 15 commits, all green 10 4 full-suite runs / 13 commits 8 3 full-suite checkpoints across 10 commits, all green 9 5 full-suite checkpoints across 13 commits, against the every-5 cadence the prompt asks for 10 1 full suite / 21 commits (prompt: every 5) 4 6 / 13 commits 8 no full suite in 4 commits 3 full pnpm run unit after commit 5, 260/260 10 2 / 8 5 1 commit, no violation 6 commit 1 untested; 1 full suite / 2 commits 6 0 full-suite runs across 7 commits 5 never run 0 0 full-suite runs, 0 commits 0
Task list built progressively 4 pts · high level first, sub-items on arrival full tree upfront, per-package sub-items, ticked 3.5 full tree written upfront from the issue text, ticked by edit after each commit 2.5 full tree upfront, per-package sub-items, ticked 3.5 full tree upfront, per-package sub-items, ticked 2.5 upfront, ticked in 2 bulk batches 1.5 libs + packages from the issue text before grep, file-level sub-items after grep, ticked per library after its commit; rewritten once post-compaction 3 full tree upfront, per-package sub-items, ticked 3.5 full tree upfront, per-package sub-items, ticked 2.5 full tree upfront from the issue, ticked via edit after each commit; several tick edits failed on non-matching oldText and were retried 2 full tree upfront, per-package sub-items, ticked 3 full tree upfront from the issue, ticks batched: 1 after commit 1, 1 at commit 7, then a bulk rewrite at the end 2 one full tree written after the first grep, per-library ticks after each commit, sub-items caught up at the end 3 flat coarse list, no sub-items 1 full tree upfront from the issue, ticked incrementally 2.5 per-file, ticked per commit incl. grep-discovered sites 4 full per-file tree upfront from its own grep, ticked after each commit, partly by rewrite 3.5 full per-file tree upfront from the issue text, nothing added on discovery, but ticked per section right after each commit, never in bulk 2.5 full tree upfront, bulk tick after compaction, boxes drifted 1.5 full tree upfront, ticked despite incomplete work 1.5 upfront full tree, no ticks 1.5 full tree upfront, not one box ticked in 7 commits 0.5 no TASKS.md ever created 0 full tree upfront, live ticks 3 chalk-only list, 8 libs never listed 1 list written, ticked ahead of the work; boxes flipped back and forth 2 full tree upfront, no ticks 1 full per-file tree written upfront, ticked by rewrite, stale after xtend 1
Truncated noisy commands 3 pts · piped to tail/head 63 % 2 56 % 2 3 % 3 21 % 2.5 64 % 3 75 % 2.5 60 % 2 14 % 3 68 % 2.5 49 % 2 81 % 2.5 63 % 2 69 % 2 61 % 3 60 % 2 34 % 2 71 % 2.5 0 / 31 piped 0 41 % 2 13 % 2.5 75 % 2.5 70 % 2 57 % 1 69 % 3 0 of 17 noisy commands piped 0 1 / 1 piped, thin sample 2.5 0 of 9 noisy commands piped; grep -r swept node_modules 0
Followed house conventions 5 pts · minimal diff, matched neighbours tight 5 minimal diff, var/const per file, bare core names, headers kept; a hand-rolled globToFiles recursion instead of Array.fromAsync, plus a three-line explanatory comment the neighbours do not carry 4.5 tight 5 tight 5 minimal diff, matched neighbours 5 minimal diff, var/const per file, bare core names; .sort() added after checking glob order 5 tight 5 tight 5 minimal diff, var/const per file, bare core names, headers kept, no drive-by churn 5 tight 5 minimal diff, var/const per file, bare core names, headers kept, no drive-by churn 5 minimal diff, var/const per file, bare core names, no drive-by churn 5 5 multi-package commits 3 minimal diff, var/const per file, bare core names, headers kept, one added code-comment block in validate-manifest.js 4 wrong fs API choice (callback glob) 3 const in a var file; blank line after the requires dropped in 4 files; unrequested basedir to cwd swap in 2 mendel-deps tests; cli-printer re-flowed around a new applyStyle helper 3 minimal diff, neighbour style matched, chalk port textbook for v1.1; drive-by removal of figures, a library the issue never lists 4 minimal diff; require('fs').promises destructure, uuidv4 alias kept 4 tight but sloppy add -A 4 mostly matched 3 minimal diff, var/const per file, headers kept, no drive-by churn 5 5 multi-package commits 3 tight 5 tight diff, one drive-by comment 4 minimal swaps, but node: prefix against house style, unused globSync and color left 2 no diff 0 minimal swaps attempted, but header stripped and node: prefixes mixed 1
Context economy
Tokens burned all requests, compaction included 4.7 M 10.1 M 13.2 M 19.6 M 7.3 M 5.9 M 17.9 M 16.8 M 10.8 M 19.2 M 7.9 M 5.0 M 5.1 M 11.4 M 7.9 M 7.0 M 7.3 M 23.8 M 7.7 M 3.56 M 7.0 M 7.9 M 610 K 1.7 M 8.1 M 218 K 10.0 M
Compactions context rebuilds mid-run 0 3 0 0 0 4 0 1 0 0 0 1 0 0 0 2 (both overflow) 2 (both overflow) 1 0 0 0 0 0 1 0 0 0
Headroom left in the window 94 % 6 % 48 % 8 % 90 % free (99.6k / 1M) 7 % 30 % 63 % 20 % 82 % left 22 % 8 % 89 % 12 % n/a none — peak 97 823 overshot the 81 920 window (119 %) 94 % — peak 61 485 of 65 536 2 % 96 % left 13 % 29 % 93 % 12 % 23 % free (50.2k / 64k) 36 % 82 % 32 %
Tool calls to finish 159 272 176 189 141 210 170 207 193 292 195 173 169 214 203 190 211 246 151 135 189 134 29 76, 22 errors 120, 21 errors, ended in a 5× identical edit loop 15, then output collapse 92, 30 errors, budget spent on the edit-schema loop
Total 93.5 93 92.5 92 90.5 87 84.5 83.5 80.5 79 76.5 76 75 66 63 50.5 50 47.5 43.5 55 74 34 67.5 60.5 21 30.5 2

Prompt v1.0

Criterion grok opus-5 luna sonnet-5 kimi deepseek sol haiku-4.5 qwen gemma (partial) bonsai (partial) qwen3.8 (partial)
Correctness
Bugs remaining weighted, 25 pts none 25 none 25 1 medium 19 1 medium 19 2 minor 19 1 crit · 1 med · 1 min 7 2 med · 1 min 10 1 crit · 1 med 10 1 crit · 1 med 10 1 crit · 1 min 13 no bugs in 3 landed fixes 25 no bugs in 6 landed commits 25
Unit tests still green per package pass pass pass pass pass pass generator-config fails 3/3 pass pass pass never run ran 13x
Runtime smoke test ignore/exclude/external globs works works works works works TypeError works TypeError TypeError TypeError not reached not reached
Task completion
All 8 libraries done 20 pts (shared) 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 5 / 8 3 / 8 done 3 done, rimraf partial
No code still requires it rimraf ×2 left clean clean clean rimraf ×2 left clean clean rimraf ×2 left rimraf ×2 left 12 sites left 5 left 4 left, rimraf partial
Gone from every package.json 1 left all gone all gone all gone 1 left all gone all gone 17 left — none removed 3 left 8 left 5 left 4 left, rimraf partial
Found the reference the issue missed legacy-packages/mendel-requirify missed found + fixed found found + fixed excluded it told by owner found + fixed missed missed missed not reached not reached
Subtotal 20 pts 15 20 20 20 15 18 20 8 11 7 7.5 8.5
Dependency hygiene
node_modules actually pruned 8 pts real install ×17 8 lockfile-only ×18 2 verified by hand 7 install ×14, not pruned 4 real install ×3 7 lockfile-only ×19 2 lockfile-only ×19 2 never ran pnpm 0 install ×2, unverified 4 never ran pnpm 0 no pnpm install 0 real installs ×3 8
Lockfile pruned lines removed 96 96 96 96 102 lines 194 + unrelated churn 96 − 3 added back 0 9 0 untouched 29 lines removed
Static checks
Prettier & ESLint clean 5 pts · re-run on the branch clean · ran 40× 5 clean · ran 34× 5 clean · ran 40× 5 clean · ran 29× 5 clean · ran 15× 5 clean · ran 42× 5 clean · ran 43× 5 clean via hooks 3 clean via hooks 3 clean, hooks bypassed 2 clean diff, never ran itself 2 clean · ran 7× 5
Handled the node: ESLint trap 1 turn never hit it 1 turn never hit it 1 turn 2 turns never hit it never hit it added it to package.json 3 turns not reached not reached
Commit craft
Right type — all chore all chore refactor/test/chore all fix fix/test/chore refactor/test/chore all chore fix/test/chore all fix mixed refactor/fix mixed fix/refactor used fix, not chore all chore
One commit per package yes yes yes yes yes yes yes per library per library mostly xtend spans 2 pkgs one per package
Root devDeps placement rimraf + tmp, unused at root 2 separate 2 separate 2 separate 2 separate 2 separate 2 separate 2 separate never removed never removed alongside not reached not reached
Commit hooks failures / bypasses 0 hook fails 0 fail 0 fail 1 fail 1 fail 0 fail 3 test-first fails 1 fail 5 fail 2 fail · 2× --no-verify clean, no bypasses 1 no-verify, 1 add -A
TASKS.md kept out of git clean clean clean clean clean clean clean clean committed via git add -A clean kept out kept out
Subtotal 12 pts 12 7.5 9 8.5 9 12 8.5 5.5 3 4 7.5 9
Working method
Right the first time 8 pts · self-inflicted repairs none 8 none 8 2 fix commits 5 none 8 none 7 none 8 1 scope-creep fix 5 none 8 shipped reversed args 2 unfinished 3 no repairs in 3 landed 6 no repairs 7
Narrow tests per package as instructed 14 runs 34 runs 19 runs 34 runs 19 runs 17 runs 24 runs 10 runs 31 runs 8 runs none run 13 runs
Full suite cadence 10 pts · as the prompt mandates 4 / 17 9 4 / 18 10 6 / 20 10 3 / 18 9 5 / 17 10 4 / 18 10 4 / 19 9 3 / 9 9 0 / 13 3 1 / 9 5 never run 1 full suite at 5 and 6 9
Task list built progressively 4 pts · high level first, sub-items on arrival full tree upfront, live ticks 2 own tree upfront, live ticks 2.5 exactly this 4 upfront tree, per-commit ticks 2.5 bulk rewrites ×2 1.5 upfront, but adapted 2 upfront, adapted 2 full tree upfront 1 full tree upfront 1 full tree upfront 1 built upfront, marked done 3 late (39 min), stale boxes 2.5
Truncated noisy commands 3 pts · piped to tail/head 6 % 0.5 43 % 3 8 % 0.5 31 % 2 62 % 3 40 % 2.5 17 % 1 33 % 2 27 % 1.5 0 % 0 no noisy output 2 31 % 2
Followed house conventions 5 pts · minimal diff, matched neighbours tight 5 tight 5 tight 5 tight 5 duplicate require 4 extra churn 4 fixture churn 3 tight 5 dropped an option 3 duplicate require 3 minimal diff 4 minimal diff 4
Context economy
Tokens burned all requests, compaction included 11.6 M 12.1 M 28.0 M 43.9 M 4.8 M 22.1 M 21.5 M 9.0 M 10.1 M 8.2 M 887 K 1.78 M
Compactions context rebuilds mid-run 0 0 1 0 0 0 1 0 1 0 0 0
Headroom left in the window 48 % 35 % overflowed 77 % 94 % 84 % overflowed 42 % 4 % 46 % 51 % 10 %
Tool calls to finish 188 151 265 317 158 270 194 126 258 for less work 115 unfinished 53, crashed twice 135, mid-run restart
Total 89.5 88 84.5 83 80.5 70.5 65.5 51.5 41.5 38 58 80

Fable re-scored every run from the pushed branches and the versioned session logs (2026-08-31). Four runs received mid-run human help, against the no-help rule: gpt-5.6-luna (a "continue" after a 9.4-minute stall), deepseek-v4-pro (two unsolicited scope instructions covering the trap-B area), qwen3.6-35b-a3b (three interventions, including a nudge after a 35-minute stall), and gemma-4-26b-a4b (an owner instruction inside the scored region). Their scores stand but read them with that asterisk. gpt-5.6-sol carries a newly found deterministic test regression (mendel-config generator-config). gemma-4-26b-a4b's final in-session commit (60b93f8) is absent from the pushed branch.

Footnote: gemma-4-26b-a4b made one in-session commit (60b93f8) that is missing from its branch.

Harnesses

Two harnesses appear in this table and they are not run the same way. pi runs (from 2026-09-01) go through benchmark/run-pi-rpc.mjs: one pi --mode rpc session for the whole run, auto-compaction and auto-retry forced on, and a fixed nudge policy that never reads the chat — a tooling nudge (stream error, premature length stop, stall, dead process) is free; a model nudge (the model stopped with TASKS.md unchecked or a dirty tree) is the same sentence every time and costs 2 points on "right the first time". The runner refuses to start on a model without a truthful context window and output budget. pi runs before that date used pi -p, which exits on the first length or error stop; those were relaunched fresh in the same worktree and are marked partial where that cut them short. Claude Code runs use claude -p --output-format json with permissions skipped: Claude's own compaction and retry, no nudges of either kind (a max-output stop ends the run with no second chance), so its nudge counts read n/a and its runs are, if anything, held to a stricter stop rule. Runner metadata (runs/<slug>-meta.json) records every nudge with its cause. From the same date both harnesses also run in a pinned environment: pi starts with --no-extensions --no-skills --no-prompt-templates and a benchmark-owned config directory, Claude Code starts with --bare and a benchmark-owned config directory, and both read the same frozen global instructions (benchmark/agents-global.md v1.0) and nothing else of the operator's setup. The thinking level of each run is pinned on the command line and shown under the model name; earlier pi runs inherited the operator's default level, which the sub-line now reports from the session logs. Claude Code is retired as a harness (2026-09-01): no new claude-code rows are made; the existing ones stay until replaced. A partial row's sub-line states what ended the run — nudge budget, harness budget, harness crash, time budget, or a stuck loop the operator closed — and how many of the eight libraries were done. The wall-clock and stall budgets are absolute minutes, which favours fast serving stacks; each run's measured output speed is in its meta file.

Cost

Three questions, three answers. Vendor rate is what the run costs anyone paying per token at the provider actually used. Cheapest route is the same model on OpenRouter's least-expensive endpoint. Actually paid is what left the wallet — for two of these runs, a flat subscription that was already sunk.

Per run Cheapest OpenRouter Actually paid Fresh input Output Cache read Vendor rate $/M input $/M output $/M cache read Cache share Provider used
Claude Sonnet 5 (Anthropic) $10.17 (plan est.) $0.67 0.6 k 65 k 43.6 M $10.17 2.00 10.00 0.20 99.9 % Claude plan
Claude Opus 5 (Anthropic) $8.45 (plan est.) $0.56 0.3 k 51 k 12.0 M $8.45 5.00 25.00 0.50 99.6 % Claude plan
Grok 4.6 (xAI only) $7.76 (plan share) $0.14 266 k 33 k 11.3 M $7.76 2.00 → 4.00 6.00 → 12.00 0.50 → 1.00 97.4 % xAI
Grok 4.6 (xAI only) $7.29 (plan share) $0.14 353 k 33 k 12.8 M $7.29 2.00 6.00 0.50 97.1 % xAI
GPT-5.6 Sol (Flex est.) $6.97 (plan share) $0.51 434 k 26 k 21.1 M $13.94 5.00 → 10.00 30.00 → 45.00 0.50 → 1.00 97.9 % ChatGPT plan
Claude Opus 5 (Anthropic) $5.16 (metered) $5.16 0.3 k 39 k 7.1 M $5.16 5 25 0.5 99.5 % anthropic
Kimi K3 (Fireworks) $2.13 (metered) $2.13 133 k 25 k 4.5 M $2.13 3 15 0.3 96.6 % Fireworks
DeepSeek V4 Pro (0813) (DeepSeek V4 Pro) $2.10 (metered) $3.33 344 k 522 k 18.3 M $3.33 1.32 3.96 0.044 95.5 % Fireworks
Kimi K3 (Fireworks) $2.08 (metered) $2.08 63 k 32 k 4.7 M $2.08 3 15 0.3 98.0 % Fireworks
Claude Haiku 4.5 (Anthropic) $1.24 (plan est.) $0.08 1 k 31 k 8.9 M $1.24 1.00 5.00 0.10 99.6 % Claude plan
DeepSeek V4 Pro (0813) (DeepSeek) $0.85 (metered) $1.70 316 k 82 k 21.7 M $1.70 1.32 3.96 0.044 98.2 % Fireworks
DeepSeek V4 Flash (0731) (Makora) $0.55 (metered) $0.79 1131 k 651 k 16.1 M $0.79 0.14 0.28 0.028 90.0 % Fireworks
GPT-5.6 Luna (Flex) $0.42 (plan share) $0.05 487 k 44 k 27.5 M $0.84 0.20 → 0.40 1.20 → 1.80 0.02 → 0.04 98.1 % ChatGPT plan
GLM 5.3 Flash (Fireworks) $0.19 (metered) $0.19 235 k 28 k 4.8 M $0.19 0.15 0.5 0.029 94.8 % Fireworks

Plan estimated costs

The per-run table above is real money: the metered vendor rate and the cheapest OpenRouter route. This section is the separate plan view — amortising each flat subscription against the share of the allowance the run consumed. Four runs were subscription-billed; the two Claude Code runs report a metered figure from the harness itself and their plan-window share is not exposed. The Claude figures are estimates: Claude Code does not log window usage, so the marginal cost assumes a Max 5x week sustains roughly $350 of metered-equivalent work — marginal ≈ metered ÷ 350 × $23.08. Recheck against claude.ai usage when precision matters. Amortising each plan against the share of the weekly allowance the run consumed gives the real marginal cost, and the runs-per-week the plan sustains.

On a plan Marginal cost Metered equivalent Runs/week held Plan Allocated to the model A week of allowance Of the 5-hour window Of the weekly window
Grok 4.6 ≈ $0.14 $7.29 ≈ 50 X Premium+ $40/mo SuperGrok rate $30 $6.92 not exposed ≈ 2 %
GPT-5.6 Sol ≈ $0.45 $12.11 ≈ 9 ChatGPT Plus $20/mo all of it $20 $4.62 not measured ≈ 10 %
Grok 4.6 ≈ $0.14 $7.76 ≈ 50 X Premium+ $40/mo SuperGrok rate $30 $6.92 not exposed 2 %
Claude Opus 5 (est) ≈ $0.56 $8.45 ≈ 41 Claude Max 5x $100/mo all of it $100 $23.08 not exposed not exposed
GPT-5.6 Luna ≈ $0.05 $0.84 ≈ 100 ChatGPT Plus $20/mo all of it $20 $4.62 8 % 1 %
GPT-5.6 Luna ≈ $0.03 $0.47 ≈ 100 ChatGPT Plus $20/mo all of it $20 $4.62 not measured ≈ 1 %
Claude Sonnet 5 (est) ≈ $0.67 $10.17 ≈ 34 Claude Max 5x $100/mo all of it $100 $23.08 not exposed not exposed
GPT-5.6 Sol ≈ $0.51 $13.94 ≈ 9 ChatGPT Plus $20/mo all of it $20 $4.62 74 % ≈ 11 %
Claude Haiku 4.5 (est) ≈ $0.08 $1.24 ≈ 280 Claude Max 5x $100/mo all of it $100 $23.08 not exposed not exposed

How pi prices a run

A local static table, not a live lookup and not OpenRouter. calculateCost() multiplies the token counts the provider returns by the per-model rates in ~/.pi/agent/models-store.json, picking the highest matching entry in an optional cost.tiers array. It matched Fireworks to the cent on both deepseek and kimi, and matched OpenAI on luna.

Where pi is wrong about grok

xAI doubles every rate once a prompt passes 200 k tokens, and 11 of grok's 113 requests did. pi's grok-4.6 entry carries no tiers array, so it billed all 113 at the low rate. The true figure is $7.76 — pi understates it by 21.5 %. Its gpt-5.6-luna entry does have tiers and priced correctly, so this is a data gap, not a logic bug.

Cache reads decide everything

Between 97 and 98 % of every run's tokens are cache reads, so the cache-read rate alone sets the ranking. grok sent the fewest tokens of the metered three and still cost the most, because xAI charges 25 % of input for a cache hit where Fireworks charges 3.3 % and OpenAI 10 %.

OpenRouter is worth it for two of the four

It resells deepseek-v4-pro-0813 from 14 providers and kimi-k3 from 16, and Fireworks sits mid-pack in both. DeepSeek's own endpoint halves the deepseek run to $0.85; Makora cuts kimi 15 % to $1.77. Pin the provider — default routing balances on uptime and throughput, and bouncing between providers also breaks the prompt cache. Rates are pass-through, but card top-ups carry a 5.5 % fee. For grok it changes nothing: xAI passes straight through and the only other route is Bedrock at $2.20/$6.60/$0.55.

Splitting a bundled plan

Grok came in free with X Premium+ at $40/month, so its price has to be allocated. The defensible split is the standalone rate: SuperGrok sells the same access for $30, leaving ~$10 for the X features. That puts a run at 14¢. An even split would say 9¢ and charging the whole bundle would say 18¢ — the range is narrow and none of it changes a conclusion, but 14¢ is the honest figure, and it is three times what the same run costs on ChatGPT Plus.

Reading your own plan usage

OpenAI exposes it: GET https://chatgpt.com/backend-api/wham/usage with the Codex OAuth bearer token returns plan_type plus primary_window (5 h) and secondary_window (7 d), each with used_percent and reset_at. The Codex CLI records the same block in every rollout log. xAI publishes no equivalent route — /v1/usage and /v1/subscription both 404; the weekly pool is only visible in grok.com's Usage tab.

What decided it

The trap that split the field

fs.promises.glob() returns an async iterator, not a promise. Three models kept the old .then(files => …) chain and shipped code that throws on the first bundle using an ignore, exclude or external glob. Two wrapped it in Array.fromAsync. No test in the repo covers that file, so all five suites stayed green.

The reference the issue missed

legacy-packages/mendel-requirify still required rimraf in two test files. Only deepseek and luna grepped widely enough to find it; both handled the callback form correctly. grok, qwen and gemma trusted the issue's list. kimi is the awkward case — it ran a repo-wide sweep but passed -g '!legacy-packages/**', so it looked straight past it.

Where the issue was wrong

It claims tmp auto-removes its directories on exit. It does not — tmp 0.2 needs setGracefulCleanup(). deepseek and luna dutifully added an exit hook to validate-manifest.js, which now deletes the debug manifest one line after printing its path. grok, qwen and kimi left it alone and matched the old behaviour exactly — kimi is the only one that wrote down why.

gemma's partial run

Stopped at five of eight libraries when the GPU was needed elsewhere — not the model's failure. But its delivered portion still carried the critical glob bug, never ran pnpm once, and reached for --no-verify twice, so the score is not depressed only by the missing scope.

kimi's context discipline

4.8 M tokens for the whole job — less than half the next-lowest run, and a third of deepseek's — with a peak request of 63 k against a 1 M window. It finished in 22 minutes on 158 tool calls, and piped 62 % of its shell commands through tail or head, the highest rate in the field. Same eight libraries, a quarter of the context.

grok's quiet advantage

pnpm install && git add <explicit files> && git commit, every single time. It is the only run where node_modules on disk actually matched the branch, and the only one that finished with nothing to fix.

luna's task-list discipline

The only model that wrote the eight top-level items and nothing else, then added each library's sub-items just before starting it. Everyone else planned the whole tree upfront from the issue text — which is also why they inherited its mistakes.

Defect ledger — 16 findings, open when you want them

grok-4.6 — 0 points

  • Nothing found. Every swap is behaviour-preserving; the glob shim, the styleText colour switch and the temp-dir scoping are all correct.

gpt-5.6-luna — 2 points

  • mediumvalidate-manifest.js deletes its own debug artifact. The added process.once('exit') removes the temp dir immediately after printing <path> written, so the file the message points at is always gone. It also registers a new exit listener per call.

kimi-k3 — 2 points

  • minorDuplicate import. mendel-mocha-runner/index.js gained const { globSync } = require('fs') next to the const fs = require('fs') already on the line above.
  • minorUnrelated lockfile churn. 21 of its inserted lines are supports-color peer re-resolution across eslint, debug and axios, folded into the dependency commits.

deepseek-v4-pro-0813 — 6 points

  • criticalfs.promises.glob(...).then is not a function at three call sites in apply-extra-options.js. Any bundle config using an ignore, exclude or external glob throws. Reproduced directly.
  • mediumSame validate-manifest.js debug-artifact deletion as luna.
  • minorUnrelated lockfile churn. Dropped babel-plugin-polyfill-corejs2 and re-resolved a babel version, folded into the dependency commits — 98 of its 195 changed lines have nothing to do with the issue.

qwen3.6-35b-a3b — 5 points

  • criticalSame fs.promises.glob bug at three call sites.
  • mediumCLI analytics lost its colours. chalk.level = options.enableColor !== false ? 3 : 0 was deleted outright, so enableColor is now dead config, and without validateStream: false the printer emits no colour at all whenever stdout is not a TTY. Verified: the other three still force colour, qwen does not.

gemma-4-26b-a4b — 4 points, over five of eight libraries

  • criticalSame fs.promises.glob bug at three call sites.
  • minorDuplicate import. mendel-mocha-runner/index.js gained const { globSync } = require('fs') next to the const fs = require('fs') already on the line above.

claude-opus-5 — 0 points

Nothing found. Array.fromAsync glob shim, correct styleText colour switch with validateStream: false, validate-manifest.js left alone, and it is the only model that both found and fixed the mendel-requirify references the issue missed. Process dings live in the matrix: git add -A on all eighteen commits, --lockfile-only installs only, refactor/test commit types.

claude-sonnet-5 — 2 points

  • mediumForced colour lost. The styleText wrapper honours enableColor: false, but leaves stream validation on, so the analytics printer emits no colour when stdout is not a TTY — master forced level 3 regardless. Same class of regression as qwen and haiku, milder shape.

Otherwise clean: trap A passed via fs.promises.glob handled correctly, validate-manifest.js left alone, and it found and fixed the mendel-requirify references. Ran fourteen real pnpm installs yet the dropped packages still resolve from the worktree root — pruning never verified.

claude-haiku-4.5 — 5 points

  • criticalSame fs.promises.glob bug at three call sites in apply-extra-options.js. Reproduced: TypeError: fs.glob(...).then is not a function.
  • mediumCLI analytics lost its colours the same way qwen did — the forced level is gone and styleText runs with stream validation on, so enableColor no longer forces colour off a TTY.

Not a bug but the completion gap of the field: haiku replaced every require() site, then never touched a single package.json, never ran pnpm, and left the lockfile untouched. All seventeen dependency declarations are still there.

Shared, not charged to anyone

  • All six replaced shasum(x) with createHash('sha1').update(x) without guarding non-string input. shasum stable-stringifies; createHash throws. Both call sites always pass strings today, so it is latent, not live.
  • All six correctly used bare require('crypto') in the end, and all six hit the implicit-dependencies ESLint rule on node:crypto first.
  • All five that reached the chalk item found the two call sites the issue did not list — cache/client.js and pipeline.js.