EvalPlus (HumanEval+) — M1 Max 32 GB
The quality gate: pass@1 at temperature 0, output budget calibrated per model — see the method. Scores are shared across serving configs when thinking mode, effort, and quant match.
These runs are shallow, a few thousand tokens each, so the KV cache type does not move a score and the rows name it only where it is part of the quant, as the fork's calibrated q4_0 KV is. Each model page names the KV type its config serves.
An empty completion is reasoning that exhausted the output budget; a high empty rate is a real model limit, not a harness bug. Completion is the share of the 164 problems that got an answer at all. Read the two columns together: a low score with a low completion rate measures delivery, not code quality — the withdrawn Gemma-4-12B thinking-on row passed 99% of the answers it gave and never answered 37% of the problems (historical). Raw runs and per-problem results live in the repo's benchmarks/ run kits; each model's data page (first column) carries its scoring notes.