Skip to content

EvalPlus (HumanEval+) — M1 Max 32 GB

The quality gate: pass@1 at temperature 0, output budget calibrated per model — see the method. Scores are shared across serving configs when thinking mode, effort, and quant match.

These runs are shallow, a few thousand tokens each, so the KV cache type does not move a score and the rows name it only where it is part of the quant, as the fork's calibrated q4_0 KV is. Each model page names the KV type its config serves.

configbudgetpass@1 basepass@1 plusemptycompletion
Qwen3.8-27B 4-bitmlx-community,mlx_lm.serverf16medium81920.9820.9390/164100%
Qwen3.8-27B IQ3_S-mtpISTA-DASLab,llama-servermtp/3f16medium81920.9760.9451/16499%
Qwen3.8-27B AD-IQ3_SAtomicChat,llama-servermtp/3f16medium88860.9880.9270/164100%
Qwen3.8-27B IQ3_S-mtpISTA-DASLab,llama-serverf16low81920.9760.9331/16499%
Qwen3.8-27B IQ3_S-mtpISTA-DASLab,llama-serverf16xhigh300000.9450.9215/16497%
Qwen3.6-35B-A3B UD-Q4_K_XLunsloth,llama-servermtp/3q8_0on266240.9390.9215/16497%
Qwen3.6-35B-A3B UD-Q4_K_XLunsloth,llama-servermtp/3q8_0off81920.9510.9150/164100%
Ternary-Bonsai-27B Q2_g64prism-ml,prism-llamaq4_0+biason102400.9270.8904/16498%
Ternary-Bonsai-27B 2-bitprism-ml,mlx_lm.serverf16on102400.9150.8845/16497%
Gemma-4-12B Q4_K_XLunsloth,llama-serverf16off81920.9760.9390/164100%
Gemma-4-12B 4-bitlmstudio-community,lmsf16off300000.9090.8720/164100%
Gemma-4-26B-A4B UD-Q4_K_XLunsloth,llama-servermtp/2f16off81920.9760.9450/164100%
Gemma-4-26B-A4B UD-Q4_K_XLunsloth,llama-servermtp/2f16on300000.8840.86018/16489%
Gemma-4-26B-A4B 4-bitmlx-community,mlx_lm.serverf16on300000.7130.70146/16472%
Qwen3.8-27B Q4_K_Mbartowski,llama-servermtp/3f16xhigh300000.9570.9396/16496%

An empty completion is reasoning that exhausted the output budget; a high empty rate is a real model limit, not a harness bug. Completion is the share of the 164 problems that got an answer at all. Read the two columns together: a low score with a low completion rate measures delivery, not code quality — the withdrawn Gemma-4-12B thinking-on row passed 99% of the answers it gave and never answered 37% of the problems (historical). Raw runs and per-problem results live in the repo's benchmarks/ run kits; each model's data page (first column) carries its scoring notes.