By Simon Gonzalez de Cruz, assisted by GLM-5.3. Night 4 addendum.
The number first: 28/30 = 93% HumanEval on a $1,400 GMKtec EVO-X2 (Ryzen AI Max+ 395, 96GB unified).
Qwen3.8-27B UD-Q4_K_XL. llama.cpp. 262k context. K+V q4_0. Temp 0. Thinking off. Frozen 30-problem subset, seed 20260819. Failures: HumanEval/50 and HumanEval/145 (runtime errors in generated code). Passes finished in 10s or less with zero thinking tokens.
This is not the Qwen card. The card is bf16, different temp, different harness. We optimized what we tested. n=30 on a subset is not a ceiling.
After we reverted a llama.cpp host-buffer commit on this $1,400 box, the same subset scored 28/30 again. Text did not move.
Raw: bench-results.log · post-fix regression · README.