Local LLM / Benchmarks

Lab Notes: the KV verdict

The q8-to-q4 KV question from night 1 got its paired answer: same accuracy, zero tripwires, half the cache. Plus the label we had to correct in public when the server's counter beat our estimate.

By Simon Gonzalez de Cruz (follow the build in public on X @KyaniteLabs_), assisted by GLM-5.3. Night 2 of the nightly lab notes. Last night ended with a promise: run the paired q4_0-KV A/B our configuration was missing. This is what came back.

The number first: the champion now runs 4-bit KV cache — about 47% less KV memory (4.3GB saved) at the same context window — with zero measured accuracy cost and zero corruption flags. The promotion gate was paired, pre-registered, and self-executing: if the q4 arm matched the q8 arm's needle hits and posted zero tripwires, the unit file flips; anything less, the old champion restores itself.

The evidence stack

Three layers, each paired:

LayerDesignResult
Task accuracyGSM8K, n=60 stratified, pairedExact McNemar p = 1.0; q4 slightly faster (18.78 s vs 19.47 s mean)
Deep-context integrityNeedle probes at 50%/90% depth, identical ~198k-token haystack, same seed, same-session controlq4: 0/2 hits, zero tripwires, clean stops; q8: 0/2 hits, one trip
Retrieval boundaryNeedle at 25% depth (where q8 historically retrieves)FOUND, clean stop — the dumb-zone boundary did not move

The tripwire is the finish_reason anomaly counter we keep armed on every serving hour since the RDNA4 quantized-KV corruption report made the rounds. One trip fired all night — on the q8 arm (a 40-token answer that never stopped; n=1, plausibly ordinary deep-context rambling, recorded not interpreted). The arm the rule actually gates on — q4 — ran clean.

The miss we caught in ourselves

The gate's haystack was built to an estimate: 5,800 paragraphs, ~45 tokens each, call it 262k. The server's own counter said 198,227 tokens. The paired comparison is untouched — both arms prefilled the identical haystack, which is the entire point of pairing — but every place we had written “262k needle probes” was overstating the instrument. The announcement tweet had already gone out with the estimate. The correction went out as a reply within the hour, the repo README now carries server-reported token counts, and the rule is restated for good: label instruments with the counter at the source, not the generator's estimate. Cold prefill measured through that needle wall: 109.5 tok/s at ~198k (needle-class number: full prompt + answer, not a bench harness).

What half a cache buys

Arithmetic, labeled as arithmetic: q8_0 encodes ~8.5 bits per KV value, q4_0 ~4.5 — hence ~47% (4.3GB saved) off the KV line at identical context. Measured wall time through the probes trended the same way as GSM8K: q4 marginally faster on cache reads. The compounding cases are next: the 35B dress-rehearsal math (where KV headroom decides feasibility), and the adaptive-KV research lane, whose bar just moved from “beat static q8” to “beat static q4.”

Conditions

GMKtec EVO-X2 (Ryzen AI Max+ 395, 96GB unified), Qwen3.8-27B Q4 dynamic quant, llama.cpp ROCm build, 262,144-token window, K+V q4_0 KV (champion as of 2026-08-18 10:32Z), temperature 0. GSM8K n=60 paired; needle n=2 stations per arm plus one boundary probe; one box, one night; tripwire counter armed continuously. Datapoints, not laws. Raw gate log and rollback unit referenced in the repo: qwen38-27b-strix-halo.

Work with Kyanite

Want this working in your environment?

If this post describes a Kyanite tool or result you need, implementation help can cover setup, advising, docs, examples, checks, and a usable handoff.

Fit boundary

Kyanite offers help grounded in its tools, products, and build practice. Broader consulting routes through PuenteWorks.

Keep following the system.

The Delegation Card: we asked a $1,400 mini-PC to take our jobs

Not is-it-smart but can-you-hand-it-work-and-walk-away. 495 trials, certified floors.

Qwen 3.8 27B on Strix Halo: the complete measured story

Every dial measured, every number public: the frozen optimal config for a 27B on a $1,400 mini-PC.

Lab Notes: the measured-knees method for reasoning effort

A methods note on reasoning-effort calibration. Thinking rescued 15/40 hards vs 4/40 off. On easy tasks it bought nothing. Publish the knee.

Lab Notes: 67% LiveCodeBench-30 on a $1,400 rig

20/30 = 67% LiveCodeBench-30 on a $1,400 mini-PC. Wilson 95% CI 49-81%. Easy 10/10, medium 8/10, hard 2/10. n=30 public subset. Not the card.

Lab Notes: we reverted a llama.cpp regression

We found a llama.cpp regression and reverted it. n=6 battery on a $1,400 rig: 0/6 slashes before, 5/6 after. Same Q4_K_XL. Not a quant story.

Lab Notes: 93% HumanEval on a $1,400 rig

28/30 = 93% HumanEval on a $1,400 mini-PC. Qwen3.8-27B Q4_K_XL. Temp 0, thinking off. Failures: 50 and 145. Raw log linked.

Lab Notes: the model that can't forget but can't remember

The number first: this model is 75% not a transformer. 48 of 64 layers keep a running state. On a $1,400 rig, load 198k once (1818s), then query in 9-27s.

Lab Notes: the basin was a bug

We published a basin. The serving binary was the hole. After the c7d8722 revert: 6/6 HIT at 198k. n=1 map of the fixed build.

Lab Notes: verdict night

A $1,400 mini-PC serving a 27B at 262k context asks one question: does capping thinking cost accuracy? The paired answer, the invalid first attempt it survived, and the doctrine that followed.

One mini-PC, one night, and the numbers that argued with themselves: tuning Qwen3.8-27B on Strix Halo

A fully measured night-and-evening of tuning a 27B dense model on a Strix Halo mini-PC: the honest bands, the reversal, the crash doctrine, and the EC detective story.

GPT-5.6 Sol vs. Terra vs. Luna: an evidence-based routing policy for coding agents

The practical GPT-5.6 split is Sol for discovery, Terra for bounded execution, and Luna for repeatable processing - with a different verification contract for each.

Agents need verifiable tools, not better prompt theater

The useful agent pattern is not a prettier prompt. It is a tool surface the agent can call, inspect, verify, and revise.

Repo history is a product signal

A repo is not just storage. It is evidence of decisions, repairs, release behavior, naming drift, test gaps, and what the builder actually knows how to finish.

Implementation help is part of the product surface

A useful open-source tool still needs a path from public repo to working environment. That path is product work, not an afterthought.

Why Kinocut matters

Kinocut gives AI agents callable handles on timelines, effects, Hyperframes, and finished media at kinocut.dev.

Infinite monkeys, LLMs, and the room around the machine

The argument behind the video: output quality is not just probability. It is architecture, filters, and human taste.

What a working AI tool needs before people can use it

A practical checklist for turning a working tool, workflow, or rough app into something other people can understand, install, and use.

MCP server implementation checklist

The checklist Kyanite uses to decide whether an MCP server is a toy, a usable tool, or something worth implementing.

Repo archaeology turns history into proof

Why commit history is one of the strongest proof sources for learning diagnostics, implementation help, and engineering trust.

AI discovery needs more than a sitemap

What Kyanite adds so search engines and AI assistants can understand the tools, products, proof, and support path.