Research report in English · KyaniteLabs Research. First published September 28, 2026. Original public report. This edition consolidates the report and preserves its original limits.
The useful question is what this image stack actually did on this machine. The September 28 report describes an approximately $1,000 consumer APU, an 8-core Strix Halo, 80 GiB of GTT shared memory and a 7B diffusion transformer in Q4_K.
Measured timing and memory
- Co-residency: 5.84–5.88 seconds per iteration across the measured memory modes, with both resident LLM services live. Peak GTT was 58.8 GiB out of 80, with approximately 15.5 GiB RAM free.
- Sampler cost: at 30 steps, euler and euler-a were speed-identical. dpm++2m added 5.6% at guidance 4 and 15.9% at guidance 6. At 6 and 10 steps, the report found no speed separation.
- Guidance timing: 191.7 seconds at guidance 4 versus 211.4 at guidance 6 per 768×512 image, with n=10+10. That is a timing comparison; paired quality assessment remained pending.
- Reload overhead: approximately 15–20 seconds per fresh CLI invocation. A 20-image batch spent about 6 minutes of roughly 65 minutes reloading. The suggested approximately 9% throughput improvement was an estimate, not a measured adoption result.
- Full-quality run: roughly 3.2 minutes per 768×512 image at about 54 W GPU-package power. The reported CPU-only comparison was roughly 29 minutes per image and failed under sustained load.
Quality was a human judgment
One internal human judge preferred dpm++2m at 30 steps and found it usable at 6 steps. The same judge reported successful stitched lettering. Those are labeled single-judge observations. An empty OCR result on those images was reported as an instrument artifact; this page does not extend that finding to other images.
Limits and unresolved work
One machine class, one model family, no blind scoring and no inter-rater statistics. The source still listed paired guidance-quality judgments and like-for-like comparison cells as unresolved. These historical findings are not a current model ranking, a hardware guarantee or permission to use particular weights commercially.
The transferable method is a frozen run plan, paired seeds, append-only indexes, per-image sidecars and exact output hashes. The numbers stay attached to the conditions that produced them.