Measured Research / Local Image Generation

Local image generation, measured on a consumer APU

Measured sampler costs, guidance timing and model reload overhead. Quality judgments remain single-judge observations, not a leaderboard.

Research report in English · KyaniteLabs Research. First published September 28, 2026. Original public report. This edition consolidates the report and preserves its original limits.

The useful question is what this image stack actually did on this machine. The September 28 report describes an approximately $1,000 consumer APU, an 8-core Strix Halo, 80 GiB of GTT shared memory and a 7B diffusion transformer in Q4_K.

Measured timing and memory

  • Co-residency: 5.84–5.88 seconds per iteration across the measured memory modes, with both resident LLM services live. Peak GTT was 58.8 GiB out of 80, with approximately 15.5 GiB RAM free.
  • Sampler cost: at 30 steps, euler and euler-a were speed-identical. dpm++2m added 5.6% at guidance 4 and 15.9% at guidance 6. At 6 and 10 steps, the report found no speed separation.
  • Guidance timing: 191.7 seconds at guidance 4 versus 211.4 at guidance 6 per 768×512 image, with n=10+10. That is a timing comparison; paired quality assessment remained pending.
  • Reload overhead: approximately 15–20 seconds per fresh CLI invocation. A 20-image batch spent about 6 minutes of roughly 65 minutes reloading. The suggested approximately 9% throughput improvement was an estimate, not a measured adoption result.
  • Full-quality run: roughly 3.2 minutes per 768×512 image at about 54 W GPU-package power. The reported CPU-only comparison was roughly 29 minutes per image and failed under sustained load.

Quality was a human judgment

One internal human judge preferred dpm++2m at 30 steps and found it usable at 6 steps. The same judge reported successful stitched lettering. Those are labeled single-judge observations. An empty OCR result on those images was reported as an instrument artifact; this page does not extend that finding to other images.

Limits and unresolved work

One machine class, one model family, no blind scoring and no inter-rater statistics. The source still listed paired guidance-quality judgments and like-for-like comparison cells as unresolved. These historical findings are not a current model ranking, a hardware guarantee or permission to use particular weights commercially.

The transferable method is a frozen run plan, paired seeds, append-only indexes, per-image sidecars and exact output hashes. The numbers stay attached to the conditions that produced them.

Trabajar con Kyanite

¿Quieres que esto funcione en tu entorno?

Si esta nota describe una herramienta o resultado de Kyanite que necesitas, la ayuda de implementación cubre setup, asesoría, docs, ejemplos, checks y un handoff usable.

Límite de fit

Kyanite offers help grounded in its tools, products, and build practice. La consultoria mas amplia se enruta por PuenteWorks.

Sigue el sistema.

Play-1: the acceptance-inversion result

The measured nmax2 acceptance inversion: 72.2% versus 22.7% for the draft-12 control. Throughput and engine comparisons, with single-lab limits.

Adversarial custody chains: keep the attacks in the record

Signed permissions and hash-chained receipts catch real bypasses. Two acceptance-path defects were still open in the source report.

GPT-6 Sol vs. Luna vs. Astra: a routing policy for coding agents — and when local still wins

A routing policy for the actual GPT-6 lineup: Sol for discovery and coding, Luna for bounded high-volume processing, Astra when nothing else holds — each with a verification contract, and the lanes where a local 27B still beats all three.

The Equalizer Bench: a 3B that can't write ffmpeg, the same 3B shipping video edits, and the bug our own benchmark caught

We benchmarked our own thesis: tiny model + deterministic guardrail layer vs raw capability. The curve is textbook, the trim trap is real, and the bench indicted our own product before anyone else could.

The One-Line Bug That Crashed Our Fast Lane: finding, fixing, and measuring a speculative-decoding crash on a $1,400 mini-PC

Spec decoding crashed our fastest lane on day one. The bug was one missing line in our fork; the fix bought +17.7% and a surprise about draft quantization.

Dos modelos, una mini-PC de $1,400: los números emparejados, fracasos incluidos

Un modelo de razonamiento de 35B ya corre junto a nuestro 27B diario en una caja de $1,400, al mismo tiempo. Cada número emparejado, mismos problemas, misma máquina. Los fracasos también están aquí.

How I became a forward deployed engineer without a software engineer title

The title is new; the work is old. Twelve years of enterprise deployments plus public, measured AI work. The honest path, artifacts included.

Evals are the FDE skill nobody lists: my 495-trial public benchmark

The market says it cannot find people who can build AI evals. The skill is learnable and I published mine: 495 trials, certified floors, sabotage cell, open source.

Forward deployed vs solutions engineer vs implementation vs customer engineer: el decodificador

Cuatro títulos, una familia de trabajo, barras de código distintas. Un decodificador para leer cualquier vacante y saber en qué te estás metiendo.

Qué hace realmente un forward deployed engineer

La respuesta directa y luego los recibos: todo el trabajo de un FDE sobre un mini-PC de $1.400, con evals públicas.

The Delegation Card: we asked a $1,400 mini-PC to take our jobs

Not is-it-smart but can-you-hand-it-work-and-walk-away. 495 certified trials, then re-validated at deeper n after the product changed: 965 total, floors to 92.8%.

Qwen 3.8 27B en Strix Halo: la historia completa, medida

Cada dial medido, cada numero publico: la configuracion optima congelada para un 27B en un mini-PC de $1.400.

Notas: the measured-knees method for reasoning effort

A methods note on reasoning-effort calibration. Thinking rescued 15/40 hards vs 4/40 off. On easy tasks it bought nothing. Publish the knee.

Notas de lab: 67% LiveCodeBench-30 en un rig de $1.400

20/30 = 67% LiveCodeBench-30 en un mini-PC de $1.400. IC Wilson 95% 49-81%. Easy 10/10, medium 8/10, hard 2/10. Subset n=30. No es la card.

Notas de lab: revertimos una regresión de llama.cpp

Encontramos una regresión de llama.cpp y la revertimos. Batería n=6 en un rig de $1.400: 0/6 slashes antes, 5/6 después. Mismo Q4_K_XL. No es una historia de quant.

Notas de lab: 93% HumanEval en un rig de $1.400

28/30 = 93% HumanEval en un mini-PC de $1.400. Qwen3.8-27B Q4_K_XL. Temp 0, thinking off. Fallos: 50 y 145. Log crudo linkeado.

Notas de lab: el modelo que no puede olvidar pero no puede recordar

El número primero: este modelo es 75% no transformer. 48 de 64 capas guardan un estado. En un rig de $1.400, carga 198k una vez (1818s) y consulta en 9-27s.

Notas de lab: la cuenca era un bug

Publicamos una cuenca. El hueco era el binario. Tras revertir c7d8722: 6/6 HIT a 198k. Mapa n=1 del build arreglado.

Notas de lab: el veredicto del KV

La pregunta q8-a-q4 del KV de la noche 1 ya tiene respuesta pareada: misma accuracy, cero tripwires, la mitad del cache. Y el label que tuvimos que corregir en público cuando el contador del server le ganó al estimado.

Notas de lab: la noche del veredicto

Un mini-PC de $1.400 sirviendo un 27B a 262k de contexto hace una pregunta: ¿capear cuánto piensa el modelo te cuesta accuracy? La respuesta pareada, el primer intento inválido, y la doctrina que quedó.

Un mini-PC, una noche y los numeros que discutian entre si: afinando Qwen3.8-27B en Strix Halo

Una noche y una tarde de medicion afinando un modelo denso de 27B en un mini-PC Strix Halo: las bandas honestas, la reversion, la doctrina de crashes y la historia detectivesca del EC.

Los agentes necesitan herramientas verificables, no mejor teatro de prompts

El patron util no es un prompt mas bonito. Es una superficie de herramienta que el agente puede llamar, inspeccionar, verificar y revisar.

El historial del repo es una señal de producto

Un repo no es solo almacenamiento. Es evidencia de decisiones, reparaciones, releases, cambios de nombre, huecos de pruebas y oficio real.

La ayuda de implementacion es parte de la superficie del producto

Una herramienta open source util todavia necesita una ruta desde repo publico hasta entorno funcionando. Esa ruta es producto.

Por que importa Kinocut

Kinocut da a los agentes de IA herramientas llamables sobre timelines, efectos, Hyperframes y medios terminados en kinocut.dev.

Monos infinitos, LLMs y el cuarto alrededor de la maquina

El argumento detras del video: la calidad no es solo probabilidad. Es arquitectura, filtros y gusto humano.

Lo que una herramienta de IA necesita antes de que alguien pueda usarla

Una guia practica para convertir una herramienta, flujo o app medio cruda en algo que otros puedan entender, instalar y usar.

Checklist de implementacion para servidores MCP

El checklist que Kyanite usa para distinguir un servidor MCP de juguete, una herramienta usable y algo que vale la pena implementar.

La arqueologia de repos convierte historia en evidencia

Por que el historial de commits es una de las fuentes de prueba mas fuertes para diagnosticos, implementacion y confianza tecnica.

El descubrimiento por IA necesita mas que un sitemap

Lo que Kyanite agrega para que buscadores y asistentes de IA entiendan herramientas, productos, prueba y rutas de soporte.

La Contribution Economy de ResonantDAO, explicada con recibos

Lo contrario de un DAO normal: primero la contribución verificada, el token después. El primer análisis en español del movimiento, con recibos.