Local LLM / Servingspeculative decoding llama.cpp draft model iGPU speedup
The One-Line Bug That Crashed Our Fast Lane: finding, fixing, and measuring a speculative-decoding crash on a $1,400 mini-PC
Spec decoding crashed our fastest lane on day one. The bug was one missing line in our fork; the fix bought +17.7% and a surprise about draft quantization.
Local AI Infrastructurerun two local llms on one mini pc paired benchmark
Dos modelos, una mini-PC de $1,400: los números emparejados, fracasos incluidos
Un modelo de razonamiento de 35B ya corre junto a nuestro 27B diario en una caja de $1,400, al mismo tiempo. Cada número emparejado, mismos problemas, misma máquina. Los fracasos también están aquí.
Forward Deployed Engineeringhow to become a forward deployed engineer without degree
How I became a forward deployed engineer without a software engineer title
The title is new; the work is old. Twelve years of enterprise deployments plus public, measured AI work. The honest path, artifacts included.