EngramHalo.cpp: Qwen 3.8 Flash Next at 38 t/s on Strix Halo
Following up on the hipCUB fix post, I built and benchmarked EngramHalo.cpp, a llama.cpp fork with RDNA 3.5 kernel patches and MTP speculative decoding. The result: 28-38 tok/s decode on the same hardware and quant, nearly doubling the previous 15-20 tok/s baseline.
read more →