
LLM Engineering •
Fix slow LLM inference in macOS VMs: 12.6 → 207 tok/s
LLM inference in macOS VMs collapses to 12.63 tok/s because the guest reports GPU family 5 and llama.cpp drops its matrix kernels. The check, and its limits.
9 min read
Read more →