
LLM Engineering •
Qwen3.8 27B VRAM: how to fit 262K context in 16 GiB, not 64
Qwen3.8 27B VRAM math: 25.9 GiB of FP8 weights plus 16 GiB of KV cache at 262,144 tokens, not 64. The arithmetic, and where a 48 GB card breaks.
10 min read
Read more →

