---
author: "Umesh Malik"
canonical: "https://umesh-malik.com/blog/tag/linear-attention"
description: "Explore articles tagged with Linear Attention by Umesh Malik — AI Engineer, LLM & GenAI Developer. Learn Linear Attention best practices, practical tips, and in-depth guides."
title: "Umesh Malik's Blog - Linear Attention Articles | Linear Attention Tutorials"
tokens: 142
generator: "scripts/generate-page-markdown.mjs"
---

[← Back to Blog](https://umesh-malik.com/blog)

# Linear Attention

1 article

 [![Qwen3.8 27B VRAM budget: FP8 weights plus KV cache at 262K context on a single GPU](https://umesh-malik.com/blog/qwen3-8-27b-vram-kv-cache-math-cover.png)

LLM Engineering • Aug 15, 2026

### Qwen3.8 27B VRAM: how to fit 262K context in 16 GiB, not 64

Qwen3.8 27B VRAM math: 25.9 GiB of FP8 weights plus 16 GiB of KV cache at 262,144 tokens, not 64. The arithmetic, and where a 48 GB card breaks.

9 min read

Read more →](https://umesh-malik.com/blog/qwen3-8-27b-vram-kv-cache-math)
