DeepSeek-V4-Flash (284B-A13B) has 284 billion parameters. At standard 4-bit quantization with 8K context, it needs roughly 192.3 GB of VRAM — weights plus cache and runtime overhead.
VRAM by quantization
| Precision | Weights | Cache/Buffer | Total VRAM |
|---|---|---|---|
| 2-bit (IQ2_XXS) | 90.9 GB | 34.1 GB | 127.0 GB |
| 4-bit (Q4_K_M) | 156.2 GB | 34.1 GB | 192.3 GB |
| 8-bit (Q8_0) | 298.2 GB | 34.1 GB | 334.3 GB |
| 16-bit (FP16) | 568.0 GB | 34.1 GB | 604.1 GB |
Which GPU can run DeepSeek-V4-Flash (284B-A13B) (at 4-bit)?
| GPU class | VRAM | DeepSeek-V4-Flash (284B-A13B) (192.3 GB) |
|---|---|---|
| 8 GB · RTX 5060 / 4060 | 8 GB | Won’t fit |
| 12 GB · RTX 5070 / 3060 | 12 GB | Won’t fit |
| 16 GB · RTX 5070 Ti / 4080 | 16 GB | Won’t fit |
| 24 GB · RTX 4090 / 3090 | 24 GB | Won’t fit |
| 32 GB · RTX 5090 | 32 GB | Won’t fit |
| 48 GB · 2×24 / RTX 6000 Ada | 48 GB | Won’t fit |
| 128 GB · M-series / RTX Spark | 128 GB | Won’t fit |
Apr 2026 MoE (13B active) with 1M context. ~570GB weights. MIT licensed.
Get the exact number for your setup
Pick your model, quantization, and context length — the calculator shows the full VRAM math and tells you precisely which hardware fits.
Open the Local AI Calculator →
VRAM figures are reproducible estimates (weights + KV cache + overhead) and vary by runtime and quant format. Data current as of 2026-07-05.