DeepSeek-V4-Pro (1.6T MoE) has 1600 billion parameters. At standard 4-bit quantization with 8K context, it needs roughly 1074.0 GB of VRAM — weights plus cache and runtime overhead.

VRAM by quantization

PrecisionWeightsCache/BufferTotal VRAM
2-bit (IQ2_XXS)512.0 GB192.0 GB706.0 GB
4-bit (Q4_K_M)880.0 GB192.0 GB1074.0 GB
8-bit (Q8_0)1680.0 GB192.0 GB1874.0 GB
16-bit (FP16)3200.0 GB192.0 GB3394.0 GB

Which GPU can run DeepSeek-V4-Pro (1.6T MoE) (at 4-bit)?

GPU classVRAMDeepSeek-V4-Pro (1.6T MoE) (1074.0 GB)
8 GB · RTX 5060 / 40608 GBWon’t fit
12 GB · RTX 5070 / 306012 GBWon’t fit
16 GB · RTX 5070 Ti / 408016 GBWon’t fit
24 GB · RTX 4090 / 309024 GBWon’t fit
32 GB · RTX 509032 GBWon’t fit
48 GB · 2×24 / RTX 6000 Ada48 GBWon’t fit
128 GB · M-series / RTX Spark128 GBWon’t fit

Apr 2026 flagship MoE (49B active), 1M context. ~3.2TB weights. MIT licensed.

Get the exact number for your setup
Pick your model, quantization, and context length — the calculator shows the full VRAM math and tells you precisely which hardware fits.
Open the Local AI Calculator
Related guides
Best GPU for Llama 3 70B How Much VRAM for DeepSeek-R1 Q4 vs Q8 Quantization Explained Apple Silicon for Local AI RTX Spark: 128GB Unified Memory

VRAM figures are reproducible estimates (weights + KV cache + overhead) and vary by runtime and quant format. Data current as of 2026-07-05.