Qwen 3 (30B-A3B MoE) has 30.5 billion parameters. At standard 4-bit quantization with 8K context, it needs roughly 22.4 GB of VRAM — weights plus cache and runtime overhead.

VRAM by quantization

PrecisionWeightsCache/BufferTotal VRAM
2-bit (IQ2_XXS)9.8 GB3.7 GB15.4 GB
4-bit (Q4_K_M)16.8 GB3.7 GB22.4 GB
8-bit (Q8_0)32.0 GB3.7 GB37.7 GB
16-bit (FP16)61.0 GB3.7 GB66.7 GB

Which GPU can run Qwen 3 (30B-A3B MoE) (at 4-bit)?

GPU classVRAMQwen 3 (30B-A3B MoE) (22.4 GB)
8 GB · RTX 5060 / 40608 GBWon’t fit
12 GB · RTX 5070 / 306012 GBWon’t fit
16 GB · RTX 5070 Ti / 408016 GBWon’t fit
24 GB · RTX 4090 / 309024 GBFits
32 GB · RTX 509032 GBFits
48 GB · 2×24 / RTX 6000 Ada48 GBFits
128 GB · M-series / RTX Spark128 GBFits

Speed of 3B, brain of 30B. MoE ideal for fast replies on 16GB GPUs.

Get the exact number for your setup
Pick your model, quantization, and context length — the calculator shows the full VRAM math and tells you precisely which hardware fits.
Open the Local AI Calculator
Related guides
Best GPU for Llama 3 70B How Much VRAM for DeepSeek-R1 Q4 vs Q8 Quantization Explained Apple Silicon for Local AI RTX Spark: 128GB Unified Memory

VRAM figures are reproducible estimates (weights + KV cache + overhead) and vary by runtime and quant format. Data current as of 2026-07-05.