How Much VRAM for Qwen 3 in 2026? Full Size Breakdown

VRAM for Qwen 3 in 2026 — 4B needs 2.5GB, 14B needs 9.3GB, 32B needs 20GB and the 30B-A3B MoE 19GB at Q4. Sizes from Ollama's own listing.

Quick read: This guide is built to help you match model size, VRAM, and budget before you buy.

Quick answer: Qwen 3 4B needs ~2.5GB at Q4_K_M, 8B ~5.2GB, 14B ~9.3GB, the 30B-A3B MoE ~19GB and 32B ~20GB. A 16GB card covers everything up to the 14B; the 30B-A3B and the 32B are both 24GB models. There is no Qwen 3 72B — if you are looking for one, you are thinking of Qwen 2.5 72B, which is ~47GB at Q4.

Best for Qwen 3 14B

NVIDIA GeForce RTX 4060 Ti 16GB

16GB GDDR6

16GB VRAM fits Qwen 3 14B at Q4_K_M with headroom, and handles Q8 on the 4B comfortably. The most efficient entry point for Qwen 3 inference.

Affiliate links — we may earn a commission at no extra cost to you. Amazon ships globally; Shopee SG covers Singapore & ASEAN.

Qwen 3 VRAM requirements by model size

The Model Size column is the published download from Ollama’s Qwen 3 library, verified 2026-09-13: 4B at 2.6GB / 4.4GB / 8.1GB, 14B at 9.3GB / 16GB / 30GB, 32B at 20GB, 30B-A3B at 19GB and 235B-A22B at 142GB. Rows Ollama does not publish (Q6_K, Q3_K_M) keep a tilde. The VRAM column adds our own KV-cache allowance and is an estimate, not a published figure.

Qwen 3 is Alibaba’s third-generation open model series. The published sizes run 0.6B, 1.7B, 4B, 8B, 14B, 30B-A3B, 32B and 235B-A22B — note that unlike Qwen 2.5 there is no 72B, which is the single most common mix-up with this family. Here are the VRAM numbers at each quantization level:

Qwen 3 4B

QuantizationModel SizeVRAM NeededMinimum GPU
FP16~8GB10GBRTX 3060 12GB
Q8~4.5GB6GBAny 8GB GPU
Q4_K_M~2.5GB~3GBAny 4GB GPU
Q3_K_M~2GB3GBAny 4GB GPU

The 4B model is extremely lightweight. Even an 8GB laptop GPU handles it at Q8.

Qwen 3 14B

QuantizationModel SizeVRAM NeededMinimum GPU
FP1630GB32GB+Multi-GPU
Q816GB18GB+24GB card — not a 16GB one
Q6_K~10.5GB12GBRTX 3060 12GB
Q4_K_M9.3GB~11GBRTX 3060 12GB
Q3_K_M~6.5GB8GBRTX 4060 8GB

Qwen 3 14B at Q4_K_M is the most popular local setup: a 9.3GB download leaves a 12GB card roughly 2.7GB for context, which covers 8K comfortably and not much beyond. Note the Q8 row — the file is 16GB, so a 16GB card cannot hold it at all, context or no context.

Qwen 3 32B

QuantizationModel SizeVRAM NeededMinimum GPU
FP16~64GB68GB+Multi-GPU
Q8~32GB34GB+Multi-GPU
Q6_K~24GB26GBRTX 5090 32GB
Q4_K_M20GB~22GBRTX 4090 24GB
Q3_K_M~14GB16GBRTX 4060 Ti 16GB (tight)

Qwen 3 32B at Q4_K_M is a 20GB download, so a 24GB card holds it with about 4GB left for context — enough, but not generous. The RTX 4090 is the natural home for this model size.

Check RTX 4090 PriceBuy on Shopee SG

Qwen 3 30B-A3B (MoE)

QuantizationModel SizeVRAM NeededMinimum GPU
BF16~61GB65GB+Multi-GPU
Q8~33GB35GB+Multi-GPU or RTX 5090 (tight)
Q4_K_M~19GB~21GBRTX 3090 / RTX 4090 24GB

The MoE variant activates ~3B parameters per token, so it generates far faster than the dense 32B — but all 30B parameters stay resident, so it needs the same class of card. If you see it described as a model that fits 16GB, that figure came from the active parameter count rather than the file Ollama ships.

Qwen 3 235B-A22B

Listed for completeness: ~142GB at Q4_K_M. That is data-centre hardware or cloud, not a consumer card at any quantization Ollama publishes.

Context length moves these figures more than people expect; the VRAM calculator lets you set it and see where the total lands.

VRAM capacity vs memory bandwidth
RTX 5090 32GB · 1792 GB/s RTX 4090 24GB · 1008 GB/s RX 7900 XTX 24GB · 960 GB/s RTX 3090 (used) 24GB · 936 GB/s RTX 5080 16GB · 960 GB/s RTX 5070 Ti 16GB · 896 GB/s RTX 4070 Ti Super 16GB · 672 GB/s RX 7800 XT 16GB · 624 GB/s RTX 5060 Ti 16GB 16GB · 448 GB/s RTX 4060 Ti 16GB 16GB · 288 GB/s RTX 5070 12GB · 672 GB/s Intel Arc B580 12GB · 456 GB/s RTX 3060 12GB (used) 12GB · 360 GB/s RTX 4060 8GB · 272 GB/s

VRAM capacity memory bandwidth Specs are manufacturer figures. Bar lengths are scaled independently per metric.

GPU recommendations by Qwen 3 model size

ModelBest GPUWhy
Qwen 3 4BRTX 4060 (8GB)Overkill but cheap
Qwen 3 14BRTX 4060 Ti 16GBQ4_K_M fits with headroom
Qwen 3 14B (Q8)RTX 4060 Ti 16GBExact fit, minimal context
Qwen 3 32BRTX 4090 (24GB)Q4_K_M with comfortable headroom
Qwen 3 30B-A3BRTX 3090 24GB (used)Q4_K_M at ~19GB, 3B active so it runs fast
Qwen 3 235B-A22BCloud~142GB at Q4_K_M — not a consumer card

Which GPU should YOU buy for Qwen 3?

  • Running Qwen 3 4B or 14B? Get the RTX 4060 Ti 16GB ($400). It handles both at Q4_K_M to Q8 with zero issues. Best value for Qwen 3’s most popular sizes.
  • Running Qwen 3 32B? Get the RTX 4090 ($2,200). The 24GB sits the 32B at Q4_K_M comfortably, and it’s the only single consumer card that does it reliably.
  • Want the 30B-A3B MoE? Get a used RTX 3090 (~$820) or an RTX 4090. At ~19GB it needs 24GB, the same as the dense 32B — the MoE saves you time per token, not memory.
  • Already have a 12GB card (RTX 3060)? It handles Qwen 3 14B at Q4_K_M fine. Upgrade only if you need 32B.
Check RTX 4060 Ti 16GB PriceBuy on Shopee SG Check RTX 4090 PriceBuy on Shopee SG

Don’t forget context overhead

Model size is just the floor. Every 8K tokens of active context adds roughly:

  • 4B model: ~0.3GB extra VRAM
  • 14B model: ~1GB extra VRAM
  • 32B model: ~2.5GB extra VRAM

If you’re doing long-context tasks — document summarization, long code review — pad 2-4GB onto the numbers above.

Common mistakes to avoid

  • Buying a 16GB card for Qwen 3 32B. The 32B model needs ~20GB at Q4_K_M. 16GB cards cannot fit it even at Q3_K_M reliably. The 32B is a 24GB model.
  • Confusing Qwen 3 with Qwen 2.5 requirements. Qwen 3 models are slightly larger than their Qwen 2.5 counterparts at the same parameter count. Don’t reuse old VRAM estimates — and the same caution applies forward: Qwen 3.8’s 27B is an 18GB download, which moves the entry point from 16GB to 24GB.
  • Looking for a Qwen 3 72B. There isn’t one. Qwen 2.5 had a 72B (~47GB at Q4_K_M, so a dual-24GB or 48GB setup); Qwen 3’s large end is the 32B dense and the 235B-A22B MoE, with nothing in between.
  • Ignoring quantization mix. Q4_K_M (4-bit with mixed 6-bit for critical layers) is meaningfully better than plain Q4. Always prefer Q4_K_M over Q4_0 when both fit.

Final verdict

Your Qwen 3 targetGPU neededPrice
4B daily driverRTX 4060 8GB~$479
14B at Q4_K_MRTX 3060 12GB (used)~$250
14B at Q8RTX 4060 Ti 16GB~$425
32B at Q4_K_MRTX 4090 24GB~$2,200
30B-A3B MoE at Q4_K_MRTX 3090 24GB (used)~$820
Best for Qwen 3 32B

NVIDIA GeForce RTX 4090

24GB GDDR6X

24GB VRAM fits Qwen 3 32B at Q4_K_M with room for long context. The definitive single-card solution for Qwen 3's most capable local-runnable size.

Affiliate links — we may earn a commission at no extra cost to you. Amazon ships globally; Shopee SG covers Singapore & ASEAN.

For GPU picks specific to Qwen 3, see our best GPU for Qwen 3 guide. For a broader look at VRAM across all model families, see how much VRAM for local LLM. Running an older Qwen release? Check our best GPU for Qwen guide, or for the 14B variant specifically see how much VRAM for Qwen 14B.

VRAM is the only hard constraint for Qwen 3. Get enough for your target model size and quantization level — everything else is secondary.

Affiliate Disclosure: This article may contain affiliate links. If you purchase through these links, we may earn a commission at no extra cost to you. Learn more
← Back to all guides