Quick answer: Qwen 3 4B needs ~2.5GB at Q4_K_M, 8B ~5.2GB, 14B ~9.3GB, the 30B-A3B MoE ~19GB and 32B ~20GB. A 16GB card covers everything up to the 14B; the 30B-A3B and the 32B are both 24GB models. There is no Qwen 3 72B — if you are looking for one, you are thinking of Qwen 2.5 72B, which is ~47GB at Q4.
NVIDIA GeForce RTX 4060 Ti 16GB
16GB GDDR616GB VRAM fits Qwen 3 14B at Q4_K_M with headroom, and handles Q8 on the 4B comfortably. The most efficient entry point for Qwen 3 inference.
Affiliate links — we may earn a commission at no extra cost to you. Amazon ships globally; Shopee SG covers Singapore & ASEAN.
Qwen 3 VRAM requirements by model size
The Model Size column is the published download from Ollama’s Qwen 3 library, verified 2026-09-13: 4B at 2.6GB / 4.4GB / 8.1GB, 14B at 9.3GB / 16GB / 30GB, 32B at 20GB, 30B-A3B at 19GB and 235B-A22B at 142GB. Rows Ollama does not publish (Q6_K, Q3_K_M) keep a tilde. The VRAM column adds our own KV-cache allowance and is an estimate, not a published figure.
Qwen 3 is Alibaba’s third-generation open model series. The published sizes run 0.6B, 1.7B, 4B, 8B, 14B, 30B-A3B, 32B and 235B-A22B — note that unlike Qwen 2.5 there is no 72B, which is the single most common mix-up with this family. Here are the VRAM numbers at each quantization level:
Qwen 3 4B
| Quantization | Model Size | VRAM Needed | Minimum GPU |
|---|---|---|---|
| FP16 | ~8GB | 10GB | RTX 3060 12GB |
| Q8 | ~4.5GB | 6GB | Any 8GB GPU |
| Q4_K_M | ~2.5GB | ~3GB | Any 4GB GPU |
| Q3_K_M | ~2GB | 3GB | Any 4GB GPU |
The 4B model is extremely lightweight. Even an 8GB laptop GPU handles it at Q8.
Qwen 3 14B
| Quantization | Model Size | VRAM Needed | Minimum GPU |
|---|---|---|---|
| FP16 | 30GB | 32GB+ | Multi-GPU |
| Q8 | 16GB | 18GB+ | 24GB card — not a 16GB one |
| Q6_K | ~10.5GB | 12GB | RTX 3060 12GB |
| Q4_K_M | 9.3GB | ~11GB | RTX 3060 12GB |
| Q3_K_M | ~6.5GB | 8GB | RTX 4060 8GB |
Qwen 3 14B at Q4_K_M is the most popular local setup: a 9.3GB download leaves a 12GB card roughly 2.7GB for context, which covers 8K comfortably and not much beyond. Note the Q8 row — the file is 16GB, so a 16GB card cannot hold it at all, context or no context.
Qwen 3 32B
| Quantization | Model Size | VRAM Needed | Minimum GPU |
|---|---|---|---|
| FP16 | ~64GB | 68GB+ | Multi-GPU |
| Q8 | ~32GB | 34GB+ | Multi-GPU |
| Q6_K | ~24GB | 26GB | RTX 5090 32GB |
| Q4_K_M | 20GB | ~22GB | RTX 4090 24GB |
| Q3_K_M | ~14GB | 16GB | RTX 4060 Ti 16GB (tight) |
Qwen 3 32B at Q4_K_M is a 20GB download, so a 24GB card holds it with about 4GB left for context — enough, but not generous. The RTX 4090 is the natural home for this model size.
Check RTX 4090 Price→Buy on Shopee SG→Qwen 3 30B-A3B (MoE)
| Quantization | Model Size | VRAM Needed | Minimum GPU |
|---|---|---|---|
| BF16 | ~61GB | 65GB+ | Multi-GPU |
| Q8 | ~33GB | 35GB+ | Multi-GPU or RTX 5090 (tight) |
| Q4_K_M | ~19GB | ~21GB | RTX 3090 / RTX 4090 24GB |
The MoE variant activates ~3B parameters per token, so it generates far faster than the dense 32B — but all 30B parameters stay resident, so it needs the same class of card. If you see it described as a model that fits 16GB, that figure came from the active parameter count rather than the file Ollama ships.
Qwen 3 235B-A22B
Listed for completeness: ~142GB at Q4_K_M. That is data-centre hardware or cloud, not a consumer card at any quantization Ollama publishes.
Context length moves these figures more than people expect; the VRAM calculator lets you set it and see where the total lands.
VRAM capacity memory bandwidth Specs are manufacturer figures. Bar lengths are scaled independently per metric.
GPU recommendations by Qwen 3 model size
| Model | Best GPU | Why |
|---|---|---|
| Qwen 3 4B | RTX 4060 (8GB) | Overkill but cheap |
| Qwen 3 14B | RTX 4060 Ti 16GB | Q4_K_M fits with headroom |
| Qwen 3 14B (Q8) | RTX 4060 Ti 16GB | Exact fit, minimal context |
| Qwen 3 32B | RTX 4090 (24GB) | Q4_K_M with comfortable headroom |
| Qwen 3 30B-A3B | RTX 3090 24GB (used) | Q4_K_M at ~19GB, 3B active so it runs fast |
| Qwen 3 235B-A22B | Cloud | ~142GB at Q4_K_M — not a consumer card |
Which GPU should YOU buy for Qwen 3?
- Running Qwen 3 4B or 14B? Get the RTX 4060 Ti 16GB ($400). It handles both at Q4_K_M to Q8 with zero issues. Best value for Qwen 3’s most popular sizes.
- Running Qwen 3 32B? Get the RTX 4090 ($2,200). The 24GB sits the 32B at Q4_K_M comfortably, and it’s the only single consumer card that does it reliably.
- Want the 30B-A3B MoE? Get a used RTX 3090 (~$820) or an RTX 4090. At ~19GB it needs 24GB, the same as the dense 32B — the MoE saves you time per token, not memory.
- Already have a 12GB card (RTX 3060)? It handles Qwen 3 14B at Q4_K_M fine. Upgrade only if you need 32B.
Don’t forget context overhead
Model size is just the floor. Every 8K tokens of active context adds roughly:
- 4B model: ~0.3GB extra VRAM
- 14B model: ~1GB extra VRAM
- 32B model: ~2.5GB extra VRAM
If you’re doing long-context tasks — document summarization, long code review — pad 2-4GB onto the numbers above.
Common mistakes to avoid
- Buying a 16GB card for Qwen 3 32B. The 32B model needs ~20GB at Q4_K_M. 16GB cards cannot fit it even at Q3_K_M reliably. The 32B is a 24GB model.
- Confusing Qwen 3 with Qwen 2.5 requirements. Qwen 3 models are slightly larger than their Qwen 2.5 counterparts at the same parameter count. Don’t reuse old VRAM estimates — and the same caution applies forward: Qwen 3.8’s 27B is an 18GB download, which moves the entry point from 16GB to 24GB.
- Looking for a Qwen 3 72B. There isn’t one. Qwen 2.5 had a 72B (~47GB at Q4_K_M, so a dual-24GB or 48GB setup); Qwen 3’s large end is the 32B dense and the 235B-A22B MoE, with nothing in between.
- Ignoring quantization mix. Q4_K_M (4-bit with mixed 6-bit for critical layers) is meaningfully better than plain Q4. Always prefer Q4_K_M over Q4_0 when both fit.
Final verdict
| Your Qwen 3 target | GPU needed | Price |
|---|---|---|
| 4B daily driver | RTX 4060 8GB | ~$479 |
| 14B at Q4_K_M | RTX 3060 12GB (used) | ~$250 |
| 14B at Q8 | RTX 4060 Ti 16GB | ~$425 |
| 32B at Q4_K_M | RTX 4090 24GB | ~$2,200 |
| 30B-A3B MoE at Q4_K_M | RTX 3090 24GB (used) | ~$820 |
NVIDIA GeForce RTX 4090
24GB GDDR6X24GB VRAM fits Qwen 3 32B at Q4_K_M with room for long context. The definitive single-card solution for Qwen 3's most capable local-runnable size.
Affiliate links — we may earn a commission at no extra cost to you. Amazon ships globally; Shopee SG covers Singapore & ASEAN.
For GPU picks specific to Qwen 3, see our best GPU for Qwen 3 guide. For a broader look at VRAM across all model families, see how much VRAM for local LLM. Running an older Qwen release? Check our best GPU for Qwen guide, or for the 14B variant specifically see how much VRAM for Qwen 14B.
VRAM is the only hard constraint for Qwen 3. Get enough for your target model size and quantization level — everything else is secondary.