LLM GPU Comparison
Side-by-side comparison of 15 GPUs for running local LLMs. Filter by model size, price, or brand, then sort any column. Each row links to our in-depth guide.
Figures reviewed September 15, 2026.
If you just want the answer
Runs 34B models, fast Ollama tok/s, proven compatibility.
Best used value RTX 3090 24GB · ~$820 usedThe 4090's 24GB at roughly 37% of its price. Still the king of VRAM-per-dollar.
Best budget RTX 4060 Ti 16GB 16GB · ~$425Runs 7B-13B LLMs comfortably. The entry point into local LLM work.
Best for 34B RTX 4090 24GB · Q4_K_M fitsMinimum practical VRAM for 34B at Q4. 5090 is faster but pricier.
Best for 70B RTX 5090 or 2× RTX 3090 32GB / 48GB combined5090 fits 70B at Q3. Dual 3090s hit Q4 at lower total cost.
Or filter the full table
15 of 15 GPUs match your filters
| GPU | VRAM | Bandwidth | Price (new) | Price (used) | 7B Q4 | 13B Q4 | 34B Q4 | Best for |
|---|---|---|---|---|---|---|---|---|
| RTX 5090 NVIDIA | 32GB GDDR7 | 1792 GB/s | $4,900 | — | 95 tok/s | 55 tok/s | 35 tok/s | 34B+ LLMs, flagship AI work, future-proofing |
| NVIDIA A6000 NVIDIA | 48GB GDDR6 | 768 GB/s | $3,500 | $2,500 | 55 tok/s | 35 tok/s | 25 tok/s | Single-card 70B LLM inference; ECC memory |
| RTX 4090 NVIDIA | 24GB GDDR6X | 1008 GB/s | $2,200 | — | 65 tok/s | 40 tok/s | 22 tok/s | Best overall for most AI users; 24GB covers 34B LLMs |
| RTX 5080 NVIDIA | 16GB GDDR7 | 960 GB/s | $1,400 | — | 55 tok/s | 32 tok/s | won't fit | 16GB + GDDR7 speed for 7B-13B; cheaper than 4090 |
| RTX 5070 Ti NVIDIA | 16GB GDDR7 | 896 GB/s | $1,050 | — | 48 tok/s | 28 tok/s | won't fit | 16GB Blackwell, priced a little over $1,000 |
| RX 7900 XTX AMD | 24GB GDDR6 | 960 GB/s | $900 | $850 | 42 tok/s | 25 tok/s | 15 tok/s | 24GB AMD — Linux-first ROCm users only |
| RTX 5070 NVIDIA | 12GB GDDR7 | 672 GB/s | $875 | — | 42 tok/s | 20 tok/s | won't fit | Fast 7B models; 12GB limits 13B headroom |
| RTX 3090 (used) NVIDIA | 24GB GDDR6X | 936 GB/s | — | $820 | 55 tok/s | 35 tok/s | 20 tok/s | Best VRAM per dollar — 24GB at roughly 37% of the 4090 |
| RTX 4070 Ti Super NVIDIA | 16GB GDDR6X | 672 GB/s | $800 | $737 | 40 tok/s | 24 tok/s | won't fit | Best sweet-spot GPU for AI under $1000 |
| RTX 5060 Ti 16GB NVIDIA | 16GB GDDR7 | 448 GB/s | $630 | — | 32 tok/s | 22 tok/s | won't fit | Blackwell's cheapest entry, 16GB included |
| RTX 4060 NVIDIA | 8GB GDDR6 | 272 GB/s | $479 | — | 22 tok/s | — | won't fit | Entry-level; 8GB limits model sizes |
| RX 7800 XT AMD | 16GB GDDR6 | 624 GB/s | $450 | $380 | 28 tok/s | 18 tok/s | won't fit | 16GB AMD — ROCm-comfortable users with Linux |
| RTX 4060 Ti 16GB NVIDIA | 16GB GDDR6 | 288 GB/s | $425 | $350 | 25 tok/s | 18 tok/s | won't fit | Cheapest new NVIDIA route to 16GB, for 7B-13B models |
| Intel Arc B580 INTEL | 12GB GDDR6 | 456 GB/s | $310 | — | 18 tok/s | 10 tok/s | won't fit | Experimental; lowest cost per GB of any new card, software gaps aside |
| RTX 3060 12GB (used) NVIDIA | 12GB GDDR6 | 360 GB/s | — | $250 | 20 tok/s | 12 tok/s | won't fit | Cheapest way into real AI — 12GB at $250 |
How to read this table
- VRAM is the hard gate for local LLMs. A model must fit in VRAM plus KV cache overhead — no VRAM fit means it cannot run at that quantization. 16GB is the 2026 practical minimum.
- Memory bandwidth (GB/s) directly drives tok/s. Same VRAM + higher bandwidth = faster inference. GDDR7 beats GDDR6X beats GDDR6 at comparable tiers.
- 7B / 13B / 34B tok/s estimates are at Q4_K_M quantization with short context (2-4K tokens). Llama-family behavior; Mistral, Qwen, Gemma will land close. Actual numbers vary with driver, framework (Ollama vs llama.cpp vs vLLM), and context length.
- "Won't fit" means the target model at Q4 exceeds available VRAM after realistic KV cache overhead — running it would OOM or fall back to painful CPU offload.
- Price (used) reflects typical eBay / r/hardwareswap pricing at time of publish. Adjust ±15% for seller condition and mining wear.
How to choose — the short version
VRAM first, then bandwidth, then price. Local LLM headaches come from picking a card that can't hold your target model. Match the model size (7B / 13B / 34B / 70B) to the VRAM tier first, then use bandwidth to decide within that tier. For 34B and above, a used RTX 3090 at $820 often beats a new mid-range card at the same price because 24GB of VRAM is the hard gate.
Performance numbers are representative ranges synthesized from manufacturer specifications, Tom's Hardware and TechPowerUp benchmarks, Ollama and llama.cpp community reports, LM Studio benchmark threads, and LocalScore data. This is not first-party testing. See our methodology for how we evaluate GPUs and editorial policy for how we make recommendations.