LLM GPU Comparison

Side-by-side comparison of 15 GPUs for running local LLMs. Filter by model size, price, or brand, then sort any column. Each row links to our in-depth guide.

Figures reviewed September 15, 2026.

If you just want the answer

Or filter the full table

Any

15 of 15 GPUs match your filters

GPU VRAM Bandwidth Price (new) Price (used) 7B Q4 13B Q4 34B Q4 Best for
RTX 5090 NVIDIA 32GB GDDR7 1792 GB/s $4,900 95 tok/s 55 tok/s 35 tok/s 34B+ LLMs, flagship AI work, future-proofing
NVIDIA A6000 NVIDIA 48GB GDDR6 768 GB/s $3,500 $2,500 55 tok/s 35 tok/s 25 tok/s Single-card 70B LLM inference; ECC memory
RTX 4090 NVIDIA 24GB GDDR6X 1008 GB/s $2,200 65 tok/s 40 tok/s 22 tok/s Best overall for most AI users; 24GB covers 34B LLMs
RTX 5080 NVIDIA 16GB GDDR7 960 GB/s $1,400 55 tok/s 32 tok/s won't fit 16GB + GDDR7 speed for 7B-13B; cheaper than 4090
RTX 5070 Ti NVIDIA 16GB GDDR7 896 GB/s $1,050 48 tok/s 28 tok/s won't fit 16GB Blackwell, priced a little over $1,000
RX 7900 XTX AMD 24GB GDDR6 960 GB/s $900 $850 42 tok/s 25 tok/s 15 tok/s 24GB AMD — Linux-first ROCm users only
RTX 5070 NVIDIA 12GB GDDR7 672 GB/s $875 42 tok/s 20 tok/s won't fit Fast 7B models; 12GB limits 13B headroom
RTX 3090 (used) NVIDIA 24GB GDDR6X 936 GB/s $820 55 tok/s 35 tok/s 20 tok/s Best VRAM per dollar — 24GB at roughly 37% of the 4090
RTX 4070 Ti Super NVIDIA 16GB GDDR6X 672 GB/s $800 $737 40 tok/s 24 tok/s won't fit Best sweet-spot GPU for AI under $1000
RTX 5060 Ti 16GB NVIDIA 16GB GDDR7 448 GB/s $630 32 tok/s 22 tok/s won't fit Blackwell's cheapest entry, 16GB included
RTX 4060 NVIDIA 8GB GDDR6 272 GB/s $479 22 tok/s won't fit Entry-level; 8GB limits model sizes
RX 7800 XT AMD 16GB GDDR6 624 GB/s $450 $380 28 tok/s 18 tok/s won't fit 16GB AMD — ROCm-comfortable users with Linux
RTX 4060 Ti 16GB NVIDIA 16GB GDDR6 288 GB/s $425 $350 25 tok/s 18 tok/s won't fit Cheapest new NVIDIA route to 16GB, for 7B-13B models
Intel Arc B580 INTEL 12GB GDDR6 456 GB/s $310 18 tok/s 10 tok/s won't fit Experimental; lowest cost per GB of any new card, software gaps aside
RTX 3060 12GB (used) NVIDIA 12GB GDDR6 360 GB/s $250 20 tok/s 12 tok/s won't fit Cheapest way into real AI — 12GB at $250

How to read this table

  • VRAM is the hard gate for local LLMs. A model must fit in VRAM plus KV cache overhead — no VRAM fit means it cannot run at that quantization. 16GB is the 2026 practical minimum.
  • Memory bandwidth (GB/s) directly drives tok/s. Same VRAM + higher bandwidth = faster inference. GDDR7 beats GDDR6X beats GDDR6 at comparable tiers.
  • 7B / 13B / 34B tok/s estimates are at Q4_K_M quantization with short context (2-4K tokens). Llama-family behavior; Mistral, Qwen, Gemma will land close. Actual numbers vary with driver, framework (Ollama vs llama.cpp vs vLLM), and context length.
  • "Won't fit" means the target model at Q4 exceeds available VRAM after realistic KV cache overhead — running it would OOM or fall back to painful CPU offload.
  • Price (used) reflects typical eBay / r/hardwareswap pricing at time of publish. Adjust ±15% for seller condition and mining wear.

How to choose — the short version

VRAM first, then bandwidth, then price. Local LLM headaches come from picking a card that can't hold your target model. Match the model size (7B / 13B / 34B / 70B) to the VRAM tier first, then use bandwidth to decide within that tier. For 34B and above, a used RTX 3090 at $820 often beats a new mid-range card at the same price because 24GB of VRAM is the hard gate.

Performance numbers are representative ranges synthesized from manufacturer specifications, Tom's Hardware and TechPowerUp benchmarks, Ollama and llama.cpp community reports, LM Studio benchmark threads, and LocalScore data. This is not first-party testing. See our methodology for how we evaluate GPUs and editorial policy for how we make recommendations.