Local Inference Hardware Guide

Choose the right GPU for local LLM inference.

Best GPU for LLM helps you match Ollama, Llama, and private local inference workloads to the right hardware. We focus on VRAM, throughput, and model fit so you do not overspend or buy the wrong tier.

92 guides 15 GPUs tracked updated 2026-09 ollama + llama.cpp focus

Start here

Choose a path based on your actual workload

These are the pages most readers should hit first before browsing the full archive.

Browse by topic

Find a guide for what you're actually doing

Grouped by model family, hardware, and workflow — not by date.

$ tools

Calculators that answer the hard questions

Plug in your model and get an answer instead of guessing.

Full archive

All 92 guides

Every guide we've published, filterable by category.

Best GPU for Qwen 3.8 in 2026: Why 16GB Is Not Enough Qwen 3.8 27B ships as an 18GB Q4_K_M download, so a 16GB card cannot hold it. VRAM tiers, GPU picks, and what the vision encoder costs. buyer-guide
Best GPU for DeepSeek V4: The Honest VRAM Math (81GB Minimum) DeepSeek V4-Flash needs roughly 81-96GB for its smallest quants. The real numbers for 4x RTX 3090 rigs, 128GB Mac Studio, and cloud H200s. buyer-guide
Can You Run Kimi K3 Locally? No — Here's the Exact Math Kimi K3's 2.8T open weights need roughly 1.4TB of GPU memory — no consumer rig comes close. The honest math, and what to run at home instead. guide
Rent a GPU for LLM Fine-Tuning: The $30 Weekend Project Fine-tune a 7B-13B model on a rented A100 for roughly $20-40 a weekend instead of buying a $2,200 RTX 4090. When renting wins and when owning pays off. guide
Best Cloud GPU for LLM in 2026: What to Rent by Model Size Rent an RTX 4090 from ~$0.35/hr for 7B-13B models, an H100 at ~$2-3/hr for 70B. The exact cloud GPU to rent for every LLM size in 2026. buyer-guide
Best GPU for Nemotron TwoTower in 2026: 5 GPUs Ranked NVIDIA's first diffusion LLM: 60B total, only 3B active per tower. Real VRAM is 32-48GB, not 120GB. RTX 5090 32GB works with Q4; 5 GPUs ranked. buyer-guide
Best GPU for LongCat 2 in 2026: 1.6T MoE, 1M Context LongCat 2.0 is a 1.6T MoE with 33-56B active and 1M context. Q4 lands near 990GB and even Q2 near 540GB. No consumer path — rented multi-GPU only. buyer-guide
Best GPU for MiniMax M3? Why 427B Won't Fit a 5090 MiniMax M3 is 427B — its Q4 GGUF is 264GB, not 32GB. What running it actually takes, why KV cache is the second problem, and what to run instead. buyer-guide
RTX 5090 vs H100 for LLM in 2026 ($5K vs $30K Debate) RTX 5090 32GB is enough for 90% of local LLM users. H100 80GB earns its 6× price only for FP8 training or long-context serving. Full compare 2026. comparison
Best GPU for MLX in 2026: Apple Silicon Ranked for Local LLM MLX + Ollama 0.30.8 makes Apple Silicon competitive with CUDA. M4 Max 64GB runs 70B Q4. Ranked M3/M4/Pro/Max/Ultra RAM tiers for local LLM 2026. buyer-guide
Qwen3-Coder-Next Needs 48.5GB at Q4: No Single 24GB Card Qwen3-Coder-Next's Q4_K_M download is 48.5GB, so no 24GB card runs it. Dual RTX 3090 at $1,640 holds Q3 with 128K context. Five setups ranked. buyer-guide
Best GPU for Kimi K2: Why It Won't Run on Consumer Cards Kimi K2 is a 1T MoE — roughly 600GB at Q4, and Ollama offers it cloud-only. What running it actually takes, and what to buy for agents instead. buyer-guide
Cloud GPU vs Self-Hosted LLM: Real TCO Breakdown Full cost comparison — RunPod, Vast.ai, and Lambda vs buying your own GPU for local LLM inference. Break-even analysis included. comparison
How Much VRAM for Gemma 4? The 26B MoE Wants 18GB Gemma 4 sizes do not match their names: the 26B-A4B is an 18GB download and the 12B Dense only 7.6GB. Published sizes for every variant, with GPU picks. guide
Best Motherboard for Dual GPU LLM in 2026 (PCIe 5) Top motherboards for running two GPUs for local LLM inference in 2026 — PCIe slot spacing, lane allocation, and budget picks. buyer-guide
Best Quantization for Local LLM in 2026 (Q4 to Q8) Q4_K_M vs Q5_K_M vs Q6_K vs Q8 in 2026 — which quantization gives the best quality-to-VRAM tradeoff for local LLM inference? guide
GPU Shortage 2026: Should You Buy Now for LLM? GPU prices are surging in 2026 due to GDDR7 shortage and AI demand. Here's whether to buy now or wait for local LLM use. guide
Dual RTX 3090 for 70B LLMs: 48GB Build Guide 2026 Two used 3090s reach 48GB and run Llama 70B at Q4 around 18-22 tok/s. No NVLink needed: motherboard, PSU sizing, how to split, and real costs. guide
PSU for Dual GPU LLM: 1200W for 3090s, 1500W for 4090s Dual RTX 3090s draw ~700W and want a 1200W PSU; dual 4090s need 1500W. Per-pair wattage table, rail requirements and what the minimum really is. guide
Llama 4 Maverick Hardware Guide (400B MoE) for 2026 What hardware do you need for Llama 4 Maverick 400B? Multi-GPU requirements, cloud options, and whether it's worth self-hosting. guide
Best GPU for Gemma 4: The 26B MoE Needs a 24GB Card Gemma 4 spans 4GB to 64GB across its variants, and the 26B-A4B MoE is an 18GB model despite its name. What each one needs, and what to buy. buyer-guide
Best GPU for Qwen 3.6 in 2026 (35B-A3B MoE Guide) Qwen 3.6 35B-A3B is a 24GB model, not a 16GB one. VRAM by build, why MoE sizing surprises people, and the GPUs that actually run it. buyer-guide
RTX 3060 Relaunch: What Budget LLM Buyers Got NVIDIA relaunched the RTX 3060 12GB at $329 in June 2026 and street prices have climbed since. What that means for budget local LLM builds. guide
RTX 5070 Ti vs RTX 3090 for LLM: New $1,050 vs Used $820 RTX 5070 Ti (16GB GDDR7) vs used RTX 3090 (24GB GDDR6X) for local LLMs in 2026 — tok/s, VRAM, and which is the better buy. comparison
Best GPU for Continue.dev (Local AI Coding) in 2026 Best GPU for Continue.dev in 2026 — run a local Copilot with Ollama. RTX 4060 Ti 16GB for 14B, RTX 4090 for 33B code models. buyer-guide
Best GPU for Llama 4 Scout (109B MoE) in 2026 Ranked Llama 4 Scout is 67GB at Q4, so no single consumer GPU runs it. The multi-card builds that do, the low-bit ones that fit less, and rental costs. buyer-guide
Best GPU for Running a Local Coding LLM in 2026 Best GPUs for local AI coding in 2026 — run DeepSeek Coder, Qwen Coder, and other code LLMs as a private Copilot alternative. buyer-guide
Can the RTX 4060 Ti Run 13B Models in 2026? (Honest) Can the RTX 4060 Ti run 13B models in 2026? 16GB version: yes (Q4-Q6). 8GB version: barely (Q3 only). Full VRAM breakdown. guide
LM Studio vs Ollama in 2026: Which Local LLM Tool Should You Use? LM Studio vs Ollama compared — GUI vs CLI, MLX vs CUDA performance, ease of use, and which tool fits your workflow in 2026. comparison
RTX 5090 vs RTX 3090 for LLM: New Flagship vs Used Value King RTX 5090 ($4,900) vs used RTX 3090 ($820) for LLM in 2026 — 32GB GDDR7 vs 24GB GDDR6X. Is the new flagship worth 6x the price? comparison
Can a Mac Mini Run Local LLMs in 2026? M6 vs M5 Pro The M6 Mac mini starts at 16GB and 153 GB/s; the M5 Pro brings 24GB and 307 GB/s. Which local models each one runs, and where each one stops. guide
Intel Arc B580 for Local LLM: Can Intel's Budget GPU Run Models? Intel Arc B580 for LLM in 2026 — 12GB at $310. Runs 7B via llama.cpp Vulkan, but Ollama support is limited. Honest verdict here. guide
RTX 5070 vs RTX 4090 for LLM in 2026: 12GB vs 24GB RTX 5070 vs RTX 4090 for local LLM in 2026 — 12GB vs 24GB VRAM. The 4090 wins because VRAM is king for LLMs. Full comparison. comparison
Best GPU for LLM Summarization in 2026 (5 Picks) Best GPU for LLM summarization — long context needs extra VRAM for KV cache. RTX 4090 is the sweet spot for 32K context. buyer-guide
Local LLM Under $300 in 2026: What 12GB Actually Loads One used RTX 3060 12GB is all that stays under $300 in 2026. Model by model: what its 12GB loads at Q4, and where the 2026 MoE tier stops it. buyer-guide
Best GPU for LM Studio 2026: RTX 4090, or a Used 3090 LM Studio's MLX path is real but CUDA still wins tok/s per dollar. The RTX 4090 is our NVIDIA pick; Apple's M5 Max is the 34B answer. buyer-guide
Best GPU for Qwen 3 in 2026 (4B to 72B Compared) Best GPUs for running Qwen 3 locally in 2026 — from 4B to 72B variants. VRAM requirements, speed comparisons, and hardware picks. buyer-guide
Mac M5 vs NVIDIA for Local LLM: 512GB at 1.2 TB/s Apple's M5 Ultra pairs 512GB of unified memory with 1.2 TB/s of bandwidth. What that changes against an RTX 5090 for local LLM inference, and what it does not. comparison
Ollama vs llama.cpp vs vLLM: Start, Speed, or Serve Ollama to get running in seconds, llama.cpp for the extra tokens per second, vLLM to serve other people. Which of the three fits your workflow in 2026. comparison
RTX 5060 Ti vs RTX 4060 Ti for LLM Inference in 2026 RTX 5060 Ti vs 4060 Ti for LLM — both 16GB, but GDDR7 is 55% faster bandwidth. Speed vs value comparison with benchmarks. comparison
Best GPU for Gemma 3 in 2026 (4B-27B Picks Ranked) Best GPUs for running Google's Gemma 3 locally in 2026 — from 4B to 27B variants. VRAM needs, speed, and hardware picks by budget. buyer-guide
Best GPU for Microsoft Phi-4 in 2026 (5 Picks Ranked) Best GPUs for running Microsoft Phi-4 locally in 2026 — small 14B model that runs comfortably on $400 GPUs. VRAM and speed picks. buyer-guide
How Much VRAM for Llama 4 in 2026? Scout vs Maverick Llama 4 Scout needs ~67GB at Q4 and Maverick ~245GB — far past any consumer GPU. The real numbers, why they surprise people, and what to run. guide
Best GPU for Llama 4 in 2026: Scout & Maverick Guide Llama 4 Scout is 67GB at Q4, so two 24GB cards fall short and Maverick needs 245GB. The builds that actually load them, and what they cost. buyer-guide
How Much VRAM for a 70B LLM in 2026? (Q4-Q8 Table) Llama 3.3 70B ships as a 43GB Q4_K_M download — two 24GB cards hold it, one 32GB card does not. Published sizes for every quantization level. guide
How Much VRAM for Qwen 3 in 2026? Full Size Breakdown VRAM for Qwen 3 in 2026 — 4B needs 2.5GB, 14B needs 9.3GB, 32B needs 20GB and the 30B-A3B MoE 19GB at Q4. Sizes from Ollama's own listing. guide
Used RTX 3090 for LLM: Inspection Checklist and Red Flags A used RTX 3090 runs about $820 for 24GB. What to test before you pay, the red flags that mean walk away, and the ex-mining faults to expect. guide
Best GPU for Gemma 2B-27B in 2026 (6 Picks Ranked) Run Google Gemma locally — VRAM needs for 2B, 7B, and 27B models. Inference speed comparisons and budget-friendly GPU picks. buyer-guide
Best GPU for LLM Fine-Tuning in 2026 (Ranked Picks) Best GPUs for LoRA, QLoRA, and full fine-tuning of LLMs. VRAM requirements, speed benchmarks, and practical recommendations. buyer-guide
Windows vs Linux for Local LLM: Which OS Wins in 2026? Windows vs Linux for local LLM inference — performance differences, VRAM efficiency, multi-GPU support, and when WSL is good enough. comparison
Best Budget GPU for Local LLM 2026: RTX 3060 to $350 RTX 3060 12GB at $250 runs 7B models. RTX 4060 Ti 16GB at $425 handles 13B. 5 budget GPU picks ranked for Ollama + llama.cpp in 2026. buyer-guide
Llama 70B 2026: Why 24GB Isn't Enough (Real Builds) 24GB can't run Llama 70B at usable quality. Dual RTX 3090 at $1,640 is the floor. 4 working builds ranked by tok/s + total cost for 2026. buyer-guide
Best GPU for Local LLM Under $2000 in 2026 (Ranked) The RTX 5090 now runs ~$4,900. Under $2,000 in 2026, dual used RTX 3090s give 48GB and 70B at Q4 — what actually fits the budget, ranked. buyer-guide
Local LLM VRAM 2026: The 12GB Trap Most Buyers Hit Most '16GB is enough' advice misses what breaks at 34B+. Full Q4-Q8 VRAM tiers + the budget mistake that costs you a year of upgrades. guide
Best GPU for DeepSeek Models in 2026 (Picks Ranked) Best GPUs for running DeepSeek-R1, DeepSeek Coder, and DeepSeek V3 locally. VRAM needs, speed benchmarks, and top picks. buyer-guide
Best GPU for Open WebUI in 2026 (5 Picks Compared) Best GPUs for running Open WebUI with Ollama in 2026 — fast local chat interface, practical hardware recommendations from $250. buyer-guide
Best GPU for Microsoft Phi-3 in 2026 (Picks Ranked) Best GPUs for running Phi-3 Mini, Small, and Medium locally in 2026 — VRAM needs, speed comparisons, and budget-friendly picks. buyer-guide
Best GPU for Text Generation WebUI in 2026 (Ranked) Best GPUs for running Oobabooga's Text Generation WebUI locally. VRAM needs for popular models, speed benchmarks, and buying advice. buyer-guide
How Much VRAM for Qwen 14B in 2026? (Q4-Q8 Guide) Exact VRAM requirements for Qwen 2.5 14B and Qwen 3 14B in 2026 at every quantization level — Q4, Q5, Q6, Q8 — plus GPU picks. guide
Best GPU for Local LLM Under $1,500 in 2026 (Ranked) The RTX 4090 left this tier at ~$2,200. A used RTX 3090 is the only 24GB card under $1,500, with tok/s from 7B to 32B compared. buyer-guide
Best GPU for Local Whisper Transcription in 2026 Best GPUs for running Whisper locally in 2026 for private audio transcription. Real-time speed comparisons and hardware picks. buyer-guide
Can the RTX 4060 Ti Run Llama 70B in 2026? (Honest) Can a 16GB RTX 4060 Ti actually run Llama 70B in 2026? Honest answer with quantization analysis and practical alternatives. guide
How Much VRAM Do You Need for Llama 3 8B in 2026? Exact VRAM requirements for Llama 3 8B at every quantization level, including context length overhead and practical GPU recommendations. guide
Can an RTX 3060 Run Ollama in 2026? (Honest Guide) Which LLMs an RTX 3060 12GB runs through Ollama, which models to pull in 2026, and why the two new MoE releases need far more than its 12GB. guide
Best GPU for AI Agents in 2026 (5 Picks Ranked) Which GPU runs local AI agents well in 2026? VRAM, speed, and hardware picks for autonomous agent workflows from $400 to $2,000. buyer-guide
How to Run a 70B LLM on a Single GPU in 2026 (Q3-Q4) Run Llama 3 70B and other 70B models on one GPU using aggressive quantization. VRAM requirements, quality trade-offs, and practical setup guide. guide
RTX 4090 vs RTX 3090 for Ollama: Worth 2.7x the Price? RTX 4090 vs RTX 3090 for Ollama compared in 2026. Same 24GB VRAM, different speed — is the 4090 worth 2.7x the used 3090 price? comparison
RTX 5080 vs RTX 4090 for LLM: Which Is Better in 2026? RTX 5080 16GB vs RTX 4090 24GB for local LLM inference. Benchmarks, VRAM analysis, and which card wins for your model size. comparison
Best GPU for 7B Parameter Models in 2026 (Ranked) Best GPUs for running 7B LLMs in 2026 — Llama 3 8B, Mistral 7B, and Qwen 7B locally with Ollama. RTX 3060 12GB anchors at $250. buyer-guide
Can the RTX 5070 Run 34B Models in 2026? (Analyzed) Can the RTX 5070's 12GB VRAM handle 34B parameter LLMs in 2026? Honest analysis with quantization breakdowns and alternatives. guide
RunPod vs Vast.ai for LLM Inference in 2026 (Compared) RunPod vs Vast.ai compared for LLM inference. Pricing, reliability, GPU availability, and which cloud provider wins for your workflow. comparison
Ollama GPU Requirements 2026: VRAM for Every Model What VRAM each Ollama model actually needs, 1B to 70B. At Q4 an 8B is 4.9GB, a 13B 7.9GB and a 70B 43GB, with KV cache on top of that. guide
RTX 5090 vs RTX 4090 for LLM: 32GB vs 24GB in 2026 RTX 5090 vs RTX 4090 for local LLM inference in 2026. 32GB GDDR7 vs 24GB GDDR6X — is the extra VRAM worth the price premium? comparison
Best GPU for 34B Models: Yi, CodeLlama & Qwen Run 34B parameter models locally in 2026 — Yi-34B, CodeLlama, and Qwen 34B compared. VRAM needs, real-world speeds, and top GPU picks. buyer-guide
Best GPU for Ollama 2026: 4090 Speed vs 3090 Value The RTX 4090 is the fastest Ollama card, but a used RTX 3090 reaches 88% of its 13B speed on the same 24GB for a third of the price. buyer-guide
Best Multi-GPU LLM Setup 2026: Dual 3090s and Splitting Which multi-GPU LLM rig to build: dual 3090s for 48GB at about $1,640, how layer and row splitting differ, whether NVLink earns its price, and PSU sizing. guide
Best Used GPU for Local LLM in 2026 (3090 Top Pick) Top used GPUs for running local LLMs in 2026 on a budget — RTX 3090, 3080, and others. Pricing, VRAM, and what to avoid buying. buyer-guide
Best GPU for 13B Parameter Models in 2026 (Ranked) Top GPU picks for running Llama 13B, CodeLlama 13B, and other 13B LLMs locally in 2026 — with VRAM, tok/s, and budget tiers from $300. buyer-guide
Local LLM Under $1,000: Why 24GB Used Beats 16GB New The RTX 5070 Ti and 5080 both left this tier in 2026. Under $1,000 the choice is 24GB used against 16GB new, and only 24GB loads a 34B model. buyer-guide
Best GPU for Private AI in 2026 (5 Picks for Local) Top GPUs for running private, local AI inference with no cloud data sharing. Keep your prompts and data completely offline. buyer-guide
Best GPU for vLLM Serving in 2026 (5 Picks Ranked) Best GPU for vLLM inference serving. Covers PagedAttention, throughput benchmarks, and top GPU picks for production LLM deployment. buyer-guide
Best GPU for Local LLM Under $500 in 2026 (5 Picks) Top 5 budget GPUs under $500 for running local LLMs with Ollama and llama.cpp in 2026, ranked by VRAM, speed, and value at $/GB. buyer-guide
Cloud vs Local GPU for LLM: Real Cost Breakdown Cloud GPU vs buying your own for LLM in 2026 — RunPod, Vast.ai, and local costs compared. See the break-even point for your usage. guide
Best GPU for Code LLMs in 2026 (Qwen Coder, DeepSeek) Best GPU for running CodeLlama, DeepSeek Coder, and Qwen Coder locally in 2026 — 16GB for 14B, 24GB for 33B code models. buyer-guide
Best GPU for RAG Workloads in 2026 (Ranked Picks) Top GPUs for RAG in 2026 — embedding, vector search, and LLM inference compared. See which cards handle the full pipeline well. buyer-guide
Best GPU for LLM Inference Server in 2026 (vLLM) Top GPUs for serving LLMs to multiple users in 2026 with vLLM, TGI, and Ollama. Production inference server hardware guide. buyer-guide
ROCm vs CUDA for Local LLM 2026: Is AMD Usable Yet? The RX 7900 XTX's 960 GB/s sits between a used 3090 and a 4090, but its 24GB caps it at 34B. Where AMD is fine, and where it costs you a weekend. comparison
How to Choose a GPU for Ollama in 2026 (Step Guide) Step-by-step guide to picking the right GPU for Ollama in 2026 — match your model size, budget, and use case to the ideal card. guide
Best GPU for Llama 3 in 2026 (8B-70B Picks Ranked) Find the best GPU for running Llama 3 8B, 70B, and 405B locally. VRAM requirements, benchmarks, and top picks for every budget. buyer-guide
RTX 4090 vs RTX 3090 for LLM: New vs Used Value in 2026 RTX 4090 vs RTX 3090 head-to-head for local LLM inference in 2026. Same 24GB VRAM, very different performance and pricing. comparison
Best GPU for Mistral Models in 2026 (5 Picks Ranked) Mistral 7B is a 4.4GB download at Q4; Mixtral 8x7B wants 28GB. 5 GPUs ranked from RTX 3060 to 5090 for Mistral and Mixtral in 2026. buyer-guide
Best GPU for Qwen Models in 2026 (Qwen 3 + 3.6 Picks) Best GPU for running Qwen 2.5 models locally, from 0.5B to 72B. VRAM requirements, benchmarks, and top GPU picks by budget. buyer-guide

Coverage

What this site is designed to help you decide

VRAM

Model fit

Can your target model fit fully in GPU memory, or are you heading toward painful CPU offload?

TOK/S

Inference speed

We care about usable tokens-per-second, not just broad gaming or synthetic benchmark claims.

$

Budget discipline

Most readers do not need datacenter hardware. We bias toward practical consumer and used-market value.

DIY

Real workflows

Recommendations are framed around Ollama, llama.cpp, quantized models, and local workstation constraints.