Quick answer: The RTX 3090 at ~$800-900 used is the best value GPU for local LLM inference. Its 24GB VRAM and 936 GB/s bandwidth handle models up to 34B quantized, and nothing else comes close at that price.
NVIDIA GeForce RTX 3090
24GB GDDR6XKing of the used GPU market — 24GB VRAM at ~$850 runs CodeLlama 34B and Qwen 2.5 32B that 16GB cards can't touch.
Affiliate links — we may earn a commission at no extra cost to you. Amazon ships globally; Shopee SG covers Singapore & ASEAN.
Why buy used for LLM work?
LLM inference does not stress GPUs the way mining or gaming does. Inference is mostly memory reads, not sustained compute at max temperature. A used card that spent years gaming or even mining still has plenty of life for LLM workloads.
The key advantage: VRAM-per-dollar on the used market is dramatically better than new. The RTX 3090 offers 24GB for under $900 — the cheapest new card with 24GB (RTX 4090) costs $2,200.
Best used GPUs ranked
| GPU | VRAM | Bandwidth | Tok/s (7B Q4) | Used Price | VRAM/$ |
|---|---|---|---|---|---|
| RTX 3090 | 24GB GDDR6X | 936 GB/s | ~65 tok/s | ~$800-900 | 27-30 GB/k$ |
| RTX 3090 Ti | 24GB GDDR6X | 1,008 GB/s | ~70 tok/s | ~$900-1,000 | 24-27 GB/k$ |
| RTX 3080 12GB | 12GB GDDR6X | 912 GB/s | ~45 tok/s | ~$400-450 | 27-30 GB/k$ |
| RTX 3080 10GB | 10GB GDDR6X | 760 GB/s | ~40 tok/s | ~$350-400 | 25-29 GB/k$ |
| RTX 3060 12GB | 12GB GDDR6 | 360 GB/s | ~30 tok/s | ~$250 | 48 GB/k$ |
| RTX 3070 Ti | 8GB GDDR6X | 608 GB/s | ~38 tok/s | ~$280-320 | 25-29 GB/k$ |
| RTX 4090 (used) | 24GB GDDR6X | 1,008 GB/s | ~65 tok/s | ~$1,300-1,400 | 17-18 GB/k$ |
VRAM capacity memory bandwidth Specs are manufacturer figures. Bar lengths are scaled independently per metric.
RTX 3090 — the king of used LLM GPUs
The RTX 3090 dominates the used market for one reason: 24GB VRAM at ~$850.
What you can run:
- Llama 3 8B — FP16 full precision, blazing fast
- Mistral 7B / Gemma 2 9B — any quantization level
- Llama 2 13B — Q8 quantization, near-lossless
- CodeLlama 34B — Q4_K_M, usable quality
- Qwen 2.5 32B — Q4_K_M, fits well
What you can’t run:
- 70B models at any decent quality (need 48GB+)
- Multiple large models simultaneously
The 936 GB/s memory bandwidth is still competitive with current mid-range cards. For pure LLM inference, a used RTX 3090 beats a new RTX 5080 on model capacity despite being two generations older. If you are seriously considering a 3090 purchase, our used RTX 3090 buying guide for LLM covers exactly what to inspect, which sellers to trust, and the common failure modes to test for. For a head-to-head Ollama speed comparison between the 3090 and the newer 4090, see RTX 4090 vs 3090 for Ollama.
RTX 3060 12GB — the ultra-budget champion
At around $250, the RTX 3060 12GB is the cheapest way into local LLM:
- 12GB VRAM runs all 7B models and 13B at Q4
- Huge community support — this was the default budget LLM card for years
- Low power draw (170W) means no PSU upgrade needed
- Slower bandwidth (360 GB/s) is the main limitation
If your budget is under $300, this is the only used card worth considering. The 12GB VRAM matters far more than the slower inference speed. See our budget GPU guide for more on this card.
RTX 3080 12GB — the overlooked middle ground
The 12GB variant of the RTX 3080 is often forgotten:
- 12GB VRAM matches the RTX 3060 but with 2.5x the bandwidth
- 912 GB/s bandwidth means noticeably faster inference
- Used prices around $400-450 make it solid value
- Handles 7B models at full speed and 13B at Q4-Q6 comfortably
If you can afford $400 but not $850, this card fills the gap between the 3060 and 3090 nicely.
What to watch for when buying used
| Risk | How to check |
|---|---|
| Dead VRAM chips | Run nvidia-smi and a VRAM stress test immediately |
| Thermal damage | Check for thermal pad residue, warped PCB |
| Mining wear | Ask for usage history; mining cards often have replaced thermal pads |
| BIOS mods | Verify stock BIOS with GPU-Z before paying |
| Warranty | Most used cards have no warranty — factor this into pricing |
A used mining card is not necessarily bad. Miners typically ran cards at stable temperatures with undervolts. Gaming cards that thermal-cycled constantly may actually have more wear.
Used vs new: when to buy new instead
Buy new if:
- You want warranty coverage
- You need 32GB VRAM (RTX 5090 is the only consumer option)
- You value power efficiency (30-series draws significantly more power)
- Your PSU can’t handle 350W cards
Buy used if:
- Budget is your primary constraint
- You need 24GB VRAM without spending $2,200
- You’re comfortable with basic hardware testing
- You have adequate PSU and cooling
NVIDIA GeForce RTX 4090
24GB GDDR6XBest new option for used GPU shoppers who want warranty and longevity — 24GB at 1,008 GB/s with current driver support.
Affiliate links — we may earn a commission at no extra cost to you. Amazon ships globally; Shopee SG covers Singapore & ASEAN.
Which used GPU should you buy?
- Want to run 34B models and have $800-900? Get the RTX 3090. Its 24GB VRAM is unmatched at this price and handles CodeLlama 34B, Qwen 2.5 32B, and every 7B/13B model at any quantization.
- Budget under $300? Get the RTX 3060 12GB (~$250). The 12GB VRAM runs all 7B models and gets you into local LLM for less than three months of API costs.
- Want a middle ground at $400-450? Get the RTX 3080 12GB. Same 12GB VRAM as the 3060 but with 2.5x the memory bandwidth for noticeably faster token generation.
- Already saving for a new RTX 5090? Buy a used RTX 3090 now. It becomes an excellent second GPU later, giving you 56GB combined VRAM alongside the 5090. For a direct comparison of what the 5090 gains over the 3090 across real LLM workloads, see RTX 5090 vs 3090 for LLM.
Common mistakes to avoid
- Buying a used RTX 3070 Ti for LLM work. Its 8GB VRAM is too restrictive for anything beyond 7B models. Spend less on a 3060 12GB (more VRAM) or more on a 3090 (far more capable).
- Skipping hardware testing before committing to a purchase. Run
nvidia-smiand a VRAM stress test within the return window. Dead VRAM chips are the most common used GPU failure mode. - Paying over $1,000 for a used RTX 3090. At that price, you are close to RTX 4090 territory on sale. Set a hard ceiling of $900 and be patient for the right listing.
- Avoiding ex-mining cards on principle. Mining cards ran at stable temperatures with undervolts. A heavily thermal-cycled gaming card may actually have more wear. Judge by condition, not history.
Our recommendation
For most people entering the local LLM space on a budget, the RTX 3090 at ~$850 used is the single best hardware purchase you can make. It handles the vast majority of models people actually want to run, and the 24GB VRAM gives you room to grow as you explore larger models.
If that’s too much, start with an RTX 3060 12GB and upgrade when you hit the VRAM wall — our can the RTX 3060 run Ollama guide shows what that card handles and where it starts to struggle. Considering the newer RTX 5070 Ti as an alternative to a used 3090? See our RTX 5070 Ti vs 3090 for LLM comparison for the VRAM-vs-speed tradeoff. Check our VRAM guide to plan your upgrade path.
With the 2026 GPU shortage tightening supply, used GPU prices may climb further — buying now while availability is good is a smart move.
The used GPU market is the best-kept secret in local LLM. Last generation’s flagship is this generation’s best value.