Set it up, size it, ship it.

How-to Guides

How-to guides are the implementation companion to the buyer's guides. Once you've chosen hardware, these answer the operational questions: how much VRAM does this model actually use, which quantization is worth it, how do you wire two RTX 3090s without thermal-throttling, and what does each runtime do differently in 2026.

27guides in this category
2026refresh window

How this category works

We focus on the questions where the answer changes month-to-month: KV cache math after the latest llama.cpp release, Flash Attention 3 support across frameworks, MoE memory behaviour with Llama 4 and Gemma 4, and how MLX on Apple Silicon now compares to CUDA for the same model.

Every guide has been authored against a real reference setup we can actually run — Ollama on a single consumer GPU, llama.cpp on a dual-3090 rig, vLLM on a single H100 rental. When we cite a tokens-per-second number, the body of the article links to the underlying chart so you can audit the methodology rather than trust a headline number.

If you're new to local inference, the VRAM planning, runtime comparison, and quantization guides are the trio to read first — they remove ninety percent of the uncertainty before you spend on hardware.

Browse 27

Every article in How-to Guides

Filter or search the full list. Sorted newest-first by default.

guide Jul 26, 2026

Can You Run Kimi K3 Locally? No — Here's the Exact Math

Kimi K3's 2.8T open weights need roughly 1.4TB of GPU memory — no consumer rig comes close. The honest math, and what to run at home instead.

Read guide →
guide Jul 13, 2026

Rent a GPU for LLM Fine-Tuning: The $30 Weekend Project

Fine-tune a 7B-13B model on a rented A100 for roughly $20-40 a weekend instead of buying a $2,200 RTX 4090. When renting wins and when owning pays off.

Read guide →
guide Apr 21, 2026

How Much VRAM for Gemma 4? The 26B MoE Wants 18GB

Gemma 4 sizes do not match their names: the 26B-A4B is an 18GB download and the 12B Dense only 7.6GB. Published sizes for every variant, with GPU picks.

Read guide →
guide Apr 20, 2026

Best Quantization for Local LLM in 2026 (Q4 to Q8)

Q4_K_M vs Q5_K_M vs Q6_K vs Q8 in 2026 — which quantization gives the best quality-to-VRAM tradeoff for local LLM inference?

Read guide →
guide Apr 20, 2026

GPU Shortage 2026: Should You Buy Now for LLM?

GPU prices are surging in 2026 due to GDDR7 shortage and AI demand. Here's whether to buy now or wait for local LLM use.

Read guide →
guide Apr 19, 2026

Dual RTX 3090 for 70B LLMs: 48GB Build Guide 2026

Two used 3090s reach 48GB and run Llama 70B at Q4 around 18-22 tok/s. No NVLink needed: motherboard, PSU sizing, how to split, and real costs.

Read guide →
guide Apr 18, 2026

PSU for Dual GPU LLM: 1200W for 3090s, 1500W for 4090s

Dual RTX 3090s draw ~700W and want a 1200W PSU; dual 4090s need 1500W. Per-pair wattage table, rail requirements and what the minimum really is.

Read guide →
guide Apr 17, 2026

Llama 4 Maverick Hardware Guide (400B MoE) for 2026

What hardware do you need for Llama 4 Maverick 400B? Multi-GPU requirements, cloud options, and whether it's worth self-hosting.

Read guide →
guide Apr 14, 2026

RTX 3060 Relaunch: What Budget LLM Buyers Got

NVIDIA relaunched the RTX 3060 12GB at $329 in June 2026 and street prices have climbed since. What that means for budget local LLM builds.

Read guide →
guide Apr 12, 2026

Can the RTX 4060 Ti Run 13B Models in 2026? (Honest)

Can the RTX 4060 Ti run 13B models in 2026? 16GB version: yes (Q4-Q6). 8GB version: barely (Q3 only). Full VRAM breakdown.

Read guide →
guide Apr 11, 2026

Can a Mac Mini Run Local LLMs in 2026? M6 vs M5 Pro

The M6 Mac mini starts at 16GB and 153 GB/s; the M5 Pro brings 24GB and 307 GB/s. Which local models each one runs, and where each one stops.

Read guide →
guide Apr 11, 2026

Intel Arc B580 for Local LLM: Can Intel's Budget GPU Run Models?

Intel Arc B580 for LLM in 2026 — 12GB at $310. Runs 7B via llama.cpp Vulkan, but Ollama support is limited. Honest verdict here.

Read guide →
guide Apr 9, 2026

How Much VRAM for Llama 4 in 2026? Scout vs Maverick

Llama 4 Scout needs ~67GB at Q4 and Maverick ~245GB — far past any consumer GPU. The real numbers, why they surprise people, and what to run.

Read guide →
guide Apr 8, 2026

How Much VRAM for a 70B LLM in 2026? (Q4-Q8 Table)

Llama 3.3 70B ships as a 43GB Q4_K_M download — two 24GB cards hold it, one 32GB card does not. Published sizes for every quantization level.

Read guide →
guide Apr 8, 2026

How Much VRAM for Qwen 3 in 2026? Full Size Breakdown

VRAM for Qwen 3 in 2026 — 4B needs 2.5GB, 14B needs 9.3GB, 32B needs 20GB and the 30B-A3B MoE 19GB at Q4. Sizes from Ollama's own listing.

Read guide →
guide Apr 8, 2026

Used RTX 3090 for LLM: Inspection Checklist and Red Flags

A used RTX 3090 runs about $820 for 24GB. What to test before you pay, the red flags that mean walk away, and the ex-mining faults to expect.

Read guide →
guide Apr 6, 2026

Local LLM VRAM 2026: The 12GB Trap Most Buyers Hit

Most '16GB is enough' advice misses what breaks at 34B+. Full Q4-Q8 VRAM tiers + the budget mistake that costs you a year of upgrades.

Read guide →
guide Apr 5, 2026

How Much VRAM for Qwen 14B in 2026? (Q4-Q8 Guide)

Exact VRAM requirements for Qwen 2.5 14B and Qwen 3 14B in 2026 at every quantization level — Q4, Q5, Q6, Q8 — plus GPU picks.

Read guide →
guide Apr 3, 2026

Can the RTX 4060 Ti Run Llama 70B in 2026? (Honest)

Can a 16GB RTX 4060 Ti actually run Llama 70B in 2026? Honest answer with quantization analysis and practical alternatives.

Read guide →
guide Apr 3, 2026

How Much VRAM Do You Need for Llama 3 8B in 2026?

Exact VRAM requirements for Llama 3 8B at every quantization level, including context length overhead and practical GPU recommendations.

Read guide →
guide Apr 2, 2026

Can an RTX 3060 Run Ollama in 2026? (Honest Guide)

Which LLMs an RTX 3060 12GB runs through Ollama, which models to pull in 2026, and why the two new MoE releases need far more than its 12GB.

Read guide →
guide Apr 1, 2026

How to Run a 70B LLM on a Single GPU in 2026 (Q3-Q4)

Run Llama 3 70B and other 70B models on one GPU using aggressive quantization. VRAM requirements, quality trade-offs, and practical setup guide.

Read guide →
guide Mar 29, 2026

Can the RTX 5070 Run 34B Models in 2026? (Analyzed)

Can the RTX 5070's 12GB VRAM handle 34B parameter LLMs in 2026? Honest analysis with quantization breakdowns and alternatives.

Read guide →
guide Mar 28, 2026

Ollama GPU Requirements 2026: VRAM for Every Model

What VRAM each Ollama model actually needs, 1B to 70B. At Q4 an 8B is 4.9GB, a 13B 7.9GB and a 70B 43GB, with KV cache on top of that.

Read guide →
guide Mar 26, 2026

Best Multi-GPU LLM Setup 2026: Dual 3090s and Splitting

Which multi-GPU LLM rig to build: dual 3090s for 48GB at about $1,640, how layer and row splitting differ, whether NVLink earns its price, and PSU sizing.

Read guide →
guide Mar 21, 2026

Cloud vs Local GPU for LLM: Real Cost Breakdown

Cloud GPU vs buying your own for LLM in 2026 — RunPod, Vast.ai, and local costs compared. See the break-even point for your usage.

Read guide →
guide Mar 17, 2026

How to Choose a GPU for Ollama in 2026 (Step Guide)

Step-by-step guide to picking the right GPU for Ollama in 2026 — match your model size, budget, and use case to the ideal card.

Read guide →

Other lanes

Looking for something else?