Set it up, size it, ship it.
How-to Guides
How-to guides are the implementation companion to the buyer's guides. Once you've chosen hardware, these answer the operational questions: how much VRAM does this model actually use, which quantization is worth it, how do you wire two RTX 3090s without thermal-throttling, and what does each runtime do differently in 2026.
How this category works
We focus on the questions where the answer changes month-to-month: KV cache math after the latest llama.cpp release, Flash Attention 3 support across frameworks, MoE memory behaviour with Llama 4 and Gemma 4, and how MLX on Apple Silicon now compares to CUDA for the same model.
Every guide has been authored against a real reference setup we can actually run — Ollama on a single consumer GPU, llama.cpp on a dual-3090 rig, vLLM on a single H100 rental. When we cite a tokens-per-second number, the body of the article links to the underlying chart so you can audit the methodology rather than trust a headline number.
If you're new to local inference, the VRAM planning, runtime comparison, and quantization guides are the trio to read first — they remove ninety percent of the uncertainty before you spend on hardware.
Featured
Where most readers start
The 5 highest-leverage guides in this category — the ones that answer the most common reader questions.
- 01
- 02
- 03
- 04
- 05
Browse 27
Every article in How-to Guides
Filter or search the full list. Sorted newest-first by default.
Other lanes