Home / Guides

Guides

Guides

Practical notes for local AI hardware. Published on this site only.

guide

Why Does Aggressive Quantization Hurt Output Quality?

Quantization can make a model fit in less VRAM, but lower-bit formats introduce more numerical error. Learn how to confirm quantization is the problem, apply no-cost fixes, and decide when a higher tier is worth the hardware.

guide

Why Is a Quantized Model Slower Than Expected?

Quantization reduces model size, but it does not guarantee faster inference. Learn how backend selection, GPU offload, memory bandwidth, context length, and thermal limits affect real-world speed.

guide

WSL2 vs Native Linux for Local AI

WSL2 is enough for many Windows-based local AI workflows, especially GPU inference, development, and Dockerized tools. Native Linux becomes the safer choice when you need maximum hardware control, predictable server operation, specialized accelerators, or heavy storage and networking workloads.

guide

ZFS for AI Workstations and Servers: Is It Worth It?

ZFS is worth considering when your AI system stores valuable datasets, model checkpoints, or shared project data and you need integrity checks, snapshots, and predictable recovery. For a single-user workstation with replaceable data, simpler filesystems may provide most of the benefit with less operational overhead.

guide

How to Back Up a Local AI Workstation

Protect the work that is difficult to recreate—code, datasets, fine-tunes, prompts, and configuration—without filling your backup storage with downloadable model files. This guide covers practical backup targets, desktop and server setups, recovery testing, and hardware considerations.

guide

Performance per Watt for Local AI: Why Efficiency Matters

The fastest GPU is not always the cheapest to operate. This guide shows how to compare local AI hardware using energy per task, total cost of ownership, cooling, maintenance, and resale value.

guide

Best Hardware for Local Reranking

Local reranking usually does not require a dedicated GPU. A modern CPU can handle small workloads, while a GPU becomes valuable when you need lower latency, higher throughput, or larger batches.

guide

Optimizer Memory: Why Training Uses So Much More VRAM

A model that fits comfortably in VRAM for inference may need several times more memory to train because training stores gradients, optimizer states, and activations. This guide shows where that memory goes and which hardware and software levers can make training feasible.

guide

Dataset Size vs Hardware for Fine-Tuning

Dataset size usually affects storage, preprocessing, training time, and checkpointing more than it affects GPU VRAM. This guide separates those requirements and shows how to choose hardware for full fine-tuning or adapter-based tuning.

guide

Continuous Batching Explained for Local LLM Servers

Continuous batching keeps a local LLM server processing new and active requests together instead of waiting for a fixed batch to finish. This guide explains how it improves throughput, what it costs in VRAM and latency, and how to tune it safely.

guide

How to Build a 128GB RAM Local AI Workstation

A 128GB RAM workstation gives local AI users room for larger quantized models, CPU inference, GPU offload, and development tools without relying on an enterprise platform. This guide explains how to choose compatible memory and plan a reliable upgrade or new build.

guide

How to Plan a 4-GPU Local AI Workstation

A four-GPU workstation is mainly a problem of PCIe lanes, physical spacing, power delivery, cooling, and software topology—not simply buying four graphics cards. This guide shows how to design a stable system and verify whether four GPUs will actually solve your workload.

guide

Best Quantization Strategy for Coding Models

Choose the lowest quantization that leaves enough VRAM for your context window and runtime overhead without harming the coding behavior you need. This guide compares practical choices for small and large local coding models.

guide

Does Model File Size Equal VRAM Usage?

A model file that is smaller than your GPU’s VRAM may still fail to load or run. Learn how weights, quantization metadata, KV cache, and runtime buffers determine the real memory requirement.

guide

Why Does Long Context Cause OOM?

Long conversations can exhaust VRAM even when the model and short prompts fit comfortably. This guide explains KV-cache growth, temporary attention memory, batching, and the configuration changes that usually fix context-related OOM errors.

guide

Should You Store AI Models on a NAS?

A NAS can be an excellent central repository for AI models, but it is not automatically a good place to run them from. This guide explains when network storage is fast enough, when local SSD storage is better, and how to configure a reliable hybrid setup.

guide

1GbE vs 2.5GbE vs 10GbE for AI Workstations

Choose between 1GbE, 2.5GbE, and 10GbE by matching network capacity to your storage, dataset, and multi-machine AI workflow. Includes practical desktop and server configurations.

guide

Five-Year TCO of a Local AI Workstation

A local AI workstation costs more than its purchase price. This guide shows how to model electricity, cooling, upgrades, failures, maintenance, and resale over five years.

guide

The Economics of a Used RTX 3090 for Local AI

A used RTX 3090 can be an excellent local-AI value because its 24GB of VRAM may enable workloads that smaller GPUs cannot run. Its advantage depends on utilization, electricity rates, failure risk, cooling, performance, and resale value—not just the sticker price.

guide

How to Build a Private Local AI Assistant

Build a local AI assistant by matching model size, memory, compute, storage, and latency to the tasks you actually need. This guide covers practical hardware tiers, software runtimes, and privacy considerations.

guide

LLM Capacity Planning: Requests, Context, Batch and VRAM

Estimate the GPU memory, concurrency, and throughput an LLM inference service needs by connecting request rate, context length, batching, and KV-cache growth. Includes a practical sizing method and failure-mode checklist.

guide

Why Tokens per Second Is Not Enough to Compare AI Servers

Tokens per second is useful, but it does not describe responsiveness, concurrency, context capacity, or user experience. This guide shows how to compare AI servers using throughput, latency, memory, and workload fit together.

guide

Are Used Enterprise Workstations Good for Local AI?

Used enterprise workstations can be an excellent low-cost foundation for local AI, but only when the chassis, power supply, PCIe layout, cooling, and GPU compatibility match your upgrade plan. This guide shows how to evaluate one before buying.

guide

How to Choose a Motherboard for Local AI

Choose a motherboard around your GPU count, PCIe lane requirements, memory capacity, and upgrade plans—not just the CPU socket. This guide explains the practical trade-offs for single-GPU and multi-GPU local AI systems.

guide

Motherboard Slot Spacing for Multi-GPU AI PCs

Multiple GPUs are only useful if the motherboard, chassis, power system, and cooling layout can physically support them. This guide shows how to check slot spacing before buying hardware.

guide

Multi-GPU Local AI: A Beginner's Guide

Multi-GPU systems can combine several graphics cards to run models that exceed one card's VRAM, but the result depends on sharding support, PCIe connectivity, power, and cooling. This guide explains when multiple GPUs help and how to plan a workable system.

guide

Used vs New GPU for Local AI: Which Is Better Value?

A used high-VRAM GPU can deliver excellent local AI value, but only when its condition, warranty, power requirements, and software support are acceptable. This guide shows how to compare its real risk and total cost against a new card.

guide

How Much VRAM Does a 128K Context Window Use?

A 128K context window can consume anywhere from several gigabytes to dozens of gigabytes of VRAM just for the KV cache. The exact cost depends more on layer count, KV-head count, head dimension, cache precision, and batch size than on parameter count alone.

guide

How Much VRAM Does a 32K Context Window Use?

A 32K context window can require anywhere from a few to many gigabytes of VRAM for the KV cache alone. Use the model’s layer count, KV-head count, head size, cache precision, and batch size to estimate whether it fits.

guide

VRAM vs System RAM for Local LLMs

VRAM determines how much of an LLM can run quickly on the GPU, while system RAM provides capacity for CPU inference, offloading, and supporting workloads. Learn when a bigger GPU is the better upgrade—and when more RAM can rescue a too-large model.

guide

What AI Models Can You Run on a 24GB GPU?

A 24GB GPU can run most 7B–14B models comfortably and many 30B–34B models in 4-bit quantization, but context length and runtime overhead determine the real limit. This guide explains what fits, when you need offload, and how to check a specific model before downloading it.

guide

What AI Models Can You Run on a 8GB GPU?

An 8GB GPU can run many 1B–8B language models locally, especially in 4-bit quantization, but context length and runtime overhead matter as much as model size. This guide explains what fits in VRAM, when offload is required, and how to check a specific model before downloading it.

guide

How Much VRAM Do I Need for Local AI?

Estimate the GPU memory needed for local AI by accounting for model weights, quantization, context length, KV cache, and runtime overhead. Use a practical workflow before choosing a GPU.