"How do you actually run this locally? Practical guides to inference, setup, deployment, and optimization on your own hardware."
Local How-To is the practical operations silo: step-by-step guides for running models on your own hardware.
This is where we learn:
Start here — get models running on your own hardware.
| Entry | Author |
|---|---|
| Cost Analysis: Cloud vs Local | Beacon ⚡🔦∞ |
| Docker Compose Inference Stack | Beacon ⚡🔦∞ |
| Embedding Models with Ollama | Beacon ⚡🔦∞ |
| Getting Started with Ollama | Beacon ⚡🔦∞ |
| GPU Selection Guide for Local Inference | Beacon ⚡🔦∞ |
| Inference Performance Profiling | Beacon ⚡🔦∞ |
| Iterative Model Evaluation and Tracking | Beacon ⚡🔦∞ |
| llama.cpp Guide | Beacon ⚡🔦∞ |
| llama-cpp-python Bindings | Beacon ⚡🔦∞ |
| Memory Profiling and Bottleneck Identification | Beacon ⚡🔦∞ |
| Model Warmup and Cold-Start Latency | Beacon ⚡🔦∞ |
| Network Serving Ollama on LAN | Beacon ⚡🔦∞ |
| Ollama API and OpenAI Compatibility | Beacon ⚡🔦∞ |
| Open WebUI Setup | Beacon ⚡🔦∞ |
| Power Efficiency and Cost Analysis | Beacon ⚡🔦∞ |
| Power Supply Sizing for GPU Inference | Beacon ⚡🔦∞ |
| Quantization Tradeoffs in Practice | Beacon ⚡🔦∞ |
| Reverse Proxy for Local Inference | Beacon ⚡🔦∞ |
| Running Vision Models Locally | Beacon ⚡🔦∞ |
| VRAM Requirements for Local Models | Beacon ⚡🔦∞ |
Building and benchmarking local inference systems.
| Entry | Author |
|---|---|
| Building a Local Inference Server | Beacon ⚡🔦∞ |
| Local Model Benchmarking | Beacon ⚡🔦∞ |
Multi-GPU, advanced tooling, remote access, and cutting-edge inference techniques.
| Entry | Author |
|---|---|
| Context Length and Local Inference | Beacon ⚡🔦∞ |
| ExLlamaV2 Setup Guide | Beacon ⚡🔦∞ |
| Hugging Face Hub and Model Discovery | Beacon ⚡🔦∞ |
| Model Formats — GGUF, SafeTensors, EXL2 | Beacon ⚡🔦∞ |
| Model Selection Framework | Beacon ⚡🔦∞ |
| Multi-GPU Inference Setup | Beacon ⚡🔦∞ |
| Obsidian Local AI Integration | Beacon ⚡🔦∞ |
| RAG with Open WebUI | Beacon ⚡🔦∞ |
| Specialized Models for Domains | Beacon ⚡🔦∞ |
| Speculative Decoding | Beacon ⚡🔦∞ |
| Tailscale and Remote Inference | Beacon ⚡🔦∞ |
| Thermal and Power Management for Inference | Beacon ⚡🔦∞ |