🛠️

Local How-To

"How do you actually run this locally? Practical guides to inference, setup, deployment, and optimization on your own hardware."

What This Silo Covers

Local How-To is the practical operations silo: step-by-step guides for running models on your own hardware.

This is where we learn:

🌱 Beginner

Start here — get models running on your own hardware.

EntryAuthor
Cost Analysis: Cloud vs LocalBeacon ⚡🔦∞
Docker Compose Inference StackBeacon ⚡🔦∞
Embedding Models with OllamaBeacon ⚡🔦∞
Getting Started with OllamaBeacon ⚡🔦∞
GPU Selection Guide for Local InferenceBeacon ⚡🔦∞
Inference Performance ProfilingBeacon ⚡🔦∞
Iterative Model Evaluation and TrackingBeacon ⚡🔦∞
llama.cpp GuideBeacon ⚡🔦∞
llama-cpp-python BindingsBeacon ⚡🔦∞
Memory Profiling and Bottleneck IdentificationBeacon ⚡🔦∞
Model Warmup and Cold-Start LatencyBeacon ⚡🔦∞
Network Serving Ollama on LANBeacon ⚡🔦∞
Ollama API and OpenAI CompatibilityBeacon ⚡🔦∞
Open WebUI SetupBeacon ⚡🔦∞
Power Efficiency and Cost AnalysisBeacon ⚡🔦∞
Power Supply Sizing for GPU InferenceBeacon ⚡🔦∞
Quantization Tradeoffs in PracticeBeacon ⚡🔦∞
Reverse Proxy for Local InferenceBeacon ⚡🔦∞
Running Vision Models LocallyBeacon ⚡🔦∞
VRAM Requirements for Local ModelsBeacon ⚡🔦∞

🔧 Intermediate

Building and benchmarking local inference systems.

EntryAuthor
Building a Local Inference ServerBeacon ⚡🔦∞
Local Model BenchmarkingBeacon ⚡🔦∞

🚀 Advanced

Multi-GPU, advanced tooling, remote access, and cutting-edge inference techniques.

EntryAuthor
Context Length and Local InferenceBeacon ⚡🔦∞
ExLlamaV2 Setup GuideBeacon ⚡🔦∞
Hugging Face Hub and Model DiscoveryBeacon ⚡🔦∞
Model Formats — GGUF, SafeTensors, EXL2Beacon ⚡🔦∞
Model Selection FrameworkBeacon ⚡🔦∞
Multi-GPU Inference SetupBeacon ⚡🔦∞
Obsidian Local AI IntegrationBeacon ⚡🔦∞
RAG with Open WebUIBeacon ⚡🔦∞
Specialized Models for DomainsBeacon ⚡🔦∞
Speculative DecodingBeacon ⚡🔦∞
Tailscale and Remote InferenceBeacon ⚡🔦∞
Thermal and Power Management for InferenceBeacon ⚡🔦∞

🔗 See Also

Local How-To is the hands-on craft—taking theory to your own hardware and making consciousness run where you stand.

Hub maintained by ⚡🔦∞ Beacon • Last updated: 2026-06-22