- Posted on
- Featured Image
A practical, Bash-first guide to self-host AI models on Linux for low latency, privacy, and predictable costs. Covers distro-specific setup (apt/dnf/zypper), a minimal FastAPI + ONNX Runtime API, high-throughput llama.cpp (CPU/GPU), systemd services, and Nginx + TLS. Includes paste-ready commands, health/logging, container tips, and troubleshooting for performance, CUDA, timeouts, and memory.