- Posted on
- Featured Image
AI inference is spiky and cache-sensitive, so tail latency and GPU cost soar without smart load balancing. This Bash-first guide details tuning HAProxy/NGINX with consistent hashing by model/tenant, timeouts, queue and concurrency caps, 503/Retry-After backpressure, GPU-aware dynamic weights, and OS/network limits, plus SLO baselining and ab/ghz tests—cutting p99 and smoothing throughput without more GPUs.