Knowledge library
Engineering decisions, implementation notes, and operational field guides.
1 entry
A practical guide to building a local AI inference server, from a single RTX 3060 or 3090 to a 96 GB multi-GPU system, with Ollama, LM Studio, Hugging Face models, and an OpenAI-compatible vLLM API.