Series

Bring your own GPU

Three field notes on learning to diagnose, benchmark, and operate local LLM inference under real agent workloads.

These are my notes from running large language models on my own hardware for my coding agents and Hermes, first on one RTX 3090 and then on two joined by NVLink. Each part starts with a number that looked like it settled something, and ends with what I changed once I understood it.

The rig: two RTX 3090s (24 GB each) with NVLink, a Threadripper 2950X, 62 GiB of RAM, Ubuntu 24.04, vLLM in Docker. Earlier tests used llama.cpp on one card and rented GPUs on Vast.ai.

Index

3 entries

Bring your own GPU · Part 3

Room for two

vLLM takes 90% of each GPU before anyone connects. Working out where it goes gave me a limit on requests in flight, and the proxy in front of vLLM now enforces it.