These are my notes from running large language models on my own hardware for my coding agents and Hermes, first on one RTX 3090 and then on two joined by NVLink. Each part starts with a number that looked like it settled something, and ends with what I changed once I understood it.
- Part 1, diagnose: why the GPU showed 100% while my agents got slower.
- Part 2, decide: why a benchmark that made two cards look 21 times slower couldn’t tell me whether to buy one, and how I test now.
- Part 3, operate: where 22 GiB of idle GPU memory goes, and how it became a limit on requests in flight.
The rig: two RTX 3090s (24 GB each) with NVLink, a Threadripper 2950X, 62 GiB of RAM, Ubuntu 24.04, vLLM in Docker. Earlier tests used llama.cpp on one card and rented GPUs on Vast.ai.