Bring your own GPU · Part 3
Room for two
vLLM takes 90% of each GPU before anyone connects. Working out where it goes gave me a limit on requests in flight, and the proxy in front of vLLM now enforces it.
Topic
Index
3 entriesBring your own GPU · Part 3
vLLM takes 90% of each GPU before anyone connects. Working out where it goes gave me a limit on requests in flight, and the proxy in front of vLLM now enforces it.
Bring your own GPU · Part 2
My first benchmark said two RTX 3090s were 21 times slower per token than one. A rerun looked fine, and I had changed too many things to know why. That changed how I test.
Bring your own GPU · Part 1
nvidia-smi showed my 3090 at 100% while my agents got slower. Timing prefill and decode separately showed a card that was mostly waiting on memory.