Topic

Inference

Index

3 entries

Bring your own GPU · Part 3

Room for two

vLLM takes 90% of each GPU before anyone connects. Working out where it goes gave me a limit on requests in flight, and the proxy in front of vLLM now enforces it.