<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Capacity-Planning on /sys/admin/blog</title><link>https://kre80r.com/tags/capacity-planning/</link><description>Recent content in Capacity-Planning on /sys/admin/blog</description><generator>Hugo</generator><language>en</language><lastBuildDate>Sat, 10 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://kre80r.com/tags/capacity-planning/feed.xml" rel="self" type="application/rss+xml"/><item><title>Room for two</title><link>https://kre80r.com/room-for-two/</link><pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate><guid>https://kre80r.com/room-for-two/</guid><description>&lt;p>With vLLM running and nobody connected, &lt;code>nvidia-smi&lt;/code> shows this:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-text" data-lang="text">&lt;span class="line">&lt;span class="cl">GPU 0: 22,112 MiB / 24,576 MiB used, 0% utilisation
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">GPU 1: 22,134 MiB / 24,576 MiB used, 0% utilisation
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>My first thought was a memory leak, but vLLM takes that memory on purpose. My
config sets &lt;code>--gpu-memory-utilization 0.90&lt;/code>, so when vLLM starts it claims 90%
of each GPU and manages that space itself. &amp;ldquo;Memory used&amp;rdquo; tells me nothing about
load.&lt;/p></description></item></channel></rss>