<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Bring your own GPU on /sys/admin/blog</title><link>https://kre80r.com/series/bring-your-own-gpu/</link><description>Recent content in Bring your own GPU on /sys/admin/blog</description><generator>Hugo</generator><language>en</language><lastBuildDate>Sat, 10 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://kre80r.com/series/bring-your-own-gpu/feed.xml" rel="self" type="application/rss+xml"/><item><title>Room for two</title><link>https://kre80r.com/room-for-two/</link><pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate><guid>https://kre80r.com/room-for-two/</guid><description>&lt;p>With vLLM running and nobody connected, &lt;code>nvidia-smi&lt;/code> shows this:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-text" data-lang="text">&lt;span class="line">&lt;span class="cl">GPU 0: 22,112 MiB / 24,576 MiB used, 0% utilisation
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">GPU 1: 22,134 MiB / 24,576 MiB used, 0% utilisation
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>My first thought was a memory leak, but vLLM takes that memory on purpose. My
config sets &lt;code>--gpu-memory-utilization 0.90&lt;/code>, so when vLLM starts it claims 90%
of each GPU and manages that space itself. &amp;ldquo;Memory used&amp;rdquo; tells me nothing about
load.&lt;/p></description></item><item><title>Two cards, too many variables</title><link>https://kre80r.com/two-cards-too-many-variables/</link><pubDate>Mon, 31 Aug 2026 00:00:00 +0000</pubDate><guid>https://kre80r.com/two-cards-too-many-variables/</guid><description>&lt;p>In April I rented a machine with two RTX 3090s on Vast.ai, to see what a second
card would buy me before paying for one. On the first benchmark, two cards were
21 times slower per token than one. My write-up that day ended with &amp;ldquo;Don&amp;rsquo;t buy a
2nd 3090 for the local rig right now.&amp;rdquo;&lt;/p></description></item><item><title>Every token reads the whole model</title><link>https://kre80r.com/every-token-reads-the-whole-model/</link><pubDate>Tue, 02 Jun 2026 00:00:00 +0000</pubDate><guid>https://kre80r.com/every-token-reads-the-whole-model/</guid><description>&lt;p>For about a year I used local models the easy way, asking &lt;code>codellama:13b&lt;/code> in
Ollama a question now and then. Then I started using them for real work: coding
agents like Claude Code and Pi, and Hermes.
An agent sends big prompts with tool definitions and files in them, runs in a
loop, and carries the whole conversation in every request. The models worked,
but they were slow, and they got slower the longer a session went on.&lt;/p></description></item></channel></rss>