Bring your own GPU · Part 2 of 3

Two cards, too many variables

My first benchmark said two RTX 3090s were 21 times slower per token than one. A rerun looked fine, and I had changed too many things to know why. That changed how I test.

A Victorian oil painting of a testing laboratory with two benches that are meant to match but don't, each holding a pair of RTX 3090s wired to a gauge. One bench is overloaded and its needle is in the red, the other is calm, and an engineer between them frowns at two sheets of paper he can't compare.

In April I rented a machine with two RTX 3090s on Vast.ai, to see what a second card would buy me before paying for one. On the first benchmark, two cards were 21 times slower per token than one. My write-up that day ended with “Don’t buy a 2nd 3090 for the local rig right now.”

A second run the same weekend looked fine. But I had changed so many things between the two runs that neither result could support the decision I wanted to make, and I still can’t say what caused the 21x. What I kept from that weekend was a better way of testing, and I used it for the choices that came after.

How this was written: I ran the tests and made the decisions described here, from my own notes and logs. AI tools helped me draft and edit the text and made the illustrations. I reviewed and edited the result.

Two cards on PCIe

vLLM split the model across the two cards with tensor parallelism (TP=2), where each card holds half of every layer. Before the next layer can start, the two cards swap their partial results. The rental had no NVLink, the bridge that connects two cards directly, so those swaps went over PCIe for every layer of every token, which adds synchronisation overhead to decode.

I ran Qwen3.6-27B in 4-bit with vllm bench serve: six 40K-token prompts, three in flight at a time, first on one card and then on both, on the same host.

SetupMedian time per output tokenPer-user decode
One 3090 (TP=1)18 msabout 55 tok/s
Two 3090s, no NVLink (TP=2)385 msabout 3 tok/s

Running it again

That test started all the prompts together, so three 40K-token prefills and their decodes were competing at the same moment. That does happen, when three agents send big repositories at once, but most of the time my sessions arrive one after another. One card handled the same burst fine. My working theory was that with two cards every prefill also pushes its results across PCIe for every layer, and those big transfers held up everyone’s decode. I didn’t measure it.

So I ran it again with arrivals spread out: 8K-token prompts arriving about every 20 seconds, at most three in flight. Two cards came in at 24 ms median time per output token, about 41 tokens a second per user.

I had changed far more than the arrival pattern, though. The second run was on a different rented host with PCIe Gen 4 instead of Gen 3, the prompts were 8K instead of 40K, it used a stable vLLM release instead of a nightly build, and I didn’t run a single card on that host for comparison. I only noticed the PCIe difference when I wrote it up. I can’t say how much of the drop from 385 ms came from the arrival pattern and how much from the rest. I’ve done enough load testing on web stacks to know you change one thing at a time, and I still didn’t do it here.

Testing differently

That weekend gave me a few rules for any test that decides something. The workload is mine: long agent prompts. One thing changes at a time. What counts as a win is written down before the run. And every candidate runs between two control runs, with the raw results kept.

Buying for capacity

The “don’t buy” line came from a test of per-token speed over PCIe without NVLink, and per-token speed wasn’t why I wanted a second card. I wanted KV-cache room, for longer contexts and more agents at once. My April write-up had also said that if I kept running out of context, two 3090s with an NVLink bridge were the realistic upgrade, and that my card, an MSI Gaming X Trio, has the NVLink connector. Running out of context was where part 1 ended. So the purchase was a capacity decision, which the rental never tested.

The rig now has two 3090s joined by an NVLink bridge. nvidia-smi topo -m shows the two cards connected over NVLink, and they copy data to each other at about 26 GB/s. With one card, vLLM gave me three sessions at 48K, or one at about 72K. With two, the Qwen3.8 builds below run at 131K with room for about four full-length requests, or at 262K with room for two. I haven’t compared one card against two at home.

Picking a quantization

In August I had to choose which quantized build of Qwen3.8-27B to run for my agents, and I tested it the new way. Each model card quoted its own benchmarks on its own workloads. Speeds reported by club-3090, a good community project for 3090s, mostly came from 1,024-token prompts with 512-token answers. My agents usually send 30K to 60K-token prompts and sometimes go past 100K.

I tested three builds on the same rig with the same settings (131,072-token context, FP8 KV cache), the same long-prompt runs and the same 12-task quality suite, run once each, so the quality column is only a sanity check.

WeightsQualityTime to first token, one 120K promptDecode, one streamKV cache poolPool ÷ 131,072
AutoRound, 4-bit weights, 8-bit activations109/12058.64 s58.56 tok/s541,764 tokens4.13
AWQ, 4-bit weights, 16-bit activations107/12081.56 s60.90 tok/s610,212 tokens4.66
FP8107/12080.20 s43.24 tok/s272,338 tokens2.08

FP8 dropped out first. Its weights are twice the size of the 4-bit ones, its decode was the slowest, which is consistent with part 1, and it left about half the room for the KV cache.

The two 4-bit builds decoded at nearly the same speed, but AutoRound’s prefill was much faster. The AutoRound build runs 8-bit activations, which makes the arithmetic cheaper, and prefill is where the arithmetic is. To check that one variable, I ran another checkpoint that supports both paths, with the same weights. The 8-bit activation path cut time to first token from 81.80 to 58.70 seconds, about 28% faster, and decode went from 61.30 to 57.92 tokens a second, about 5% slower.

My agents send long prompts all day, and the 20-odd seconds saved on every big prefill mattered more to me than AWQ’s slightly faster decode and bigger cache. I picked AutoRound. For a chat workload with short prompts and long answers I’d probably have picked AWQ.

Tuning against a gate

For the settings I tested around the same time, each gate was written down before the run.

  • Raising the scheduler’s sequence limit from 4 to 8 lifted aggregate throughput at eight concurrent short requests from 208.1 to 352.8 tokens a second. Promoted.
  • Async scheduling raised aggregate throughput at six concurrent short requests from 166.8 to 210.1 tokens a second and cut median time to first token from 1.55 to 1.16 seconds. With long prompts it made no difference. Promoted.
  • Halving the scheduler’s batch token budget from 8,192 to 4,096 improved median per-stream speed by 1.44%, against a gate of 5%, and the worst single stream at five concurrent sessions got 5.32% slower. Rejected.
  • Speculative decoding with the model’s built-in multi-token prediction (MTP) made decode slower both times I tried it. At depth 4, 43.7% of draft tokens were accepted. On another checkpoint at depth 1, 70.8% were accepted and decode still fell from 67.64 to 51.53 tokens a second, with the scheduler limited to one sequence. I didn’t work out why. The vLLM docs warn that “every rejected token is compute wasted; with enough of them, throughput drops.” Off.

I still don’t know what caused the 21x. What changed after that weekend is how I test. AutoRound and async scheduling passed under those rules, and the smaller batch budget and MTP failed them.

Part 3 is about where the GPU memory goes, and the morning I admitted more requests than it could hold.

Next in the series: Part 3, Room for two