<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Llama.cpp on /sys/admin/blog</title><link>https://kre80r.com/tags/llama.cpp/</link><description>Recent content in Llama.cpp on /sys/admin/blog</description><generator>Hugo</generator><language>en</language><lastBuildDate>Tue, 02 Jun 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://kre80r.com/tags/llama.cpp/feed.xml" rel="self" type="application/rss+xml"/><item><title>Every token reads the whole model</title><link>https://kre80r.com/every-token-reads-the-whole-model/</link><pubDate>Tue, 02 Jun 2026 00:00:00 +0000</pubDate><guid>https://kre80r.com/every-token-reads-the-whole-model/</guid><description>&lt;p>For about a year I used local models the easy way, asking &lt;code>codellama:13b&lt;/code> in
Ollama a question now and then. Then I started using them for real work: coding
agents like Claude Code and Pi, and Hermes.
An agent sends big prompts with tool definitions and files in them, runs in a
loop, and carries the whole conversation in every request. The models worked,
but they were slow, and they got slower the longer a session went on.&lt;/p></description></item></channel></rss>