TrenchOps 🐎

Insights from the tech trenches

Linux

Cover Image

Your AI Agent Reads 50,000 Tokens to Learn One Thing

Compression as via negativa: the cheapest token is the one you never send

Take any coding agent session and open the raw request log. Not the pretty transcript - the actual payload. What you find is a JSON search result with 100 hits where 3 mattered, a log file where the interesting line is buried in 800 lines of "INFO: still fine," and a directory listing of a repo the model already walked twice.

Then apply the uncomfortable ratio: cost of the finished output versus the raw inputs actually required to produce it. The model needed maybe 2,000 tokens of real signal. You paid for 55,000. In manufacturing terms, that is a part made of solid gold to hold a plastic clip in place.

I spent a few weeks running everything - coding, research, boring text work - through a compression proxy called Headroom, and the interesting part was not the money. It was watching exactly how much of what I "sent" to the model was never information in the first place.

What it actually does

It sits between your agent and the API and squeezes everything the model reads - tool outputs, logs, retrieved chunks, conversation history - before it goes upstream. Compression happens locally, nothing gets shipped off to a third party to be shrunk. The model can pull the original back if it needs it.

The project's own benchmarks show roughly 20% on code search, 40-60% on incident debugging and codebase exploration, and the big numbers (86%+) on JSON arrays, which makes sense: JSON is mostly punctuation and repeated key names, a format designed for parsers and billed as if it were prose. Structured logs compress hard too. Source code mostly passes through untouched, which is the right default - you do not want your agent reasoning about a lossy version of the file it is editing.

Honest caveat, because I would rather you trust me in six months: the dashboard counts tokens it compressed, not tokens your provider stopped charging you for. Those are related but not identical numbers, and some users report the delta being smaller than advertised once retrieval round-trips are included. Treat the dashboard as a direction indicator, not an invoice.

Which brings up the part nobody markets: if you are on a flat subscription plan, you save exactly zero euros/dollars. What you save is runway. Fewer tokens consumed means you travel further before slamming into the 5-hour and weekly usage windows. For anyone who has had a long refactor cut off mid-thought by a rate limit, that is worth more than a discount.

Setup on Linux, in the time it takes to make coffee

Install and verify:

uv tool install --python 3.13 "headroom-ai[all]"
command -v headroom
headroom doctor

You want "Proxy" showing a green checkmark. If it does not, stop here and fix that first - do not proceed on hope.

Optional but sensible, run it as a user daemon so it is simply always there:

# ~/.config/systemd/user/headroom.service
[Unit]
Description=Headroom compression proxy

[Service]
ExecStart=%h/.local/bin/headroom proxy --port 8787
Restart=on-failure

[Install]
WantedBy=default.target

systemctl --user enable --now headroom
ss -lntp | grep 8787

That second command is not cosmetic. Confirm it is bound to 127.0.0.1 and not 0.0.0.0. An unauthenticated proxy that sees every prompt you write, listening on all interfaces, is not a productivity tool. It is a gift to whoever else is on your network.

Then point your client at it, in ~/.claude/settings.json:

{
Β "env": {
Β "ANTHROPIC_BASE_URL": "http://127.0.0.1:8787"
Β }
}

Restart the CLI and your IDE - VS Code or VSCodium will not pick up the new base URL otherwise. Open http://127.0.0.1:8787/dashboard and check that "Request Health" and "Live Activity" show non-zero numbers. That is the whole installation. The rest is watching a counter go up while you do the work you were going to do anyway.

When it breaks, and it will

It is a young, fast-moving project. Rarely, requests hang or die. The triage order, learned the boring way:

πŸ‘‰ Stop the current model session, restart the tool or IDE, tell the agent "continue." Fixes most of it.

πŸ‘‰ Still failing? systemctl --user restart headroom.

πŸ‘‰ Still failing? It is probably not your machine. Check status.claude.com before you debug anything else - degraded upstream service looks exactly like a broken proxy from where you are sitting, and I have wasted honest minutes proving that.

πŸ‘‰ Genuinely a regression in a new release? Pin backwards: uv tool install --python 3.13 "headroom-ai[all]==0.37.0" --force. Fast-moving projects reward people who know how to step back one version instead of filing an issue and waiting.

That last point is the real skill, by the way. Not the tool - the habit of having a rollback path before you need one.

The part that generalizes

Compression proxies are a workaround. The underlying problem is that we hand agents firehoses and call it context. We pipe in complete API responses, full log files and entire directory trees because it is easier than deciding what matters, then pay per token for the privilege of making a language model do our filtering at premium rates.

The best part is no part. The cheapest token is the one you never send. A tool that compresses your junk is strictly better than not having it - but a grep with a sane filter, a log level that is not DEBUG in production, and an API response that returns fields instead of everything would have gotten you most of the way there without any middleware at all.

πŸ‘‡ What is the dumbest thing your agent has ever been made to read in full? Mine spent real money ingesting a 400-line INFO log to find one timestamp.

Cover Image

The Opinion I Cached in 2003 Cost Me Ten Years of Better Tools

Why I moved back to SUSE after two decades - and why you should re-evaluate your stack while it still works

Somewhere around the year 2000, I bought a computer magazine because it had a CD glued to the cover. SUSE Linux 6.4. ReiserFS. I got as far as "select a mount point for root" and stopped, because I had no idea what a mount point was, and no realistic way to find out. No search engine within reach. No forum. Just me, a beige tower, and a dialog box asking me a question in a language I didn't speak.

I abandoned the installation. My first contact with Linux ended in surrender at the partitioning screen, which I suspect is the most common origin story in our industry.

A few months later, SUSE 7 came out and I ordered the boxed set. Seven CDs. A printed manual thick enough to stop a small projectile. A t-shirt. A sticker of a chameleon. And most importantly: no license key scrawled on a disc in marker by someone's cousin. I installed an operating system that nobody asked me to prove I deserved. That smell of freedom stayed with me longer than the T-shirt did.

Then came the wandering years.

Fedora: lovely, modern, and reliably broken at the worst possible moment. CentOS: rock solid, with a software repository that felt like a museum with a gift shop. So I did what everyone did and joined the Ubuntu hype. And Ubuntu worked. Mostly. For a while.

Every single installation, though, followed the same arc. Months of peace, then the onset of symptoms that cannot be described, narrowed down, or reproduced. A print dialog that takes eleven seconds. Bluetooth that pairs only on Tuesdays. Suspend that works until you look at it. Nothing you can file a bug for, because the bug report would read: "it feels sticky."

And underneath it all, APT. I have donated entire evenings of my finite lifetime to dependency conflicts that resolved themselves into a suggestion to remove approximately one third of my system. RPM and its resolvers were always the better engineering. I knew that. I kept using DEB anyway, because switching costs money you can only pay in weekends.

That is the interesting part of this story. Not the distro. The reason I stayed.

Cache invalidation is the hardest problem in computer science, including in your head

A while ago, I upgraded Ubuntu for the second or third time and did an honest inventory of everything I had learned to ignore. It was a long list. Each item individually below the threshold of action, collectively a system I was working around instead of working with. Time to reinstall.

But instead of reinstalling the same thing, I decided to actually survey the field. Mint. Rocky. A few others. Nothing clicked. So I did something slightly humiliating: I wrote down my real requirements - stability, sane package management, rollback, no drama - and handed them to an AI.

It came back with openSUSE.

My immediate reaction was physical discomfort, and the reason had nothing to do with 2025-era openSUSE. It was KDE. I was present during the era when a major desktop shipped to end users in a state its own developers described as a preview, and I mocked it with the enthusiasm of a man who had found his calling. I had also spent a decade in Xfce and before that loved an old GNOME 3. My verdict on KDE was formed, filed, and never reopened.

Here is the thing: that verdict was based on data from roughly two Olympic cycles ago, written to memory with no expiry date. My opinion had no TTL. It was a cached value being served with total confidence long after the origin server had rewritten the entire application.

Physicists have a word for this. Hysteresis: the state of a system depends on its history, not just on current conditions. Remove the force and the material does not spring back - it keeps a memory of the field it was in. That is exactly what technology opinions do inside senior engineers. We are magnetized by one bad release, one vendor who lied in 2014, one framework that burned us, one candidate who fumbled one interview question, and we stay magnetized long after the field is gone.

The genuinely useful thing the AI did for me was not intelligence. It was the absence of history. It had no grudge to protect, no war story to justify, no reputation invested in having been right about a desktop environment during the Bush administration. It just matched requirements to reality as it currently exists.

That is an uncomfortable insight if you make decisions for a living.

The universe collects on old debts

I took a backup, discovered that XFS cannot be shrunk - a fact you generally learn at the moment it becomes expensive - and reinstalled from scratch on Btrfs. Tumbleweed. Snapper. Automatic snapshots. A safety net at last.

Within days, KDE started crashing on login and refused to come up. Somewhere in the machine, karma had been accruing interest for twenty years.

So I rolled back to a previous snapshot. It didn't help.

And that turned out to be the most valuable lesson of the whole migration, because it is not a bug - it is design. On a default openSUSE setup, the snapshots cover the root filesystem. Your home directory sits in an excluded subvolume, deliberately, so that a system rollback doesn't also delete the document you wrote this morning. Perfectly sensible. Also: my breakage lived in a Qt config file inside my home directory. The safety net I had chosen the entire distribution for did not cover the layer where the failure actually occurred.

Every backup strategy you have ever approved has a version of this hole in it. The restore covers the part that was easy to snapshot, and the outage lives in the part that holds the state. You will find out which is which either during a calm Tuesday test or during an incident with an audience. Your choice.

I deleted a handful of KDE and Qt configuration files, moved to a newer Tumbleweed snapshot, and the problem never returned. Then something annoying happened: I started to like it. YaST still feels like an actual control panel for an actual system rather than a settings app that hides the switches it doesn't trust you with. Zypper resolves dependencies like an adult. The chameleon is still from Nuremberg, which means I toured the entire distro landscape for two decades to end up a few hours' drive from where I started.

I am mildly annoyed at myself. Not for leaving. For not looking back sooner.

πŸ‘‰ Re-evaluate at peak stability, not at peak pain. Decisions made during an outage default to the incumbent, because panic optimizes for familiarity. The only time you can honestly compare options is when you don't need to.

πŸ‘‰ Every technology opinion needs a TTL. Write the verdict down with a review date. "Vendor X is unusable" expires in eighteen months. So does "that tool is a toy" and "that person can't handle escalations."

πŸ‘‰ Your grudge is running on cached data. The product that burned you has shipped forty releases since. So has the colleague you wrote off. Serving a stale cache with high confidence is not experience, it's just latency.

πŸ‘‰ Test the restore, not the backup. The snapshot that covers everything except your actual state is a feeling, not a recovery plan.

πŸ‘‰ Ask the question to something with no history. A junior, an outsider, a machine - anyone not emotionally invested in having been right the first time. Their naivety is a feature.

Twenty years of loyalty to a decision I made before I understood what a mount point was. Being German, I am contractually entitled to hold a grudge - but the interest rate on that particular one turned out to be brutal.

πŸ‘‡ What is the strongest technology opinion you hold that you have not actually re-tested since the last time it was true?

This article is also available in German.

Cover Image

When My Favorite Computer Magazine Quietly Became a Shopping Catalog

Every two weeks, like clockwork, I walked to the kiosk and bought the same magazine. This was back when "the cloud" still meant rain. The pages were glorious. Linux distribution shootouts - Debian vs. SUSE vs. Fedora vs. Ubuntu, complete with benchmark tables nobody asked for and everybody loved. Firewall configs. Server hardware reviews. How to set up remote access so you could SSH into your home box from the office and play terminal Tetris when the boss wasn't looking. How to share your "totally legally acquired" movie collection across the family network. Best open-source database. Troubleshooting tips that actually worked.

It was nerd church. I read every word, including the ads.

Then something shifted. First subtle, then about as subtle as a forklift through a glass door.

The Linux comparison shrank to half a page. Then a quarter. Then a sidebar. In its place: best LCD monitor for Photoshop. Best smartphone for the outdoorsy type. Best printer for home photo printing. Music apps you "cannot live without." Best home cinema projector. Dolby surround systems that cost more than my first car.

The magazine hadn't been cancelled. It had been quietly lobotomized. Same logo, same price, same kiosk. Different soul. It went from "build it yourself" to "buy this and consume." And teenage me was annoyed because none of it helped my career. I didn't want to know which speaker thumped hardest, primarily because I didn't have any money to buy one. I wanted to know how to make the server stop falling over.

Here's the part that still fascinates me 20-something years later.

That magazine wasn't dying. It was early.

It was just an ahead-of-the-curve preview of where the entire economy was heading. Away from creation and customization, toward consumption and convenience. I think the first iPhone was a big step in this direction. The magazine didn't sell out. It read the room before the room knew it had been sold.

Fast forward to now. My robot lawnmower has an app. My vacuum has an app and, I'm fairly sure, opinions. My motorcycle has an app. My airbag jacket - the thing whose only job is to inflate before my spine meets asphalt - also has an app, and it would very much like to send me notifications. Everything is bigger screens, higher resolution, richer audio, more video, more feed, more pull. The whole online economy now runs on a single currency, and it isn't money. It's your attention. Money is just the exhaust.

This is the bit the smart people figured out a while ago. When everything becomes abundant - storage, compute, content, even decent software - the only thing left scarce is human focus. So that becomes the battlefield. Social media is just the most honest example. The lawnmower is the funniest one. In essence, people care less about if and how things work and more if they look good and are easy to operate.

Now here's where it gets interesting for anyone reading this who, like me, built a career on knowing the deep technical stuff. You don't have to like what comes next; I certainly don't, but it's your best toolset to survive this economy.

I spent years being the person who could go three levels down where the documentation ends and the screaming begins. The kind of engineer companies kept slightly hidden from HR, because if HR met us in the wild, they'd have questions about the eye contact, or lack thereof. For a long time, that was enough. You could be the network guru living happily in the server room, fluorescent tan, ashtrays, and all, and the world routed around your social quirks because your knowledge was irreplaceable.

That deal is mostly off the table now. Not entirely - a handful of companies still need humans who know one thing so deeply it borders on a personality disorder, and they'll happily pay to keep them away from meetings. But that's a shrinking list. The market reorganized itself around the attention economy, and it didn't ask the nerds first.

So if your entire career is built on pure technical depth, here's the unglamorous advice nobody wants to hear:

πŸ‘‰ Learn enough "normie" to translate. Some web dev, some basic design sense, some understanding of how things are sold and marketed. Not to become a marketer. To stop being illegible to the people who control budgets.

πŸ‘‰ Understand attention as a system. Why people click, why they don't, why a beautiful dashboard gets funded, and a brilliant, ugly one gets cancelled. You don't have to like it. You have to see it. Engineers who understand the attention game stop losing arguments to people with worse ideas and better slides.

πŸ‘‰ Make your work visible. The deepest technical save in the world is worthless to your career if it happens silently in a terminal at 3am and nobody upstairs ever hears the story. Packaging your value isn't selling out. It's self-defense. My old manager told me: "do good and talk about it". I hate it, you hate it, but this is how it works.

πŸ‘‰ Treat empathy as a technical skill. Understanding "normies" - meaning, you know, most humans - isn't a betrayal of the nerd identity. It's a hard skill that makes you better at support, better at architecture decisions, and better at not being the engineer who's technically right and professionally invisible.

And for the leaders reading this: the lesson runs in the other direction too. Your magazine-reading server gremlins 🐎 are still the ones who actually keep the lights on. The shift toward shiny and consumable doesn't mean depth stopped mattering. It means depth got harder to see from the executive floor because middle management filters out anything that doesn't fit on a slide. The engineer hiding in the server room didn't become useless. They became hard to find. There's a difference, and confusing the two is how companies lose their best people to the one competitor who still knows the difference. So if you can, try to at least protect them a little; they are likely struggling in this new world without even knowing so.

I never did find out which projector was best. But I learned the more useful thing my old magazine accidentally taught me on its way out: when the whole world pivots from building to consuming, the rare and valuable move is to stay someone who can build - and learn just enough of the consumer's language to make the builders matter again.

What did your favorite tech magazine slowly turn into? And did you broaden out, or double down on the depth? πŸ‘‡

This article is also available in German.