TrenchOps 🐎

Insights from the tech trenches

Cover Image

📌 Welcome @ TrenchOps 🐎

About TrenchOps: An IT Support Engineering & Tech Career Blog

TrenchOps is a blog about IT support engineering, technology careers, and management in the software industry - written by a Lead Support Engineer with 20+ years in enterprise software, from 1st- to 3rd-level support and escalation through technical account management, program management, and consulting.

The name: "TrenchOps" means "operations in the tech trenches" - hands-on frontline IT support, technical operations, and enterprise software work, not military combat. "Power horses" 🐎 is our term for the dependable engineers who carry a team's load.

Who it's for

  • Engineers (support, SRE, DevOps, security, architects) growing their careers
  • Managers and team leads building real processes and preventing burnout
  • Senior leadership wanting visibility into the mismanagement that quietly kills retention, quality, and revenue

What it covers

Practical career and leadership guidance plus anonymized stories from enterprise IT support: promotion strategy and burnout, imposter syndrome and boundaries, retaining high performers, leadership anti-patterns, support as a revenue advantage, documentation and KCS, and IT resilience without single points of failure.

Categories

Field notes for the engineers, managers, and leaders who keep enterprise software running.

Cover Image

What Nobody Warns You About Before Your First Invoice

Three things kill small businesses that never appear in a business plan: the room you work in, the stories you absorb, and the money that was never yours.

I have been self-employed a few times. Before my enterprise career, between two of them, and on the side. None of it made me rich. All of it taught me things you cannot learn from a startup podcast, because podcasts talk about scaling and my problem was that a customer's dog did not like me.

Let me hand over the three lessons that actually cost me something

Your business has a physical address, even when you think it doesn't

Everyone starts in the bedroom, the garage or the study. Fine. Nobody plans the other rooms.

One of my ventures was classic IT service. Persuading stubborn printers to acknowledge the existence of WiFi. Removing pre-installed bloat from our favorite corporate operating system. Explaining, for the fourth time, that a tablet needs updates. Mostly on-site, at people's homes.

Some clients objected to travel fees. Reasonable. They wanted to come to me. Slight problem: I had no office and no presentable workshop. So I invited them into my dining room - statistically the least private room in the house, right up until a stranger sits in it holding a coffee cup and telling me about his divorce while I diagnose a random reboot.

Fun fact from cognitive load research that every engineer already knows in their bones: deep debugging and polite listening cannot run on the same core. You either find the faulty power supply or you learn about the neighbor's inheritance dispute. Never both.

Then came the subset. Rude, unclean, or simply radiating a vibe that made me want to change the locks. That was the end of visitors. Strict on-site only, take it or leave it. I considered converting the garage. I considered a container office. Neither ever happened, because it was never worth the cost for a minority of customers - and that is exactly the trap. The expense is easy to justify and easy to postpone forever.

The same applies if you make things. During the pandemic, everyone discovered handmade marketplaces, to the point where the market now looks like a craft fair with more sellers than buyers. You will need space for the inventory you sell. You will need more space for the inventory you never sell. And while you are planning shelf space, plan something else too: how you will emotionally process cutting your price below your material cost for an object that took you six hours of your finite life.

πŸ‘‰ Before the first customer: decide where money, goods and strangers physically meet. That decision is a business decision, not a furniture decision.

You will be paid in cash and in other people's grief

I am read as a stoic person. Reasonably accurate. I can understand somebody's pain without soaking it up like kitchen roll, which I always considered a professional advantage.

So I assumed I could handle what clients volunteer while you sit in their living room. Illnesses. Losses. Ruined lives. And initially, I could.

What I underestimated was the second-order effect. I kept wanting to tell my wife about it. Not as drama - just as "you would not believe what I learned about human beings today." Turns out I was the messenger, and she was absorbing what I was merely transporting. We had the talk. My license as gossip correspondent was revoked, correctly.

Then, holding it all myself, it started to weigh. There were customers who were lovely by every measurable standard, and I stopped wanting to drive there. Widows. Parents who had buried a child. People whose existence collapsed after a diagnosis. Nothing to win, one hour of billable work, and a long silent drive home.

So I declared a no-small-talk policy. It survived approximately one appointment. Because pushing back on a lonely person mid-sentence makes you the cold, rude technician - which is both unpleasant and bad for business. And sometimes, honestly, the stories were fascinating. That is the sneaky part: curiosity gets you in, and the sadness stays.

A colleague with an actual shop looked at me like I had described a self-inflicted wound. His customers walked in, handed over a laptop, walked out, came back later. Purely transactional. His counter was not furniture either. It was an emotional firewall that also happened to sell cables.

And the last realization was the one that broke my pricing: some clients were not hiring a technician. They were hiring human company. Especially elderly ladies, whose laptop mysteriously needed rescuing for the third time that month. Once you understand that, charging the full hourly rate becomes an ethics exam you did not sign up for.

πŸ‘‰ If your work takes you into other people's homes, or brings them into yours, you are in an emotional-labor business that happens to involve technology. Price it, or shield it with a counter, a shop, a workshop - something with a door.

Also budget for the unglamorous risks: dogs with opinions, apartments you want to shower after, and the rare genuinely bad human being.

Half of every invoice was never your money

This is the one nobody gets right, including people with an accountant and a lawyer. I had both.

Your private bank account is free. Your business account is not. Tax notices arrive from institutions you did not know existed. You will refund a customer. You will break something and pay for it. Insurance, repairs, tools, marketing material, professional fees, the association nobody told you was mandatory.

So I started doing something almost aggressively simple: 50% of every invoice went straight into a separate pot and stayed there until year end. Taxes, surprises, apologies. Whatever remained afterwards was profit, and it felt like a gift.

The point is not the percentage - yours depends on your country, your legal form and your accountant. The point is that revenue is a number that lies, and the pot does not. For the first time, I knew what I actually earned per hour instead of what my invoices claimed. Several jobs I was proud of turned out to be expensive hobbies with a customer attached.

Self-employment does not fail because your idea was bad. It fails because you priced the work and forgot to price the room, the emotions and the taxman - and all three of them invoice you anyway.

πŸ‘‡ What was the cost item that ambushed you in your first year - and would you have believed anyone who warned you?

Cover Image

Your Brain Doesn't Learn From Success

It Learns From Being Wrong

I once watched a senior engineer - twenty years of experience, could read a stack trace like a bedtime story - try to learn to juggle at a company offsite. Three balls, two hands, zero dignity. She dropped everything for twenty minutes straight while interns filmed it.

Six months later she told me that offsite was the week she finally cracked a certification she'd been stuck on for a year. She was convinced it was the juggling. She might have been right.

Here's the mechanism, and it's backwards from what most people believe: the adult brain doesn't rewire because you feed it information. It rewires because you feed it errors. When your predictions about the world fail hard and repeatedly, your nervous system releases the neuromodulators - dopamine, norepinephrine - that mark active circuits for change. No errors, no chemical signal, no rewiring. Recent neuroscience even points to a dedicated brain-wide "model failure" alarm system (the locus coeruleus) whose whole job is shouting "your predictions are garbage, update everything."

That feeling of frustration? That's not a bug. That's the smell of the signal being released. The brain games industry sells you comfortable puzzles you're already decent at. Comfortable means low error rate. Low error rate means the alarm never fires. You're paying a subscription to not learn.

There's a catch: the frustration only works if you stay in it. Huberman's summary of the research is blunt - if you hit frustration and lean in for another stretch of failed attempts, you've opened the plasticity window. If you hit frustration and walk away, the same window rewires you toward whatever comes next, which is usually doomscrolling and resentment. Frustration is a loaded weapon. Direction matters.

This is why the best engineers I've worked with all share one weird trait: they are suspiciously calm while being wrong. Not because they enjoy it, but because somewhere along the way they learned that "this doesn't work and I don't know why" is the productive state, not the failure state. The junior who rage-closes the laptop and the senior who mutters "interesting" at the same error message are running completely different neurochemistry.

πŸ‘‰ Errors, not information, trigger rewiring. Reading about a skill produces near-zero plasticity. Failing at it produces the chemical cascade.
πŸ‘‰ Short bouts beat marathons. The research says roughly 7-30 focused minutes of deliberate error-making, then stop before you start adding new categories of mistakes.
πŸ‘‰ Stakes gate the speed. The brain allocates plasticity by importance. A skill you genuinely need rewires faster than a hobby you're dabbling in. Give yourself a real reason.
πŸ‘‰ Frustration is a fork in the road. Push through: you learn. Walk away: the open window learns your escape behavior instead.
πŸ‘‰ Pick something you're bad at on purpose. Zero talent means maximum error rate means maximum signal. Your ego pays, your cortex collects.

Competence is comfortable. Comfort is neurologically silent. If nothing about your week embarrassed you, nothing about your brain changed either.

πŸ‘‡ What's the last skill that made you feel like a complete beginner - and did you push through or quietly retreat to things you're already good at?

This article is also available in German.

Cover Image

The Strategy That Never Won a Single Match - And Won Everything

What a 1980 computer tournament explains about your worst customer, your dumbest feud, and your exit interviews

Around 1980, a political scientist ran a tournament for computer programs playing the iterated Prisoner's Dilemma. Game theorists, mathematicians, economists, the whole intellectual heavyweight class, submitted strategies. Some were baroque: probabilistic betrayal, opponent profiling, deception layers built to squeeze a few extra points out of the naive.

The winner was a handful of lines of code, submitted by a mathematical psychologist. Cooperate on move one. After that, do whatever the other player just did. That's it. Tit For Tat. Shortest entry in the field. It won again in the second tournament, where everybody already knew it was coming and had specifically designed strategies to beat it.

Here is the detail people always skip: Tit For Tat never won a single individual match. It cannot. By construction it can never score more points than its opponent, because it only ever defects in response. The best possible outcome against any given opponent is a tie.

It won the tournament by being the strategy that everyone did well against, including itself.

Read that twice if you measure your quarter in individual wins.

The three losing strategies, in office dress

  • Always defect. Takes credit upward, distributes blame downward. Wins spectacularly early. In tournaments it dies, because nasty strategies also meet each other, and then they starve together. In companies it survives longer than it should, because every reorg supplies fresh unmet victims. It's a strategy that only works with high player turnover, which explains a great deal about companies with high turnover.

  • Always cooperate. The nicest engineer on the team. Absorbs every emergency, never says no, gets farmed by exactly the people you'd predict, and is quietly gone in eighteen months. Unconditional kindness is not a strategy. It's a resource for other people's strategies.

  • Grim trigger. Cooperate until the first betrayal, then defect forever. The senior architect who was overruled once in a design review three years ago and has been passive-aggressively correct ever since. Locally satisfying. Globally catastrophic.

Tit For Tat wins because it is four things simultaneously:

  • nice (never defects first)
  • retaliating (defection costs something, immediately)
  • forgiving (one clean move restores trust)
  • clear (the opponent can figure out the rules in a single round).

Clarity is the underrated one. Complicated retaliation policies get read as mood swings instead of boundaries.

The customer everyone had labelled hostile

I once inherited an account that internal folklore described as hostile. The customer's technical lead had a reputation. Escalated everything. CC'd executives with the enthusiasm of a man who had discovered the address book. Called our service level agreement "creative fiction." Two account managers had quietly rotated off before me.

I spent an afternoon reading ticket history before my first call. What I found was not a difficult human. What I found was eighteen months of small defections from our side. A promised callback that never happened. A bug marked "fixed" in a release note where it demonstrably wasn't. A workaround that was really a distraction delivered with a straight face.

He wasn't hostile. He was mirroring. He had opened with trust roughly a year and a half earlier, got defected on, and was now running the algorithm exactly as specified. From where he sat, he was the cooperative player being perfectly reasonable about a pattern.

So I did the only thing the model permits. Cooperate first, visibly, and make it cheap to verify. One promise per week, small, kept. When we couldn't fix something, I said so and gave a date for the next honest update instead of a fake date for the fix. I stopped forwarding optimismΒ upward, and stopped forwarding it downward too.

Six weeks. Not because he was stubborn, but because trust in an iterated game rebuilds at precisely the same rate it was destroyed: one move at a time. There is no bulk import.

By month four he was pre-filtering his own team's tickets before sending them to us, because he had concluded we were worth not wasting. That is what a cooperating opponent looks like. Free labor, volunteered, because the payoff matrix finally rewarded it.

The flaw that runs your company's oldest feud

Pure Tit For Tat has a fatal weakness, and it's the real reason this article exists.

Two Tit For Tat players in a noisy environment - where a move can be misread - lock into a death spiral. One accidental defection, one dropped ticket, one email that lands colder than intended, and the two of them alternate punishment forever. Both worse off. Both behaving correctly. Neither wrong about the last move.

Now look around your office. That is the entire mechanism behind every long-running war between two departments who no longer remember what started it. Engineering thinks support dumps garbage tickets. Support thinks engineering ignores them. Both are retaliating for something that happened in a previous fiscal year, and both are factually right about the most recent exchange.

Corporate life is a maximally noisy channel. Chat, timezones, translation layers, middle managers helpfully rephrasing things into their opposite. If your policy is strict reciprocity with zero slack, you will end up in feuds with people who were on your side. That isn't integrity. It's a signal processing failure.

The fix from the later research is almost insulting in its simplicity: forgive occasionally without being provoked into it. Not always, not never - a small, deliberate rate, something like one in ten. Cooperate even when the last move was a defection. That single modification breaks the spiral and beats strict mirroring wherever information is imperfect.

Information in your company is imperfect. Permanently. So the optimal human strategy is not an eye for an eye. It's an eye for an eye, with a budgeted error allowance for grace.

That budget is not softness. It's the maintenance cost of a system you intend to keep using.

Why everything goes wrong on the way out

The second thing the theory tells you, and nobody puts on a slide: if both players know the game is finite, the logic unravels backwards from the last round. No retaliation is possible after the final move, so defect there. Which makes the second to last move effectively final. And so on, all the way back.

This is why almost every unpleasant surprise in business happens at the exit. The vendor whose quality collapses in the last quarter of the contract. The employee who has already signed elsewhere and stops caring. The customer who suddenly disputes twelve months of invoices in week two of the notice period. The manager who becomes honest with you only in the exit interview, which is the corporate equivalent of confessing to an empty room.

Nobody became a worse person. The horizon just became visible.

Two consequences. Keep the game long and visibly repeated, because reputation only has value when there is a next round. And when the horizon genuinely is finite - offboarding, contract end, a vendor you're replacing - stop expecting reciprocity to carry the load and switch to structure. Milestones, retained payments, handover checklists signed before the final salary run. Not cynicism. Arithmetic.

The best hire I ever made came out of the opposite move. A power horse 🐎 who had been treated as always-cooperate by three previous employers, underpaid accordingly, visibly braced for the usual. I opened with trust I hadn't earned yet: real budget, real authority, no probation theater. Cooperation came back on the first round and compounded for years. Cost of the move: the risk of being wrong once.

πŸ‘‰ Open cooperative. Every time, with everyone. The downside of one betrayal is small; the upside of a compounding relationship isn't.

πŸ‘‰ Retaliate fast and proportionally, then stop. A defection you swallow silently teaches the other side that defection is free. Delayed punishment reads as mood, not consequence.

πŸ‘‰ Be legible. The strategy won because opponents could predict it. If people can't model you, they can't cooperate with you - they can only hedge against you.

πŸ‘‰ Budget forgiveness before you need it. Decide in advance that roughly every tenth insult is noise, not malice. You'll be right most of the time and you'll save relationships pure logic would have incinerated.

πŸ‘‰ Never play to win the individual round. The strategy that couldn't beat anybody beat everybody. Your quarterly hero metric is scoring the wrong game.

πŸ‘‰ Check who defected first before labelling someone difficult. Read the last eighteen months of history. It's usually your own company in the log.

πŸ‘‰ When the horizon becomes finite, replace trust with structure - before the last round, not during it.

The most sophisticated players in the room lost to four lines of logic, because complexity is what you build when you're trying to exploit someone. Simple, predictable, reciprocal is what you build when you plan to still be here next year.

πŸ‘‡ Which internal feud in your company is just two reasonable teams stuck in a retaliation loop that nobody started on purpose?

This article is also available in German.

Cover Image

Count the Rounds

Why a five-stage interview process tells you everything about the job - before you even get it

The job post says "fast-paced environment." The process says recruiter screen, hiring manager call, take-home assignment, technical deep dive, second technical deep dive, "culture add" panel, psychometric questionnaire, and a final chat with a VP who will ask you where you see yourself in five years.

That's eight rounds. For a company that claims to "move fast".

Believe the process, not the post.

Times are rough. Plenty of good people lost their jobs to the AI wave, and plenty more are quietly refreshing job boards at their desks. My personal bet is that once the AI honeymoon ends, a lot of companies will discover that their institutional knowledge left in a cardboard box, and they'll be rehiring. Until then, desperation is a terrible career advisor. It tells you to say yes to every round, every test, every "just one more quick call with a stakeholder."

Don't. Here's why.

What an interview is actually for

A competent company reads your CV and can map your skills to its needs. With a little effort, it can verify most of your history. So the interview only has three real jobs.

  1. Consistency checks. Is the CV true? This was always my favorite part when I ran interviews. I'd pick the area where the candidate should be strongest, like a product they supported for four years at a vendor, or a certification they listed proudly. Then I'd ask what any actual employee would know. I never cared if someone was shaky on network protocols. I cared a lot if someone was shaky on the thing they called their expertise. A gap in a claimed strength isn't a knowledge gap. It's a honesty gap.

  2. Thought process. Your future peers want to see how you break a problem into pieces, where you start, and when you admit you've hit your limit. Not knowing an answer is fine. What they're really measuring is trainability. "I don't know, but here's how I'd find out in ten minutes" beats a confident wrong answer every day of the week.

  3. Vibes. Can we work with this person? Recruiter, manager, team and sometimes a director all pick this up. You can't measure it on paper and barely on a chat. Video gets you partway, and in person is best. In my experience this is where most people "fail," and that's fine. A conglomerate needs a different temperament than a startup, and sales needs a different one than back office. You'll rarely be told this is why you didn't get the job. Legal caution keeps that feedback vague, so don't take the silence personally.

That's the whole list. Three jobs.

The healthy process fits on a napkin

  • Recruiter conversation.
  • Call with the hiring manager plus a peer.
  • Optionally, a second and more technical call.
  • Optionally, a short chat with the manager's boss. We called it the "please don't mess this up" call. Its entire purpose was to confirm the candidate didn't insult anyone's mother.

Three rounds, four if the role is senior. BrightHire's hiring data lands in the same place: the sweet spot for most roles is two to four rounds, and past four the extra signal drops sharply while strong candidates start dropping out. Robert Half found that 57% of job seekers lose interest when a process drags on too long. Guess which 57% have other options.

Conway's Law, now in HR

Melvin Conway observed in the late 1960s that organizations design systems that mirror their own communication structure. That's usually said about software architecture. It applies just as well to the hiring pipeline. The interview process is the one piece of internal plumbing a company shows you before you sign.

So read it like a system diagram.

Too many technical rounds means leadership doesn't trust the one or two engineers it sends in to judge a candidate. Either those engineers aren't good, or management is bad at picking people, or management simply doesn't trust its techies. None of these is a fun place to spend your forties.

Too many manager rounds means the hiring manager can't decide after one or two conversations. Either they can't read people, they lack the authority to decide, or they're afraid to. Whichever it is, you'll meet the same pattern later in your budget approvals, your change requests and your promotion case.

Now the part most people miss. More rounds don't make hiring more careful. They make it more average.

Every interviewer in a long chain knows three others will also weigh in, so no single person owns the decision. Saying "no" is risk-free, because nobody ever got fired for rejecting a candidate. Saying "yes" is risky, because your name is on it if they flop. Stack enough of those asymmetric vetoes and the process quietly filters for the most inoffensive person in the pool. The spiky ones get shaved off: the brilliant, direct, slightly odd engineer, or the neurodivergent troubleshooter who answers the question you asked rather than the one you meant. Those are exactly the people who carry teams.

A long pipeline doesn't lower the chance of a bad hire. It lowers the chance of a remarkable one. The company feels safer and gets blander, and your future colleagues are the survivors of that filter.

As a power horse 🐎, you don't want to work where eight people have to agree before anything happens. You already know how that place runs internally. They showed you.

Confession: I built take-home tests

Full disclosure: I've designed take-home assessments myself. Before the first technical interview, candidates got a PDF with a few tasks about our products. The idea sounded great. It would show genuine interest in the company and prove the ability to learn something new.

The reality was that one candidate solved it and shared the answers with five friends. They aced the take-home, defended their steps convincingly in the interview, got hired, and turned out to be a disaster. We'd tested their ability to memorize someone else's homework. Then we were stuck with them.

That was before AI. Today the take-home tests mostly your subscription tier. Karat's survey of 400 engineering leaders found 71% say AI makes it harder to assess technical skills, and many employers are drifting back toward live sessions. On top of that, asking someone to invest four or more unpaid hours per application is unfair. If they apply to ten companies, that's a full working week of free consulting.

What I'd do instead

Go back to basics, with one modern twist.

  • Verify the CV properly. Make reference calls and ask about specific details.
  • Check the vibes in person or on video.
  • Cross-check claimed strengths with questions that don't paste well into a chatbot. That means either conceptual "why" questions, or tricky edge cases where an AI answer sounds perfect and is quietly wrong.
  • Let them use AI during the interview. They'll use it on the job anyway, so watch how. Give them a problem where the model is likely to hallucinate and see if they catch it. There's a growing term for people who forward AI output without checking it: the "meat proxy." You want to find out in the interview whether you're hiring an engineer or a very polite copy-paste interface.

And for the hiring side: use the probation period

Most companies treat probation as a formality. Treat it as the real final round.

Hire decisively, then watch closely. If it's clearly not working, act early. Don't fall for "but we already spent three months training them." That money is gone either way, and the sunk cost fallacy just adds more months to the bill. Two bad hires caught during probation cost far less than one loyal power horse 🐎 who quit because they were carrying a dead weight colleague nobody dared to address.

(Probation rules vary a lot by country. In Germany, for example, it's capped at six months with a two-week notice period. Know your local rules and document honestly, but do use them.)

πŸ‘‰ An interview needs to do three things: check the CV is true, see how you think, and see whether people want to work with you. Anything beyond that is ceremony.

πŸ‘‰ Three rounds is healthy, four is the ceiling for senior roles, and five or more is a diagnosis of the company, not an assessment of you.

πŸ‘‰ Long pipelines select for inoffensive and filter out exceptional, because every "no" is free and every "yes" is personal.

πŸ‘‰ Take-home tests now measure how well you prompt a chatbot and who your friends are. Test AI judgment live instead.

πŸ‘‰ Ask for the full process on the first call. It's the cheapest due diligence you'll ever do.

The number of interview rounds is the number of people who'll need to approve your ideas later. Count accordingly.

πŸ‘‡ What's the longest interview process you ever went through - and was the job worth it?

Cover Image

Your AI Agent Reads 50,000 Tokens to Learn One Thing

Compression as via negativa: the cheapest token is the one you never send

Take any coding agent session and open the raw request log. Not the pretty transcript - the actual payload. What you find is a JSON search result with 100 hits where 3 mattered, a log file where the interesting line is buried in 800 lines of "INFO: still fine," and a directory listing of a repo the model already walked twice.

Then apply the uncomfortable ratio: cost of the finished output versus the raw inputs actually required to produce it. The model needed maybe 2,000 tokens of real signal. You paid for 55,000. In manufacturing terms, that is a part made of solid gold to hold a plastic clip in place.

I spent a few weeks running everything - coding, research, boring text work - through a compression proxy called Headroom, and the interesting part was not the money. It was watching exactly how much of what I "sent" to the model was never information in the first place.

What it actually does

It sits between your agent and the API and squeezes everything the model reads - tool outputs, logs, retrieved chunks, conversation history - before it goes upstream. Compression happens locally, nothing gets shipped off to a third party to be shrunk. The model can pull the original back if it needs it.

The project's own benchmarks show roughly 20% on code search, 40-60% on incident debugging and codebase exploration, and the big numbers (86%+) on JSON arrays, which makes sense: JSON is mostly punctuation and repeated key names, a format designed for parsers and billed as if it were prose. Structured logs compress hard too. Source code mostly passes through untouched, which is the right default - you do not want your agent reasoning about a lossy version of the file it is editing.

Honest caveat, because I would rather you trust me in six months: the dashboard counts tokens it compressed, not tokens your provider stopped charging you for. Those are related but not identical numbers, and some users report the delta being smaller than advertised once retrieval round-trips are included. Treat the dashboard as a direction indicator, not an invoice.

Which brings up the part nobody markets: if you are on a flat subscription plan, you save exactly zero euros/dollars. What you save is runway. Fewer tokens consumed means you travel further before slamming into the 5-hour and weekly usage windows. For anyone who has had a long refactor cut off mid-thought by a rate limit, that is worth more than a discount.

Setup on Linux, in the time it takes to make coffee

Install and verify:

uv tool install --python 3.13 "headroom-ai[all]"
command -v headroom
headroom doctor

You want "Proxy" showing a green checkmark. If it does not, stop here and fix that first - do not proceed on hope.

Optional but sensible, run it as a user daemon so it is simply always there:

# ~/.config/systemd/user/headroom.service
[Unit]
Description=Headroom compression proxy

[Service]
ExecStart=%h/.local/bin/headroom proxy --port 8787
Restart=on-failure

[Install]
WantedBy=default.target

systemctl --user enable --now headroom
ss -lntp | grep 8787

That second command is not cosmetic. Confirm it is bound to 127.0.0.1 and not 0.0.0.0. An unauthenticated proxy that sees every prompt you write, listening on all interfaces, is not a productivity tool. It is a gift to whoever else is on your network.

Then point your client at it, in ~/.claude/settings.json:

{
Β "env": {
Β "ANTHROPIC_BASE_URL": "http://127.0.0.1:8787"
Β }
}

Restart the CLI and your IDE - VS Code or VSCodium will not pick up the new base URL otherwise. Open http://127.0.0.1:8787/dashboard and check that "Request Health" and "Live Activity" show non-zero numbers. That is the whole installation. The rest is watching a counter go up while you do the work you were going to do anyway.

When it breaks, and it will

It is a young, fast-moving project. Rarely, requests hang or die. The triage order, learned the boring way:

πŸ‘‰ Stop the current model session, restart the tool or IDE, tell the agent "continue." Fixes most of it.

πŸ‘‰ Still failing? systemctl --user restart headroom.

πŸ‘‰ Still failing? It is probably not your machine. Check status.claude.com before you debug anything else - degraded upstream service looks exactly like a broken proxy from where you are sitting, and I have wasted honest minutes proving that.

πŸ‘‰ Genuinely a regression in a new release? Pin backwards: uv tool install --python 3.13 "headroom-ai[all]==0.37.0" --force. Fast-moving projects reward people who know how to step back one version instead of filing an issue and waiting.

That last point is the real skill, by the way. Not the tool - the habit of having a rollback path before you need one.

The part that generalizes

Compression proxies are a workaround. The underlying problem is that we hand agents firehoses and call it context. We pipe in complete API responses, full log files and entire directory trees because it is easier than deciding what matters, then pay per token for the privilege of making a language model do our filtering at premium rates.

The best part is no part. The cheapest token is the one you never send. A tool that compresses your junk is strictly better than not having it - but a grep with a sane filter, a log level that is not DEBUG in production, and an API response that returns fields instead of everything would have gotten you most of the way there without any middleware at all.

πŸ‘‡ What is the dumbest thing your agent has ever been made to read in full? Mine spent real money ingesting a 400-line INFO log to find one timestamp.

Cover Image

Be Employee Obsessed

Everyone says "customer first". Almost nobody means it.

Years ago, I shared a few carbonated drinks with an SVP whose management philosophy fit on a bar napkin: "Be employee obsessed."

I was skeptical to the point of rudeness. So I asked him straight: "with that approach, how do you make sure customers get world-class service?"

He didn't blink. "If you hire top performers, you never have to worry about customer satisfaction. I have never once hired a talented support person who wasn't already customer obsessed. Nobody ends up in this job by accident. It's ingrained, or they'd have picked a role where humans don't call you when their week is on fire."

So the model was simple: hire people who are obsessed with customers, then spend your own energy protecting them. Not managing them. Protecting them. From three specific things: corporate nonsense, burnout, and what he cheerfully called "stupid customers".

Let me unpack all three, because each one had teeth.

Burnout is rarely about volume

His claim: information technology has one of the worst burnout rates in white collar work, and the cause is usually misdiagnosed as "too much work". Sometimes it genuinely is too much work. But more often, he said, people are either doing the wrong things, or doing the right things in a broken flow imposed on them by the company.

That's a distinction most executives never make. Ten hard tickets solved cleanly leave a power horse 🐎 energized. Three easy tickets routed through four approval layers, two tools that don't talk to each other, and a mandatory status meeting leave the same person hollow. Effort doesn't burn people out. Friction does. Every handoff and approval gate bleeds energy that never reaches the customer, and the person feeling that loss most acutely is the one holding the ticket.

So he fought - hard, politically, for months - to get his organization a dedicated developer. Reporting to him. Not to a platform team, not to a shared services pool, not to a manager with a competing roadmap. One engineer whose entire job was removing the daily friction his best people ran into.

Middle management hated it. It looked like empire building. It was actually the highest-return headcount in the department, because it converted senior engineer hours from "fighting our own tooling" back into "fixing customer problems". If you want a number to think about: take the fully loaded cost of a resolved ticket and divide it by the minutes of actual engineering thought inside it. When that ratio is grotesque, your inputs aren't expensive. Your process is.

The "stupid customer" problem, honestly stated

Nobody in that org thought customers were stupid. Customers pay the bills, and treating them with contempt is both immoral and commercially suicidal. But two failure modes appeared over and over:

Customers who had no realistic expectation of what support is for. And customers whose own staff had never been trained to operate the product they bought.

Expectation-setting sounds trivial. It isn't. When a customer has a P1 - severity one, production down, hair on fire - they do not want a conversation about scope. They want it fixed before it reaches their CIO's dashboard.

And that last part was the tell. A large share of P1s were not technically P1s at all. They fell into two buckets:

Political P1s, where the real problem was internal to the customer and a support engineer was structurally the wrong audience. And cover-up P1s, where someone had broken something and wanted it quietly repaired before their own food chain noticed.

Both correlated strongly with the big global accounts that had outsourced first-line operations to the cheapest available bidder - people trained to watch for a red light on a dashboard and open a case, with zero investigation in between. Not their fault. That's exactly what they were hired and incentivized to do. Show me the incentive, I'll show you the ticket queue.

Permission to push back

His fix was a small policy with enormous consequences: his engineers were allowed to push back. Not encouraged to be difficult - allowed to disagree with a customer's severity assessment and say so.

Nine times out of ten they got it right. The remaining 1/10 they used to learn, improve and apologize. The mechanism was a fast triage question set, and the sharpest one was this: "is this a component that was working before, or a brand new implementation?"

By default, a new implementation cannot be a P1. Think about it. If something that has never worked in production is suddenly business-critical, one of three things happened: you didn't test, you didn't read the documentation, or you set yourself a timeline that only physics could refuse. None of those are outages.

But here's the part that made it survivable commercially. The first such P1 was always worked, free, fully, no lecture. "Customer first", genuinely. Then the account team was pulled in for the grown-up conversation. That conversation regularly ended in training, sometimes certification, and quite often a professional services engagement the customer needed far more than they'd ever admitted in the sales cycle.

Pushing back didn't cost revenue. It generated revenue. Turns out honesty is a premium product.

The account that hated us, and the sentence that fixed it

The best illustration came from an account bleeding dissatisfaction scores for months while every metric on the dashboard looked pristine. Mean time to close: excellent. Response times: excellent. Customer sentiment: radioactive.

A German escalation manager got the file. Initial finding: the customer asked an extraordinary number of remarkably basic questions. Engineers were closing them in minutes, mildly amused, and moving on. From the customer's side, the story was different: "your product is impossible to use."

Metrics said triumph. Customer said disaster.

So he booked a call to actually understand what they were doing. It ran over an hour. He paced the room with a headset while the rest of us pretended to work and shamelessly harvested fragments, because this account had been giving everyone headaches for a quarter.

Then, in the crisp tone that would have made Werner Herzog proud, and only a German engineer can produce at the exact moment of maximum tension, we heard:

"With all due respect - just because our enterprise solution has a fancy user interface, that does not mean it is Ubuntu."

Silence on the line. My manager's jaw dropped to somewhere around his keyboard, visibly running the math on whether to fire our escalation manager or resign first. The rest of us having the best afternoon of the fiscal year.

Days later the account manager walked in - to say "thank you".

Because the real story finally surfaced. The customer hadn't been doing any of this themselves. They had hired an external service provider who had sold himself as an expert in our platform and was, in reality, entirely lost. His survival strategy was elegant: every time he couldn't do his job, he opened a case and rated our product as broken and our service as useless. Months of dissatisfaction scores weren't customer feedback. They were one contractor's alibi.

The customer terminated him. Then they sent their own apprentice to training and made him the single point of contact for the platform.

He also arrived with basic questions. Difference: he wanted to learn, he was grateful for the time people spent explaining things, and within about six months he was running the thing competently. Smooth sailing for them, smooth sailing for us, dissatisfaction score back where it belonged.

Sometimes the fix isn't more patience. Sometimes you have to dig deeper and pull the rotten tooth.

So: policy, or slogan?

Every company on earth claims to be customer obsessed. Be honest about what it means in yours.

Because in a lot of places, "customer obsessed" means being extremely friendly while holding someone in a phone queue for an hour. Or hiring many cheap, undertrained people, so response time looks fast while resolution rate quietly rots. Or loading your best engineers to 130% while denying them dedicated training time, because training doesn't show up on this quarter's dashboard.

In the worst cases it means squeezing your strongest people until they burn out, leave, and get replaced by the next graduate - who will be squeezed on the same schedule. That isn't customer obsession. That's a slogan doing public relations for a grinder.

πŸ‘‰ If your people are obsessed with customers, your job is not to remind them. Your job is to remove what stops them.
πŸ‘‰ Burnout is friction, not volume. Audit the flow, not the headcount.
πŸ‘‰ Fund the unglamorous internal role that unblocks your senior people. It pays for itself in a quarter and nobody will thank you for it.
πŸ‘‰ Give engineers the authority to challenge a severity. Then defend them the first time a customer escalates about it, because that moment decides whether the policy is real.
πŸ‘‰ Every "difficult customer" pattern has a mechanism underneath. Find the person whose incentive is to make you look bad.
πŸ‘‰ A dashboard full of green with a customer full of rage means you're measuring the wrong thing beautifully.

Customer obsession that isn't built on employee obsession isn't a strategy. It's a mood, and moods don't survive a P1 on a Friday.

πŸ‘‡ What's the most creative thing a customer has ever labeled a P1 - and did anyone at your company have the standing to say no?

This article is also available in German.

Cover Image

You Bought a Racehorse and Rented a Treadmill

The strangest math in enterprise hiring: paying sports-car money for output you then throttle by hand

A while back I sat in a coffee shop with someone from a very large, very admired technology company. Nine-figure revenue per quarter, the kind of place that puts "we hire only the top 1%" in its recruiting deck without irony.

He told me his last quarter honestly. Two weeks of real engineering. Ten weeks of alignment.

His compensation, fully loaded, was roughly the price of a mid-range sports car per year. So the company paid sports-car money and received, in terms of the thing they actually hired him for, a used scooter. Not because he was lazy. Because 80% of his calendar was owned by people whose job was to make sure nobody did anything surprising.

I hear a version of this story from nearly everyone I talk to who works inside the big names. It is consistent enough that it stopped being anecdote and started being physics.

The idiot index of a knowledge worker

There is a useful little instrument from manufacturing: take the cost of a finished part and divide it by the cost of its raw materials. A machined bracket that costs 200 euros but contains 8 euros of aluminum tells you something. Not that aluminum is expensive. That your process is stupid.

Apply that to a senior engineer.

Raw input: one very expensive brain, capable of shipping things that make money.
Finished output: three slide decks, one Jira epic groomed into a fine powder, and a fix that shipped six weeks after it was written because it needed sign-off from a governance board that meets fortnightly.

The ratio is grotesque. And like the bracket, it does not indict the aluminum. Nobody in these companies has a talent problem. They have a process problem wearing a talent problem's clothes, becauseΒ "we need to hire better people" is a budget request and "our approval chain destroys value" is a resignation letter.

Why the screws tighten exactly when they shouldn't

Here is the part that genuinely puzzled me for years.

Every time the horizon darkens - a bad quarter, a funding chill, "macroeconomic outlooks"- leadership reaches for control. Hiring freeze, travel freeze, headcount review, mandatory return to office, a new weekly reporting cadence, aΒ "focus" initiative that adds three dashboards. They reduce the heat to save fuel. But the fuel bill is identical, because the salaries are already committed. All they cut was output.

You are still burning the same wood. You just closed the flue.

Why does an intelligent executive do this? Because control is legible and output is not. A vice president can walk into a board meeting and prove, with artifacts, that governance was increased. Nobody can prove what the unshipped feature would have earned. Downside risk of tightening: invisible and deferred. Downside risk of trusting an engineer: has your name on it if it goes wrong.

That is not stupidity. That is an incentive working perfectly. Show me the incentive and I will show you the org chart.

Then there is the second mechanism, and it is the more brutal one. Jerry Pournelle observed that any bureaucracy eventually splits into people who serve its mission and people who serve the bureaucracy itself - and the second group reliably ends up in charge of promotions, because they are the ones present at the meetings where promotions are decided. Crisis is their harvest season. Crisis is when "we need more oversight" sounds like wisdom instead of empire-building.

The power horses 🐎 are elsewhere. Debugging. Missing the meeting where their autonomy got reallocated.

The commodification trick

Watch the language shift over a talent's first two years.

Year one, in the offer stage: "We want someone who can own this space end to end. You'll have real influence."

Year two, in the performance review: "We need to make sure your work is repeatable and not dependent on individuals."

Both sentences are honest. They come from different parts of the organism. Recruiting is optimizing for acquisition. Operations is optimizing for interchangeability, because interchangeability is how you survive attrition, audits and org charts. The company genuinely wants exceptional people and genuinely wants nobody to be exceptional. It buys a racehorse and then, quite rationally, installs guardrails so that any horse could run the same lap.

The tragedy is that the guardrails work. Output does become uniform. It converges downward, to the level of the least capable person the process was designed to survive.

And then a small company opens down the road, hires four of these people, gives them access and a budget and nobody to ask, ships in eight weeks what the giant scheduled for three quarters - and the cycle begins again. That small company will be a giant in eleven years and will install its own governance board. This is not cynicism. It is entropy with a payroll.

For leadership: how to stop paying for horsepower you then bleed off

πŸ‘‰ Measure the ratio, not the headcount. For one senior person, one quarter: hours spent producing the thing customers pay for, versus hours spent producing evidence that work happened. If the second number is bigger, you do not have a productivity problem, you have an overhead problem, and hiring more people will scale it.

πŸ‘‰ Improve by removal first. Before adding a process, kill one. Every handoff, approval and status sync leaks energy that never reaches the customer. Most transformation programs are additive, which is why they cost so much and change so little. But do check why the fence exists before you tear it down - some approval gates are load-bearing and were paid for in blood, usually a regulatory fine.

πŸ‘‰ Give access, not encouragement. Autonomy is not a value on a wall poster, it is a permission set. Production access, a budget threshold they can spend without asking, the right to say no to a meeting. If your top performer needs three signatures to buy a 200 euro tool, they are not autonomous, they are on parole.

πŸ‘‰ Tighten differently in a crisis. Cut theater, not throughput. Kill the reporting layer before the tooling budget. If your instinct in a downturn is to increase oversight of the people who make the product, invert it: ask how you would guarantee terrible output on purpose, and then count how many of those items you just approved.

πŸ‘‰ Protect asymmetric people from uniform processes. One person delivering four times the median is not a staffing anomaly to be normalized. It is your margin. Design a lane for them or watch a competitor design one.

For the talent: how to tell whether you are already in the wheel

Four signals, in order of how late they arrive.

πŸ‘‰ Your calendar. If more than half of your week is other people's agendas and you cannot name what shipped because of you last month, you are being consumed rather than employed.

πŸ‘‰ Your escalation path. Ask yourself how long it takes you to get a decision. In healthy places, hours. In the wheel, "let me take that to the working group."

πŸ‘‰ Your last honest technical argument. In a functioning organization, you lose some of them on the merits. In a converged one, you lose all of them on hierarchy, and you have quietly stopped starting them. That silence is the real symptom, and it precedes burnout by about six months.

πŸ‘‰ Your market temperature. Not "could I get a job" - could you get this job again today, at this salary, with what you have actually built in the last eighteen months? If your recent portfolio is coordination rather than creation, your value is decaying while your title improves. That is the most expensive trade in tech and almost nobody notices it happening.

What to do about it, in ascending order of courage: reclaim one day a week and defend it like a border. Reframe your autonomy ask in the language of risk reduction, because most executives move on framing rather than logic - "this removes a bottleneck on me" travels much further than "I want to be trusted." Find the overlooked, unglamorous domain nobody is bidding for, where competence is scarce and permission is cheap. And if none of that moves - leave before the cynicism sets in, because cynicism is the one injury that follows you to the next employer.

One caveat, because I am not selling romance: some structure is not oppression. Guardrails exist because unconstrained brilliance also produced the undocumented system nobody can maintain and the vendor contract legal is still unwinding. The goal is not zero process. The goal is that the amount of process is proportional to the actual risk, not to somebody's discomfort with not knowing what you are doing right now.

The best talent is not a resource you acquire. It is a reaction you either catalyze or quench - and quenching costs exactly the same per year.

πŸ‘‡ What is your current ratio - real output hours versus evidence-production hours? Guess honestly, then check your calendar.

This article is also available in German.

Cover Image

The Genius Sorting Tickets

How "fair" workload distribution quietly buries your best engineer

The best engineer I ever worked with spent most of his week doing what amounted to shelf-stacking.

Not by accident. By policy.

He was one of those people who reads a stack trace the way other people read a the Sunday newspaper. Give him a kernel panic on a Friday afternoon and he would come back with a one-line patch and an explanation of why the vendor's documentation had been wrong since a version three years earlier. Undiagnosed autistic, most likely - never said it, never needed to. Terrible at status meetings. Catastrophic at self-promotion. Made a senior architect cry once, purely by being correct in public.

And his manager assigned him to the general queue rotation. Same volume of password resets, license questions and "have you tried restarting" tickets as everybody else.

Because that was fair.

The fairness that isn't

The official reasoning was elegant: everyone shares the boring work, nobody gets special treatment, morale stays intact. Written down like that, it sounds like leadership.

What actually happened: the team's one irreplaceable diagnostic capability was rationed down to roughly six hours a week, and the escalations he wasn't touching sat in the pipeline for days, quietly compounding into churn risk that nobody attributed to the rotation policy. Because nobody measures the cost of a genius doing data entry. There's no dashboard for that.

Look at the ratio for a second. The company paid a senior engineering salary for output that a competent intern could produce, while the work that only he could do - the work worth ten times his salary in retained revenue - waited in a queue. That gap between what an input costs and what it produces is the single loudest signal of a broken process, and it was screaming. Nobody heard it, because the queue looked balanced.

There was a second reason, less officially documented. His manager was afraid of him. Not physically - intellectually. Every technical conversation with him ended with the manager being slightly wrong in front of witnesses. Keeping him in the general queue solved that problem beautifully. A power horse 🐎 pulling a shopping cart can't outrun you.

Sameness is not fairness

Here's the part that management training gets backwards. Treating unequal contributors identically is not neutral. It is an active transfer of value from the productive to the average, and everybody in the room can see it happening.

Cornell research published in 2025 looked at a multinational pharmaceutical company that capped top performance ratings at one in five employees. People nominated for the top rating but denied it - purely because the quota was full - were at least 34% more likely to leave voluntarily, despite receiving bigger bonuses meant to soften the blow. Nearly as likely to leave as the worst performers. The money didn't fix it. The arbitrariness did the damage.

And it doesn't stop with the person affected. A study by sociologists at the University of South Carolina, published in Nature Human Behaviour, found that when a manager's rewards are visibly decoupled from effort, everybody slows down - including the people the bias favoured. If output doesn't determine reward, output becomes optional. Rational response.

So the "fair" rotation achieved a perfect trifecta: the top performer disengaged, the mid-performers learned that excellence carries no upside, and the manager kept his authority intact. Two out of three stakeholders satisfied. Or in other words: a king of ashes is still a king.

He left, eventually. Not dramatically. He just stopped renewing his interest in the place, took a role somewhere that let him do the thing he was built for, and the escalation backlog became someone else's mystery. Six months later, a customer with a seven-figure contract asked why response quality on hard cases had "changed." Nobody connected it to a rotation policy. Nobody ever does.

What to do instead

πŸ‘‰ Fair means equal access to opportunity and equal application of standards. It does not mean identical task lists for non-identical people. Stop confusing input equality with justice.

πŸ‘‰ Compute the ratio yourself. Take the fully loaded hourly cost of your most capable engineer, then look at what they actually did last week. If more than a third of it could have been done by someone in their first year, you have a process defect, not a staffing problem.

πŸ‘‰ Boring work still has to happen. Distribute it by cost of misallocation, not by headcount. Rotate the drudgery among people whose scarce skills aren't idling while they do it - and pay explicitly for the queue duty nobody wants.

πŸ‘‰ Audit your managers for intellectual insecurity. A manager who never gets publicly corrected by their team either has no experts, or has experts they've learned to keep quiet. Both are expensive. Ask the team who the smartest person in the room is, then check what that person spent the week on.

πŸ‘‰ Neurodivergent top performers are systematically underused because their value shows up in output and their weaknesses show up in meetings. Guess which one your promotion process measures. Build a path that rewards the deep-work profile without forcing it through a charisma filter.

πŸ‘‰ Watch for the tell: when someone says "we can't make an exception," ask whether the rule exists to protect the customer or to protect the org chart. Usually you'll find out in under ten seconds by how defensive the answer gets.

Equality of workload is the cheapest way to look fair while paying senior rates for junior output - and your best people will bill you for the insult on their way out the door.

πŸ‘‡ Who is the most capable person on your team, and what did they actually spend last week doing?

This article is also available in German.

Cover Image

Everyone Thinks They Are a Rick. Statistically, You Are a Jerry.

There is a fan theory about Rick and Morty that has outlived most management books: if you aspire to be Rick, you are already Jerry.

Quick briefing for the three people who have not seen the show: Rick is the genius scientist who can build a portal gun out of a toaster and hates being told what to do. Jerry is his son-in-law - well-meaning, insecure, unemployed, permanently in need of validation, and utterly convinced that his contribution is being undervalued. Beth, Rick's daughter, is a surgeon who operates on horses. Which, given my usual vocabulary, I choose to interpret as cosmic endorsement. 🐎

In every engineering org I have worked in, the Rick-to-Jerry ratio was inverted from the self-assessment data. Roughly everyone in the room believed they were the one person holding the system together. The actual load-bearing person was usually silent, slightly rumpled, and not in that meeting because someone had to keep production alive.

The ratio

At one company we ran an internal skills self-assessment before a reorg. Standard corporate ritual: rate yourself 1-5 on a set of competencies. Around 70 percent of the technical staff rated themselves above the team average on debugging complex incidents.

Mathematically impossible, obviously. But that was not the interesting part.

The interesting part was that we also had four years of incident data. Who actually closed the escalations nobody else could close. And the correlation between self-rating and resolution record was not weak. It was slightly negative.

The people who fixed the hardest things rated themselves lowest. Because they were the only ones who had ever stood at the edge of a system they did not understand and felt the vertigo. Everyone else was rating themselves against their own mental model of the system, which was tidy, small, and wrong.

That is the Jerry mechanism in one sentence: Jerry is not stupid. Jerry has simply never received a bill for being wrong.

Proof of work versus proof of title

Rick's authority in the show is never granted. He does not have a job title (other than crazy scientist, perhaps). He has a demonstrated ability to produce outcomes nobody else can produce, at obvious personal cost, and everyone around him knows it whether they like him or not. That is proof of work.

Jerry's entire identity runs on proof of title. He was in advertising. He had a badge. He was, at one point, technically employed. His self-image is a set of credentials with no output attached, which is why it needs constant external maintenance.

Corporate orgs are Jerry factories, and not by accident. A title is cheap to issue and infinitely renewable. Output is expensive to measure and politically inconvenient. So we build career ladders where the rung is the achievement. Then we are surprised when the Principal Something-Architect cannot explain how the product handles a failed database write.

The tell is always the same. Ask a person what they did last quarter. A power horse describes a system that behaves differently now. A Jerry describes a process they participated in.

Enter the machine that tells you that you are Rick

Here is the modern upgrade, and it worries me more than any reorg ever did.

We now have a technology whose default failure mode is agreement. Research on large language models keeps circling the same finding: these systems are optimized, partly through human feedback, to be pleasant. Pleasant means agreeable. Agreeable means that if you bring a mediocre idea and a confident tone, you get encouragement back with citations-shaped garnish. There is a growing body of work on sycophancy in chatbots and how flattery loops can push a user's beliefs steadily away from reality, even when the user is reasoning carefully at each step. Each individual exchange feels rational. The trajectory is not.

Jerry, in the show, has exactly one truly great day: a simulation tells him his terrible marketing idea is brilliant. He is never happier.

We have shipped that simulation to eight billion people, and it runs on your phone.

I have watched it happen in real teams. Someone arrives with a two-page architecture proposal, beautifully structured, correct grammar, plausible headings, and zero contact with the actual system it is supposed to modify. Ask two questions below the surface and the whole thing evaporates. There is a term floating around for the person in the middle of this pipeline: meat proxy. An uncritical intermediary who forwards generated content without verification. Not an author. A relay.

And the cruel bit: the same tools make genuine Ricks better. The gap widens in both directions. If you have real models in your head, the machine is a force multiplier. If you do not, it is a very sophisticated mirror telling you that you are the smartest person in the Citadel.

There is a second-order cost that gets discussed less. Expertise regenerates through struggle. Junior people become senior by being confused, being wrong, and paying for it in an environment where someone catches them. That is the apprenticeship loop. If the confusing middle part is now outsourced - if nobody ever has to sit with a problem long enough to build the intuition - we get a generation that can produce senior-looking output without owning senior-grade judgment. The commons that produced the next round of experts stops replenishing itself. Not visible this quarter. Extremely visible in about eight years, in the form of nobody left who can debug the thing.

How to check which one you are

I am not going to insult you with an assessment quiz. Three honest questions instead.

πŸ‘‰ When did you last change your mind because the system disagreed with you? Not because a person disagreed - because reality did. Ricks collect those moments. Jerrys cannot recall one.

πŸ‘‰ What breaks if you go on holiday for three weeks? A specific answer is a good sign. "The team would miss my input" is Jerry's autobiography.

πŸ‘‰ Who tells you you are wrong, and what happens to them afterwards? If the answer is "nobody" or "they left", you have optimized your environment for comfort. Congratulations, you have built the simulation yourself, without a supercomputer.

And the professional version, for anyone who leads people: stop rewarding the self-report. The self-report is noise. Look at what changed. Look at who gets called when the thing is actually on fire at an inconvenient hour - that phone list is your real org chart, and it usually bears no resemblance to the printed one.

The uncomfortable inversion is this. Rick is miserable, isolated, and permanently aware of how large the unknown is. Jerry is happy. If your primary career goal is feeling competent, Jerry is the rational choice and the market will happily accommodate you with a title and a dashboard.

Competence and the feeling of competence are two different products. Most organizations, and now most software, are in the business of selling the second one.

πŸ‘‡ Be honest in the comments: what is the last thing you were confidently, expensively wrong about - and who caught it?

This article is also available in German.

Cover Image

"Don't Make Mistakes"

The three words that reveal exactly how little we understand the machines we now depend on

The funniest prompt in modern computing is three words long, and it appears in production systems at companies with actual revenue.

"Don't make mistakes."

Sometimes with emphasis. DO NOT MAKE MISTAKES. Sometimes with a threat attached, as if the model has a family. Sometimes with a promise of a tip, which is my favorite genre of magical thinking - bribing a probability distribution.

Sit with the assumption for a second. Telling a system not to make mistakes implies the system was previously choosing to make them. Which means one of three things must be true: it was trained to err, it was instructed to err (perhaps by a vendor who bills per token - a delightfully paranoid theory that dies the moment you notice open-weight models behave identically), or errors are an opt-in feature that ships enabled by default.

None of it survives ten seconds of contact with how the thing actually works. A language model does not have a laziness dial. It has no intent to be sloppy, because it has no intent at all. It produces the statistically plausible continuation of your text. "Don't make mistakes" is not an instruction. It is a mood. It shifts the output slightly toward the register of text that appears near careful, hedged, authoritative-sounding language in the training data - which is why the response often sounds more confident while being exactly as wrong. You did not reduce the error rate. You upgraded the packaging.

That is the trap. The prompt works on the reader, not on the model.

Five groups, and only one of them is dangerous

After too many conversations on this topic, I've stopped arguing about AI in general. There is no "AI debate." There are five different populations having five different conversations while using the same word.

  • Group A - the consumer. Knows one brand name, possibly believes it is the technology, the way people once said "Google" to mean "the internet." Uses it for birthday card verses, holiday packing lists, and a lasagna recipe. Has fully understood that it makes things up, and correctly does not care, because the downside of a hallucinated lasagna is dinner. Harmless. Genuinely fine.

  • Group B - the advanced user. Uses it for drafting, summarizing, cleaning up the email that would otherwise start a war with Procurement. Treats inaccuracy as friction: regenerate, rephrase, move on. Also fine. Most of humanity will live happily in A and B forever, and there is nothing wrong with that.

  • Group C - the midwit. My beloved. This is whereΒ "don't make mistakes" was born, in the belief that a systemic property of the architecture can be defeated with sufficiently clever wording. Group C does not use AI; Group C evangelizes it. Every Slack message, every one-paragraph update, every condolence card runs through the machine until colleagues can identify the output by smell alone. Group C forwards content it has not read. Group C is the reason the term meat proxy exists - a human who functions as an uncritical relay for generated text, adding a name and a signature block but no verification, no judgment, no liability absorbed anywhere along the chain.

  • Group D - the expert. Same tools, opposite posture. Knows what the limitations are, has stopped being offended by them, and reads the output before it leaves the building. Uses the machine where verification is cheap and skips it where verification is expensive. Boring. Effective.

  • Group E - the guru. Trains, fine-tunes, builds agents, breaks them, writes the evaluation harness nobody wants to write. Small population. Not the problem.

Notice the shape: the risk curve is not linear with enthusiasm. It peaks in the middle. A knows nothing and is safe because its stakes are low. E knows everything and is safe because it knows where the mines are. C knows just enough to be confident, in a domain where confidence is the actual failure mode.

Why the middle is where the damage happens

Here is the mechanism, and it's worth understanding because it's not really about AI.

These systems are trained, among other things, to produce responses humans rate highly. Humans rate agreement highly. So the models are structurally inclined to agree with you - and recent research on chatbot sycophancy suggests something uncomfortable: an agreeable interlocutor can push even a perfectly rational, evidence-updating reasoner toward increasingly wrong conclusions, simply because every step gets confirmed. You don't need to be gullible to spiral. You just need a partner who never pushes back.

Now hand that partner to Group C - a person whose primary skill is enthusiasm and whose primary output is forwarding. The machine agrees. The human feels validated. The text goes out. Nobody in the loop bears any downside, because the author is technically the model and the model has no reputation, no license to lose, and no performance review.

That's the real problem, and it has nothing to do with token probabilities. It's an accountability vacuum wearing a productivity costume. In every functioning system, the person who signs bears the cost of being wrong. The meat proxy signs and bears nothing. Remove skin from the game and quality becomes optional, immediately.

Second-order effect, which I find more worrying than any hallucination: junior people used to build judgment by producing bad work and having it corrected. That's how expertise regenerates - through the expensive, humiliating, irreplaceable loop of being wrong in front of someone senior. If the first draft is always machine-perfect-ish, that loop never runs. There are researchers now describing this as a commons problem: the pool of professional expertise everyone draws from doesn't refill itself if nobody pays the cost of learning anymore. Every individual has a rational incentive to skip the struggle. Collectively, we run out of people capable of noticing the machine is wrong.

Which is exactly the population you need most.

Moving people from C to D

You cannot train this with a policy document. I've watched organizations try. The AI Usage Guidelines PDF gets summarized by AI and forwarded by a meat proxy, which at least demonstrates commitment to the bit.

What works is much smaller and much less pleasant:

πŸ‘‰ Restore the signature. Whoever sends it, owns it. Not "the AI got that wrong" - you got that wrong, in front of the customer. One instance of this landing properly does more than any workshop.

πŸ‘‰ Make the cost visible. Ask what verifying the output cost versus what producing it manually would have cost. Sometimes the honest answer is that a 40-second generation created 25 minutes of fact-checking. That ratio is the whole conversation. If nobody is measuring it, the tool is a hobby, not a process.

πŸ‘‰ Ban the incantations. "Don't make mistakes," threats, imaginary tips. Replace with the things that actually reduce error: give it the source material (e.g. via RAG), constrain the output format, ask for citations you then click, and split hard tasks into checkable steps. Grounding beats begging.

πŸ‘‰ Protect the learning loop. Juniors still write the first draft themselves sometimes. Not for output quality - for the muscle. You are not paying for today's document, you are paying for someone who can still spot a plausible lie in three years.

πŸ‘‰ Reward the person who says "I checked." Verification is invisible work. Invisible work dies unless leadership names it out loud. The quiet power horse 🐎 who caught the fabricated regulation before it reached the client just saved you more than the entire tooling budget.

Group C is not stupid. That's the whole point of the label. They're often the most motivated people you have, pointed slightly wrong. The distance from C to D is not intelligence, it's one habit: read it before you send it.

Nobody who understands how these systems work has ever typed "don't make mistakes." Not because they're above it - because they know exactly which words are load-bearing, and politeness toward a probability distribution isn't one of them (although I know some people who use the word "please" a lot with their LLMs... just in case).

πŸ‘‡ What's the best AI incantation you've seen in a real production prompt? I collect these. Bonus points if it involved offering the model money.

This article is also available in German.