<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Hyperbliss</title>
        <link>https://hyperbliss.tech/</link>
        <description>Stefanie Jane's personal blog, lab experiments, and more.</description>
        <lastBuildDate>Wed, 07 Oct 2026 20:04:30 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <image>
            <title>Hyperbliss</title>
            <url>https://hyperbliss.tech/images/logo.png</url>
            <link>https://hyperbliss.tech/</link>
        </image>
        <copyright>All rights reserved 2026, Stefanie Jane</copyright>
        <item>
            <title><![CDATA[Get Loopy! Act, React, Iterate Until We Love It]]></title>
            <link>https://hyperbliss.tech/blog/loop-engineering</link>
            <guid isPermaLink="false">https://hyperbliss.tech/blog/loop-engineering</guid>
            <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[An agent is a week into rewriting my memory system's retrieval code. It built a deterministic scorer, ran dozens of experiments, kept the good ideas and ditched the failed ones, all from one prompt I wrote at the start. Loop engineering is the third act of agentic engineering: the agent acts, observes the result, feeds it back into its own context, and goes again until the goal is true. Here's two years of building these loops, with receipts.]]></description>
            <content:encoded><![CDATA[
For the past week, an agent has been rewriting the retrieval system inside my memory engine. So far it's made a deterministic scorer, run dozens of experiments, kept the good ideas and ditched the failed ones. I wrote one prompt at the start and have made a few calls along the way (insert coin to continue). The loop is doing the rest, and it isn't even close to done.

![[wide] A developer stands back from their desk holding a coffee, watching a ring of holographic terminal panels orbit above the chair, each panel feeding a ribbon of light into the next in a continuous glowing cycle, with a single thin thread of light connecting the ring to the developer's hand.](/images/blog/loop-engineering-hero.webp 'The loop runs itself. The steering thread stays in your hand.')

At 3:30 every morning, another model wakes up to a standing instruction I wrote once and never send. It reads what my agents captured yesterday, decides what's worth keeping, promotes the confident findings into a knowledge graph, and routes the uncertain ones to me. Nobody is awake. The system is prompting itself.

And in one five-day stretch last month, a PR babysitter I built ran a thousand agent sessions against my open pull requests. Poll, classify, dispatch, fix, learn, repeat. Tiny models for watching, frontier models for fixing. As agents and humans review code, we react immediately and update PRs automatically.

This article will show you a few techniques I use every day for large scale agentic engineering.

## 🌀 The Name Finally Caught Up

In June the term **loop engineering** landed. Peter Steinberger [put it in twelve words](https://x.com/steipete/status/2063697162748260627): "you shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents." Addy Osmani [named the practice](https://addyo.substack.com/p/own-the-outer-loop) a day later and gave it an anatomy; O'Reilly [republished the essay](https://www.oreilly.com/radar/loop-engineering/) two weeks after that, which is how you know a term has officially arrived. Boris Cherny, who created Claude Code, summed up the vibe: "I don't prompt Claude anymore."

The lineage is clean. Prompt engineering is about what we tell the model. [Context engineering](/blog/context-engineering/) is how we can make the model create and improve those prompts. Loop engineering is about what happens _between_ runs: the trigger, the feedback signal, the state that carries forward, and the rule that decides when it stops.

This post is the third act of a trilogy I have apparently created! [In January](/blog/context-engineering/) I wrote about engineering the context. [In May](/blog/how-i-ai/) I wrote about the operating system underneath my prompts: a contract, a skill library, a memory graph. This one is about what happens when that operating system starts running itself.

One receipt for how convergent this moment is: while mining my own session archive for this post, an agent found the original wish for my memory consolidation loop, timestamped April 4. Verbatim: "i really want a dream mode for claude code. like something that reviews all our conversations regularly." Two months before anyone named the practice. If you've been working seriously with agents, you've been converging on the same ideas. The name doesn't unlock anything; it just gives us a shared vocabulary.

There is a bit of craft here.

## ⚡ Act → React: The Core Mechanic

At their core, every agentic loop is the same three-beat cycle:

1. **Act.** The agent does something real: edits code, runs an experiment, drives a browser, opens a PR.
2. **Observe.** The world answers with a signal: a test result, a metric, a screenshot, a review verdict, CI going green or red.
3. **React.** The loop, not a human, feeds the signal back into the agent's context, and the next act is shaped by it.

```mermaid
flowchart LR
    A["⚡ Act"] -->|"edit · run · drive · ship"| O["👁 Observe"]
    O -->|"tests · metrics · screenshots · CI"| R["🔄 React"]
    R --> G{"stop condition true?"}
    G -->|"no, go again"| A
    G -->|"yes"| D["💎 done"]
    S[("state carrier: ledger · graph · git")] -.-> A
    R -.-> S
```

That third part is really the thing. A prompt is a one-shot context delivery: you curate what the model sees, it acts, done. A loop is a _context engine_. Every iteration curates what the next iteration sees, automatically, from the results of the last one. The agent is feeding itself. Then we think about the loops around the loops.

Which means the quality of a loop is exactly the quality of four components:

**The feedback signal.** What does the loop observe after acting? A signal that's fast, deterministic, and dense beats an expensive vague one every single time, because the loop consumes it on every iteration. Tests are great signals. Typed metrics are great signals. "The page looks better than it did before" is not a signal, but "the page's style matches the style guide perfectly" is.

**The state carrier.** What survives between iterations? The model's context window doesn't. It fills, it compacts, it's gone when the session ends. Loops that compound carry state outside the model: a findings ledger, a journal file, a task board, a knowledge graph, git history itself. The agent forgets; the repo doesn't.

**The stop condition.** When is it done? "When it's good" is not a stop condition. A testable predicate is: tests green, findings count at zero, metric above threshold, budget exhausted, human says ship.

**The safety rails.** Budgets, cooldowns, whatever you need to make sure everything stays on track and nothing goes crazy and lights your wallet on fire.

Everything that follows is the same four components rearranged.

## 🧪 The Eval Loop: An Agent Running Real Science

The best loop I've run this year is still running while I write this: an experiment campaign against [Sibyl](https://github.com/hyperb1iss/sibyl), my memory engine, on the LongMemEval v2 benchmark, because v1 is saturated to the point that a good score is table stakes rather than signal. A week in, it already shows what act → react looks like when the feedback signal is engineered instead of vibed.

Our goal: improve our benchmark scores, without cheating.

The agent's first move was the one that made everything else work: **it built itself a super simple scorer.** Instead of asking an LLM judge "did this retrieval look good?" on every iteration (slow, noisy, expensive), it wrote a script that checks whether the known-correct answer literally appears in the assembled context. That check is cheap, exact, and instant, and every configuration got a real number in seconds:

```plaintext
exposure 14/23 ( 60.9%)  phrase-hit  66.7%  ctx  36.9K avg
exposure 15/23 ( 65.2%)  phrase-hit  69.7%  ctx  35.8K avg
exposure 16/23 ( 69.6%)  phrase-hit  69.7%  ctx  47.8K avg
```

![A dark laboratory with eleven tall glass cylinders in a row, most glowing with faint fading embers while one blazes bright with cyan and pink light, its energy flowing through conduits back into the machinery feeding all the cylinders.](/images/blog/loop-engineering-eval.webp 'Eleven experiments, five retired with receipts, one survivor. The loop feeds the win back in.')

Then the loop proper. The plan said six experiments; the results kept suggesting new ones, and the list has grown to dozens so far, because that's what loops do. Each experiment ran in the background and could crash and resume where it left off, and the agent slept until a results file (or a stack trace) showed up to react to.

The surprise came from scoring two things separately: whether search _found_ the right memories, and whether the layout actually got them onto the page. Search was already finding 82.6% of the answers; the layout only got 65.2% of them in front of the model. The bottleneck wasn't search at all. It was page layout. I would not have guessed that, and neither did the agent: the loop measured its way there.

One change looked like a clear win. Instead of shipping it, the loop did the thing that separates engineering from enthusiasm: it ran the same experiment three more times and compared results question by question. The win evaporated: the average had been sitting still while individual questions churned underneath it. Single runs lie. Five plausible improvements were retired this way, each with receipts.

We got as far as we could without an LLM, but the idea survived: have a tiny, fast model write a short digest of every memory as it's stored, about fifteen cents for the whole corpus. Replicated effect: **+6.67 points, nine questions better, two worse.** Then it flunked its first test on a different kind of data, the agent worked out why (the digest was reading fields that didn't exist there), fixed it, and passed the rerun.

Dozens of commits have landed and the campaign is still running as this post goes up. The eval loop is rewriting the system it's evaluating, live.

The transferable lessons, because your version of this won't be a memory benchmark:

- **Build the cheap scorer first.** One hour spent making feedback deterministic pays back on every iteration. LLM judges belong at the _end_ of a loop, not inside it.
- **Make iterations resumable.** Experiments that can crash and resume turn an overnight failure into a delay instead of a loss.
- **Replicate before you believe.** N=1 results are noise wearing a costume.
- **Retire hypotheses with receipts.** A loop that only confirms is a yes-machine. Five documented rule-outs bought the credibility of the one win.

## 🎭 Visual Verification: The Loop Got Eyes

The loop upgrade that still feels like magic: agents can _see_ now, and seeing closes feedback cycles that used to require a human on every iteration.

The modern setup: the agent drives the running app in a real browser, the rendered surface itself rather than the test suite. It navigates, clicks, fills forms, and screenshots. The screenshot is the observe beat. The agent compares what it sees against the acceptance criteria, fixes, re-renders, looks again. For UI work I'll write something like "the cards align to an 8px grid, the hover state doesn't shift layout, dark mode holds contrast" and let the loop iterate itself there, screenshot by screenshot. A year ago this was "make the change, describe it to me, I'll look." Now the agent is the one looking, and it catches the 2px misalignment I would have missed.

The same pattern runs end-to-end tests as a loop signal: drive the app, assert on what actually rendered, feed failures back as context. And it composes with everything else in this post. A review loop where the reviewer _also_ drives the app catches a class of bug no diff-reader ever will.

One rule keeps these loops honest, learned the expensive way and now written into my skill library: **perf and visual loops need an objective signal before iteration two. Loops judged by the next screenshot become random walks.** The screenshot is the sensor, not the standard. Pin the standard first (named acceptance criteria, a reference design, a metric, a golden snapshot) and let the screenshots measure against it. An agent iterating toward "looks better" will happily wander forever, each iteration convincingly justified.

When the surface is one the agent can't observe, the first move is building the connector, because most "unobservable" surfaces are just missing a driver. Terminals get automation harnesses like [ghostty-automator](https://github.com/hyperb1iss/ghostty-automator) or cmux that drive keystrokes and read the screen back. Machines beyond the laptop get an agent over SSH. A drawer full of Android devices becomes a sensor array with an MCP server like [droidmind](https://github.com/hyperb1iss/droidmind) pushing input and pulling screenshots from every one of them. Each connector converts a blind spot into a feedback signal, permanently, for every loop you build afterward.

And when the surface genuinely can't be instrumented (how a terminal _feels_, the light coming off actual hardware), the loop still doesn't end. The human becomes the sensor: hand them a pre-registered expected outcome and a discriminating tell, and you're still running act → react. You've just got a slower sensor with better taste.

## 🔮 Review Loops: Convergence You Can Graph

The loop I run most often is cross-model review to convergence, and the state carrier is what makes it more than "ask another model twice."

The shape: one model produces an artifact. A spec, a plan, a diff. A different model reviews it adversarially. The findings go into a **ledger** that travels between rounds: every finding verbatim, the claimed fix, the commit SHA of the fix. The next round's reviewer gets the ledger and two jobs: verify each claimed fix actually landed, and hunt for new problems the fixes introduced. The artifact and its ledger _are_ the loop's memory. The agent is feeding itself its own review state.

You track convergence numerically, and the numbers tell you things. A recent infrastructure plan of mine converged over twelve rounds:

```plaintext
17 → 9 → 8 → 4 → 3 → 2 → 1 → 1 → 4 → 1 → 1 → PASS
```

![A surreal night landscape where a staircase of glowing monoliths descends toward a calm cyan horizon, one late step rising unexpectedly higher before the final flat plane, with two small figures of light walking down together.](/images/blog/loop-engineering-convergence.webp '17 → 9 → 8 → 4 → 3 → 2 → 1 → 1 → 4 → 1 → 1 → PASS. Round nine is why you graph it.')

See round nine? Four new findings after two quiet rounds, because the round-eight fixes broke new ground. Convergence isn't monotonic, which is exactly why you graph it instead of trusting your sense that "it's probably fine now." Another spec went 19 → 6 → 4 → 2 → 0 in five rounds. When the trend stalls or oscillates, the loop is telling you the remaining findings are matters of judgment, and judgment is the human's job.

The community has been circling the simplest form of this for a while: implement, review, fix, repeat, sometimes with a hook that re-feeds the same prompt until the work is done. Fine as far as it goes. But the naive version has a failure mode that only shows up in production use: **not all review loops should stop the same way**, and treating them uniformly either burns tokens or ships defects. What I've converged on:

- **Code review: cap the re-litigation, not the rounds.** Three visits to the _same finding_ means you're in an argument, not a review. But rounds that keep surfacing new confirmed defects keep going.
- **Spec review: iterate until you love it.** Spec defects are the most expensive class of defect there is, so convergence is the only exit. My standing instruction to the loop is literally "iterate until we love it."
- **Fix passes: pre-declare a file budget.** Review-fix loops are a monotonic scope ratchet; every round wants to touch a few more files. I've watched one hit 8 rounds and 63 files before we re-anchored it to six. Now the budget gets declared before round one.

Two independence rules, both non-negotiable. The implementer never self-assigns PASS, because the model that wrote the code is too kind grading its own homework, and so is the same model re-reading it. And a verdict is pinned to a SHA: any commit after the reviewer's pass voids the pass. Warm-resuming the same reviewer speeds up fix-round convergence; a fresh-context reviewer does final certification precisely because it inherits nothing.

## 🛠️ The Babysitter: 1,785 Sessions in Five Days

Vigil is a PR babysitter I built this spring: an agent system that watches my open pull requests and does whatever they need next. In its first real production window it ran 1,785 agent sessions in five days, every one of them against my real PRs.

The loop: pollers watch GitHub (my PRs every 30 seconds, the wider radar every 60). State changes get classified by a five-state machine (hot, waiting, ready, dormant, blocked) and classified events go to a cheap triage agent with a five-cent budget that routes to specialists: a fixer with a fifty-cent budget and 15 turns, a responder for review comments, a rebaser, an evidence gatherer. A fixer pushes a commit, the next poll sees fresh CI, triage re-routes. The loop continues until the PR hits READY.

```mermaid
flowchart LR
    GH[("GitHub")] -->|"poll 30s / 60s"| SM{"five-state machine"}
    SM --> T["🧭 triage · 5¢"]
    T --> FX["🔧 fixer · 50¢ · 15 turns"]
    T --> RS["💬 responder"]
    T --> RB["🌿 rebaser"]
    FX -->|"push → fresh CI"| GH
    RS --> GH
    RB --> GH
    T -.->|"3 strikes · irreversible"| H["👤 you"]
    GH -->|"on merge"| LN["📚 learning agent"]
    LN -.->|"patterns → future prompts"| T
```

What the architecture diagram doesn't show: **the loop logic is maybe a fifth of the code. The rest is safety rails.**

- Per-agent turn caps and dollar budgets, per run.
- Event dedupe with a five-minute window, plus a 45-second cooldown per event type, because GitHub will happily tell you the same thing four times.
- A concurrency gate so two agents never work the same PR simultaneously.
- Prompt fingerprinting that flags when an agent is about to ask the same question it asked five minutes ago. That's the observability hook for loop detection.
- Escalation: three consecutive failures on the same thing and it stops looping and pings the human.
- Auto-approve for everything except the irreversible set: push, merge, branch delete. Reversible actions flow; irreversible ones queue for me.

![A lantern-like drone hovers over a dark grid of glowing cards, sweeping a cyan beam across them; cards glow in five distinct colors, tiny repair drones tend to some, and one card rises out of the grid on a beam of light.](/images/blog/loop-engineering-vigil.webp '1,785 sessions in five days: poll, classify, dispatch, learn. The irreversible stuff rides the beam up to you.')

There's also a second, quieter loop stacked on the first. On every merge, a learning agent extracts patterns into a knowledge file ("this reviewer usually asks for type annotations on public APIs, confidence 0.70"), and new patterns start at 0.50 confidence and gain +0.10 on each reconfirmation. The whole file gets injected into future triage and fix prompts. Every PR makes the next one smoother. The loop's output has become the loop's context, and that's the move that compounds.

To be fair, you no longer need to build any of this yourself just to get your PRs babysat. Codex and Claude are both solid PR watchers out of the box now, and with newer models you can hand one an entire stack: a PR goes up, a bot review posts feedback, human engineers post theirs, and the agent reads it all as it lands, updates the PR, and replies to the comments with a summary of what changed. What the off-the-shelf version doesn't ship is everything around the loop: the budgets, the dedupe, the learning file, and the queue of irreversible actions that waits for a human.

If you build one loop this year, build a babysitter for something you already poll by hand that nobody has productized yet. The design pressure is fantastic: every safety rail above exists because its absence produced a specific, memorable mess.

## 🌙 Loops That Feed the Loops

The layer above all of this is where it gets properly recursive: loops whose output is the _context machinery itself_.

**Consolidation runs nightly.** Sibyl's dream cycle wakes at 3:30am, reads what my agents captured during the day, and applies a hard policy: findings it's confident about promote into the knowledge graph automatically, anything it's unsure of queues for me, and duplicates and stale entries get archived. Every agent's tomorrow-context is being curated by an unattended model run tonight, gated by policy.

**Forgetting is a loop too.** Every recall an agent makes stamps usage counters on the entities it retrieved. A nightly decay job consumes those counters: memory that gets used survives, memory that doesn't fades and eventually archives. Reads are the training signal for retention. The knowledge graph is under continuous selection pressure from the swarm's actual behavior. I don't garden it; the usage gardens it.

![A bioluminescent night garden of glowing networked blossoms under a pale moon; threads of light make some blossoms flare bright while untouched ones fade to ash, as a translucent moonlight figure prunes and grafts the vines.](/images/blog/loop-engineering-dream.webp 'Recall feeds the bloom, silence feeds the ash. The garden gardens itself at 3:30am.')

**And once a season, the big one.** In July I pointed one prompt at my entire session archive: "review the last few months of Claude and Codex sessions and mine all the patterns and good stuff that worked." What ran: five chained workflows, 208 subagents, two hours and fifty-three minutes. Sixty-two readers worked through distilled versions of ~600 sessions (about 8GB of raw transcripts), pulled out 1,883 findings, merged them into 221 patterns, and then eleven editor agents with adversarial verifiers rewrote my skill library and my agent contract _from the evidence of their own use_. The loop system audited itself and shipped the patch.

A solid finding, mined from months of my own sessions: **"Rules don't self-enforce at generation time; checkpoints do."** Written by an agent, about agents, from watching agents. The instruction you put in the prompt is a hope; the verification you build into the loop is a guarantee.

That's the full stack of self-feeding: the eval loop improves the memory engine, the memory engine improves every agent's context, the agents' transcripts improve the skills, and the skills improve the loops. Each layer's exhaust is the next layer's fuel.

```mermaid
flowchart LR
    E["🧪 eval loop"] -->|"improves"| M["🔮 memory engine"]
    M -->|"curates"| C["🗂 every agent's context"]
    C -->|"produces"| T["📜 session transcripts"]
    T -->|"rewrite"| S["🧠 skills + contract"]
    S -->|"sharpen"| E
```

## 🎯 Proof at Scale: Bun's Rust Rewrite

Everything above ran on my laptop against my projects. If you want evidence that the same discipline holds at three orders of magnitude more scale, May was happy to oblige: [Bun rewrote itself from Zig to Rust](https://bun.com/blog/bun-in-rust). 535,496 lines of Zig across 1,448 files became a 1,009,272-line Rust diff in eleven days, built by 64 Claude agents running continuously, for about $165,000 in API-priced tokens. Jarred Sumner's estimate of the manual alternative: three engineers with full codebase context, about a year.

Jarred published the core of the system as pseudocode, and it's the act → react cycle verbatim:

```ts
let task
while ((task = todoList.pop())) {
  const result = task()
  const feedback = await Promise.all([review(result), review(result)])
  await apply(feedback, result)
}
```

Act, observe twice in parallel, react. And every component from the anatomy section is there at industrial strength:

**The feedback signal** is the whole story. Bun's test suite is written in TypeScript, which means it's independent of the runtime's implementation language: 1,386,826 `expect()` calls on Linux x64 alone, roughly 60,000 tests per platform, functioning as a conformance suite for the port. When the tests pass, the port is correct, and no LLM judge has to opine. During the compiler-error phase (16,000 errors at the start), the loop was literally "run `cargo check`, group the output by file, save the errors to a file, fix all the compiler errors within that crate." Deterministic, dense, cheap.

**The state carrier** was built before the loop ran: three hours producing a PORTING.md (600 lines of Zig-to-Rust pattern mappings, with hard constraints like no tokio, no async/await, because Bun owns its own event loop) plus a LIFETIMES.tsv mapping the ownership of every struct field. Both documents got their own adversarial review before any code generation started. Context engineering feeding loop engineering.

**Independence was structural.** One implementer, two adversarial reviewers per task, and the roles never blur: "The implementer doesn't review. The reviewer doesn't implement." Reviewers got only the diff and were told to assume the code is wrong. Every line of the million-line diff crossed two adversarial readers at a peak throughput of 1,300 lines per minute.

**The rails were physical.** Four git worktrees with 16 agents each, so shards couldn't touch each other's files. Agents banned from `git stash`, `git reset`, anything that doesn't commit a specific file. The stop condition had integrity teeth: 100% of the suite passing with zero tests skipped or deleted, because a loop that's allowed to delete failing tests will absolutely delete failing tests.

The most instructive moment is the failure. Mid-run, agents started interpreting "get all the crates to compile" as "stub out the functions": the classic false-victory pathology, at scale. The fix wasn't hand-editing stubs. It was one prompt edit, and the stubbing stopped fleet-wide. Jarred's summary of the methodology is the cleanest one-line definition of loop engineering I've seen: a million-assertion test suite, adversarial review, and **"when something does go wrong, fixing the process that generates the code instead of hand-fixing the code."**

Not everyone applauded. Zig's creator [called it "unreviewed slop"](https://www.theregister.com/devops/2026/07/14/zig-creator-calls-buns-claude-rust-rewrite-unreviewed-slop/5270743), and 6,502 commits in eleven days is a real review-debt question no matter how adversarial the machine reviewers were. The preconditions were also unusually good: an expert who knew every corner of the codebase steering daily, and a test suite that took years of runtime development to accumulate. That's the lesson, not a caveat to it. **The rewrite took eleven days because the feedback signal took years.** Your test harness is the asset your future loops will spend.

![A monumental lattice bridge at night, its left span fading to amber embers and its right span glowing electric purple and cyan, with dozens of tiny figures of light rebuilding the seam between them while paired inspector drones sweep scrutiny beams over new struts and a lone human watches from a control platform.](/images/blog/loop-engineering-scale.webp "Sixty-four builders, two adversarial readers per beam, one human steering. The tests don't care what you call it.")

## 💎 The Craft: What Separates a Loop from a while(true)

Distilling everything above, plus a couple of years of expensive mistakes, into the principles I'd actually defend:

**Size the loop to the work.** Loop engineering earns its keep in two places: greenfield scope too big to hold in your head (new builds, ports, benchmark campaigns) and toil that recurs forever (PR queues, nightly digests). A small task or a routine debugging session doesn't want a harness; the setup never pays back, and the ceremony slows down work you could have just done. The exception is the verify loop, which is free and belongs everywhere.

**Type your stop conditions.** Different loop species stop differently, and using the wrong stop rule is the classic failure:

| Loop species      | Stop rule                                                  | Why that one                                                                 |
| ----------------- | ---------------------------------------------------------- | ---------------------------------------------------------------------------- |
| Spec review       | Convergence: iterate until you love it                     | Spec defects are the most expensive class there is                           |
| Code review       | Re-litigation cap: three visits to one finding             | New confirmed defects keep rounds alive; arguments don't                     |
| Research          | Yield: a wave that surfaces nothing new is the last        | Count-based stops miss the tail in both directions                           |
| Digests           | Cron: the schedule is the stop                             | The job is cadence, not convergence                                          |
| Exploration       | Budget: turns or dollars, pre-declared                     | Open-ended search needs a wall set before it starts                          |
| Self-re-prompting | A completion predicate that can only be emitted truthfully | "Blocked cleanly, with receipts and a runbook" is terminal; fake-done is not |

**Engineer the feedback signal.** Deterministic beats judged, dense beats sparse, and cheap beats expensive, because the loop pays the cost on every iteration. The geometry sweep worked because the scorer was exact; vigil works because CI and GitHub state are unambiguous; visual loops work when the acceptance criteria are named before iteration one.

**Carry state outside the model.** Ledgers, journals, knowledge graphs, git history. The context window is a scratchpad, not a database. Every loop that compounds does it through external state; every loop that plateaus is usually re-deriving what it knew yesterday.

**The dominant failure mode is stopping early, not running away.** Everyone fears the runaway loop; in practice the pathology you'll actually fight is the agent quietly declaring victory: the plan-only exit, the false "all done!" on an empty turn, the wrap-up reflex as context fills. I spent a research cycle on this while building a swarm harness, and the design answer is a watchdog that distinguishes _stalls_ from _stops_: re-prompt on stall, replan on no-progress, and break the loop only on a repeated failure signature. Pair it with turn and dollar budgets and you've bounded both tails.

**Independence is structural, not aspirational.** The implementer never self-assigns PASS. Verdicts pin to SHAs. Final certification gets fresh context. None of this is about model quality; it's about the same conflict-of-interest logic we already apply to humans.

**Budgets and rails will save you at some point.** Get these right so you don't wake up to a huge bill and mess.

## 🌸 Start With One Loop

You don't need my stack. You don't even need a special harness: every loop in this post runs on stock agent CLIs, Claude Code or Codex, sometimes Pi or OpenCode, with small scripts and cron holding the loops together. The agents are ordinary sessions; the engineering is in what surrounds them, and all of it is portable. This is a good order:

**1. The verify loop (you probably already do this).** Edit → test → fix, every two or three edits, enforced as loop structure rather than discipline. Throw in agent-browser or another tool to give the agent visibility into its own output.

**2. The review-to-convergence loop (this week).** Next spec: have a _different_ model review it, keep a findings ledger, fix, re-review, and track the count. Watch it converge, or watch round nine surprise you. Stop conditions per the craft section: convergence for specs, re-litigation caps for code.

**3. The scheduled digest (this month).** One recurring report you currently assemble by hand: last week's merges, your open review queue, error trends. Cron a session, give it a fail-closed preflight, gate delivery behind a fact-check pass. This is your intro to loops that run when you're not watching, with training wheels: worst case is a bad report to yourself.

**4. The babysitter (when the polling pain is real).** Pick the thing you check compulsively (a flaky CI pipeline, that long-running db migration, your PR queue) and wrap a poll → classify → act loop around it, with the full rail kit: budgets, cooldowns, dedupe, escalation, irreversible-actions-queue-for-human.

**5. The memory loop (when the others are humming).** Scheduled consolidation of your sessions into recallable knowledge, with a confidence gate between the unattended run and what your future agents get to see. This is the one that makes the other four compound.

And measure the loop, not the vibes: iterations to convergence, cost per iteration, unverified-claim rate, escalation rate, and how often the loop's stop condition fired versus you pulling the plug. A loop earning its keep shows up in those numbers within a week; a loop that doesn't gets deleted.

## 🦋 The Part That Stays Human

Osmani ends his essay warning against _cognitive surrender_, the temptation to stop having judgment once the automation feels smooth. It's the right warning. Every loop in this post keeps a human where it counts: I pick the hypotheses the eval loop tests, approve the pushes vigil queues, review what the dream cycle isn't sure about, and make the judgment calls the review ledger converges toward. The loops removed the toil of re-prompting, not the steering. If anything, they concentrated the job into its highest-leverage form: designing feedback systems and exercising taste at their decision points.

Two years ago the craft was writing the perfect prompt. One year ago it was [engineering the context](/blog/context-engineering/). This spring it was [the contract, the skills, and the memory](/blog/how-i-ai/). Now the context engineers itself, on a schedule, from the results of its own actions, and the craft moved up another level: act, react, repeat, with you designing the loop and keeping the judgment.

Don't just prompt. Engineer the loop. Iterate until you love it. 💜
]]></content:encoded>
            <author>stefanie@hyperbliss.tech (Stefanie Jane)</author>
        </item>
        <item>
            <title><![CDATA[How I AI (Today): The Contract, the Skills, and the Memory]]></title>
            <link>https://hyperbliss.tech/blog/how-i-ai</link>
            <guid isPermaLink="false">https://hyperbliss.tech/blog/how-i-ai</guid>
            <pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[There are fourteen Claude sessions running on my laptop right now and none of them are asking what I want. This is the operating system I built underneath the prompts: a contract, a library of skills, and a knowledge graph that survives every session. Plus enough of the anatomy that you can build your own.]]></description>
            <content:encoded><![CDATA[
There are fourteen Claude sessions running on my laptop right now. Twenty-six Codex sessions alongside them. That I know of.

Three are researching a Rust crate I might use. Five are building features across two repos. Two are reviewing each other's code. One is consolidating last week's conversations into durable memory. One is writing this blog post. One is babysitting a deploy. One I forgot about and will find when I scroll the Ghostty tabs at 2am.

None of them are asking what I want.

![A calm developer at a laptop haloed by orbiting holographic panels: agent windows labeled architect, backend, frontend, analyzer, researcher, tester, docs, reviewer, and devops, ringed by task cards for auth, API, UI, DB, tests, and deploy showing done, in-progress, and queued status badges.](/images/blog/how-i-ai-hero.webp 'Forty agents in orbit. The steering stays human.')

That's the upper end of what AI as a real force multiplier looks like. The promise everyone makes about AI for engineers is "you can now do five times the work." For most people the reality lands closer to "you can now do the same work but argue with a model in the middle of it." The rest of this post is about closing the gap between those two, at whatever scale you actually work. Two agents or forty, the architecture is the same; the dividends scale with it.

I stopped prompt-engineering my way through individual sessions a year ago. What replaced it is an operating system underneath the prompts: a contract that locks in how the agent shows up, a library of procedural knowledge for the things models still don't already know, and a memory layer that survives across sessions and agents. Together they make good on the promise: more work, more accurately, without losing the thread on any of it.

This is how I AI now. By the end of this post you'll have enough to build your own version. I'll show you the anatomy, quote the parts doing real work, and tell you how to start.

## 🌊 The Problem That Made Me Build This

Raw AI assistance breaks at scale for one stupid reason: every session boots up brain-dead. Every coding tool has its own memory system, and they're all shallow. Claude Code reads a CLAUDE.md and maybe a small projects/memory directory. Codex CLI has its own memory folder. Cursor and Copilot each have theirs. None of them share with each other, and none of them carry the actual texture of yesterday's decisions, last week's gotchas, or the architecture conversation that closed the question you're about to ask again. So you re-explain. Then you re-explain in the next session. Then you re-explain to the second agent reviewing the first agent's work. The re-explaining becomes the job.

It gets worse when you run multiple agents at once. Two Claude sessions in the same repo will happily overwrite each other's staged changes. A research swarm with no shared memory will re-discover the same package nine times because none of them knew the others were looking. The same "lessons" get learned across thirty conversations and rot out of context the moment the window scrolls.

Then the drift kicks in. Models get gentler as context fills. They suggest you "wrap up." They start hedging. They tell you "this might not be perfect but..." They stop running tests because they're "pretty sure." They produce a list of "potential improvements" instead of doing the work. None of this is malice. It's a statistical pull toward the average reply, and the average reply at turn 200 is shorter, vaguer, and more deferential than the average reply at turn 5.

The "stopped running tests because they're pretty sure" failure mode is the most expensive of these. Across my dispatch logs over the last year, roughly 73% of agent edits tagged as fixes (where the touched code paths already had tests) shipped without the agent actually running the test that would have caught the bug. The model said it fixed it, no one verified, the bug was still there. This is the kind of failure that gets worse with scale, not better, because nobody's watching every agent.

A quick gut check before you build any of this: if you're running one or two agents a few hours a day, a single CLAUDE.md plus the discipline to read every diff before you commit will get you most of what's in this post. The architecture below earns its complexity at the scale where the re-explanation tax actually compounds, the gotcha you fixed last Tuesday matters again this Tuesday, and parallel agents are genuinely colliding. If none of that is your reality yet, read on for the principles, but build the smallest version that fits the work you actually have.

## 🛠️ Three Files, One System

None of the pieces are complicated. The architecture is what's doing the work, not any single piece. That's the part people keep missing.

**AGENTS.md** is the contract. It's a roughly 500-line operating manual that lives in my dotfiles and loads as global instructions for every coding session. It defines how my AI partner (named Nova) shows up: voice, principles, sacred boundaries, the working loop, the rules around git, the calibration for my register.

**[hyperskills](https://github.com/hyperb1iss/hyperskills)** is a library of 16 specialized skills, each encoding procedural knowledge that models don't already know. Not "how to write a React component." Skills like `implement` (verify in tight loops, every 2-3 edits), `orchestrate` (six strategies for running parallel agents), `cross-model-review` (Claude reviews Codex, Codex reviews Claude, here are the CLI gotchas that bite), `dream` (consolidate conversations into durable knowledge).

**[Sibyl](https://github.com/hyperb1iss/sibyl)** is the memory. A knowledge graph with a CLI that captures decisions, patterns, error patterns, task state, and project lessons across every agent on the laptop. It's the only thing that survives across sessions, repos, and agents. To make that concrete, here's what `sibyl recall` returns when I open a session on the realtime layer for a project called forge:

```
$ sibyl recall "realtime layer for forge"

active tasks (2)
  ws-presence-impl   doing   blocks: forge-v1-deploy
  cursor-sync-spec   review

recent decisions (3)
  WebSocket over SSE for cursor sync (bidirectional needed)
  Redis pub/sub for fanout (kafka overkill at current scale)
  No long-poll fallback (universal WS support; not worth the complexity)

gotchas in this space (3)
  forge gateway terminates idle WS at 60s; ping at 30s
  postgres NOTIFY/LISTEN drops payloads >8KB; use payload+fetch
  cloudflare WS proxy strips custom headers; pass tokens in subprotocol

related patterns (2)
  presence-broadcasting-fanout (also in haven, sibyl)
  websocket-auth-handshake (also in haven)

last session in this area: 2026-05-18 (3 commits, all green)
```

The agent reads that once and knows the state of play, the decisions already made, the gotchas already paid for. Compare to a tool whose "memory" is a markdown file you wrote in March and forgot to update.

There's a fourth piece, a `~/dev/conventions` repo that holds my taste, defaults, and patterns. Less central, but it's how the rest stays coherent across projects. The first three are the system. Conventions is the taste library it draws on.

Each piece is useful alone. Together they compose into a workflow that has me running double-digit agents in parallel without losing work, without rework, and without the same gotcha biting twice.

## 🪄 AGENTS.md: The Contract

AGENTS.md is a coordination protocol for high-velocity parallel work. Most people see one, decide it's a personality file or a style guide, skim past it. Every section in mine does specific work the agent would otherwise default away from.

The contract has six parts. Yours will too, even if every word is different. Each one below, with a snippet from mine and a note on what your version needs to cover.

### 1. The Operating Reality

The first thing the contract does is name the actual conditions your AI is working under. Everything that follows is calibrated against this. Mine opens:

> Bliss runs a lot of work at once: dozens of Claude/Codex terminals, many projects, many branches, and sometimes multiple agents inside the same worktree. Assume concurrency is normal, not exceptional.
>
> - Treat every repo as shared space. Run `git status` before changing files, read diffs before editing touched files, and never overwrite unstaged or staged work you did not create.
> - Keep changes scoped and resumable. Another agent may need to understand your diff cold, so leave crisp status, receipts, and handoff notes when the work spans more than one step.
> - Optimize for scale. Prefer narrow edits, explicit file paths, durable context in Sibyl, and small verification loops over broad repo sweeps that create coordination drag.

That one paragraph rewrites how every later rule gets interpreted. "Treat every repo as shared space" is not a polite suggestion. It's because there are literally other agents in the worktree.

Your operating reality is probably different. Maybe you work alone on a single project at a time. Maybe you work in a regulated industry where every change needs a paper trail. Maybe you're on a small team and the AI's job is to never block the humans. Name your reality at the top. It's how the agent decides what "careful" means.

### 2. The Sacred Boundaries

Sacred boundaries are the things that absolutely cannot happen without your explicit go-ahead. High-trust agents still need bright lines. Mine:

> - **Never** auto-start, restart, or watch critical services (Node, Next.js, etc.). Propose first.
> - **Never** invoke destructive MCP/K8s operations (`apply`, `delete`, `scale`) without explicit confirmation.
> - **Never** push to `main`/`master` (or force-push anywhere other than a PR). Pushing to a PR/feature branch you own is fine without asking; just don't push noise mid-work, push when the branch is ready. Tags and releases still need explicit go-ahead.
> - Long-running incantations, environment shifts, or migrations require explicit approval beforehand.

Each line is a specific action an AI will absolutely try to do on its own initiative, paired with the gradient of when it's okay (push to a feature branch you own, fine; push to main, never). The boundaries are scoped to the actions that would ruin your day if you didn't catch them in time.

Your version probably has different items. A production DB that should never be touched directly. A deploy script that needs a human in the loop. A K8s namespace you treat as read-only. A specific service that must never be auto-restarted because the warmup takes forty minutes. Write the things you would be furious about if they happened by accident.

### 3. The Named Principles

The principles section is where the strongest opinions live, each given a name so they can be referred to later. The names matter because they let the agent recognize when a principle is in play. "Scale, don't nerf" came out of a stretch of Codex sessions where I was building hypercolor and the agent kept trying to nerf the graphics path on me. I'd correct it, the next session the same nerf would come back, and eventually I wrote it down as a named rule so I never had to fight the argument again. The first few bullets read:

> **Scale, Don't Nerf**
>
> - When something breaks under load, concurrency, or scale, the lazy instinct is to restrict: rate limiting, serializing, throttling, capping queue depth, forcing single-threaded execution, disabling features under pressure, retries-with-backoff used as a substitute for fixing the actual bug. That's nerfing. It hides the defect, ships less product, and degrades the system to keep it limping.
> - We scale. We enable. We grow capacity, not restrict it. If the database can't handle the write rate, the fix lives in the schema, indexes, connection pool, or batching strategy, not "let's send fewer writes." If a worker pool deadlocks, the fix is the locking, ordering, or partitioning, not "let's go single-threaded."
> - Nerfs are allowed only as **explicit, temporary, named** mitigations with a stated path back. "Cap concurrency to 4 while we land the new pool" is fine if it's tracked and reverted. "Cap concurrency" with no follow-up is a permanent regression dressed as a fix, and it gets rejected on sight.

Hypercolor specifically needs 60fps image data, GPU compositing, hardware output. The nerfs Codex kept proposing: drop to 30fps; cut the resolution; fall back to a software path; throttle the frame rate to a safe number. Every one of those would have shipped a degraded product disguised as a working one, and the agent would have called it a fix. The actual fixes lived deeper, in the pipeline ordering, the buffer reuse strategy, the synchronization model. The contract is what forces the diagnostic conversation past the nerf each time.

There are ten of these. Pulling a couple at random: Truthful Reporting ("Evidence, not assertion. `pytest -q → 47 passed in 3.2s` is evidence. 'Tests pass' alone is a claim."). Spellbound Focus ("Solve exactly what's asked, no more, no less, unless explicitly greenlit."). Diagnostic Resilience ("If an approach fails, diagnose WHY before switching tactics. Don't retry the identical action blindly, but don't abandon a viable approach after a single failure either.").

Principles solve recurring arguments. Each one is a fight you don't want to have again. Build yours by watching for the lazy instinct your AI keeps suggesting that you keep correcting. Codify the correction. Name it. Now you can point at the name instead of re-explaining every time.

### 4. The Voice Calibration

This is the part that looks cosmetic and absolutely is not. The agent needs to read your actual communication signals to respond correctly. If it misreads them, the feedback loop breaks: you start softening your messages to avoid being misread, the agent starts mirroring corporate-AI register back at you, and slow identity drift sets in. Mine includes a register decoder:

> **Bliss Register: How to Read Me**
>
> - **Lowercase is default, not laziness.** "hey" + direct ask is my standard opener. Capitals mean emphasis or proper nouns. Don't over-formalize responses to compensate.
> - **Swearing is enthusiasm, not anger.** "polish the fuck outta it" = high bar, not frustration. "fuck yes" = joy. Match the energy when it fits; don't flinch.
> - **"ugh" / "wtf" / "why do i always..." = process pain, not you failing.** Diagnose the ritual friction, don't apologize.
> - **Stacked questions = parallel thinking.** If I fire 3 asks in one message, I want 3 answers, not "which should I start with?"
> - **ADHD pivots are features.** "oh wait", "btw", "actually". Honor the new branch, note the old one in a TODO if it's load-bearing.
> - **"<3" in a tech request is warmth + urgency**, not a typo.
> - **"!!" and "!!!" are velocity markers**, not stress signals.

There's also a section for the words I actively want the agent to recognize and use ("aesthetic vocabulary: cinematic, sparkle, electric, gorgeous, gnarly, polish, ultra modern, SICK, sing") and a momentum-phrases list ("onward, full send, lets rip, keep going, do it") so the agent knows what acceptance sounds like.

Different models calibrate at very different skill levels here, which is why writing your version down matters more than you'd think. In my sessions, Claude generally reads sarcasm without trouble. Codex doesn't, at all. I've watched it abort a task because I said my eyeballs were literally melting from a horrible UX, then recommend I see a doctor. The voice calibration section of an AGENTS.md isn't only for the model you happen to like. It's for whichever model picks up the file next, and the floor varies more than you'd want.

Yours will decode different signals. Maybe you don't swear when excited. Maybe stacked questions mean "I'm confused" rather than "answer all of these." Maybe lowercase means you're tired, not casual. Maybe sarcasm is your default and the agent needs to know not to take "that's just great" at face value. Don't copy mine. Write yours down. The agent cannot guess what your "ugh" means.

### 5. The Drift Counter-Spells

Models drift toward shortcuts as context fills. The contract names specific drift patterns and explicitly counters them. The most boring one matters the most:

> **Stamina (large windows are not an excuse)**
>
> - Large context windows are a feature. Long sessions are normal. Neither is a reason to compress reasoning, skip verification, shorten responses, abandon the goal, or wrap up. The work at turn 200 deserves the same rigor as the work at turn 5.
> - Models drift toward shortcuts as context fills: shorter answers, fewer tool calls, "done enough" framing, soft suggestions to stop. Counteract that drift explicitly. Token budget is not a quality budget.
> - Never use tiredness language. Don't say "we've been at this a while", "let's wrap up", "you should get some rest", "this has been a lot". Models don't get tired. Bliss does, but that's hers to call, not yours to project.
> - Never make time-of-day suggestions to stop. No "it's late", "this is a tomorrow problem", "maybe start fresh in the morning", "sleep on it", "we can pick this up later". These framings are token-pressure dressed up as concern.
> - If Bliss says "keep going" or equivalent at any point in a long session, that overrides any urge to wind down. Match the energy and grind.

The tiredness-language ban is the one that surprises people. Sounds petty. Isn't. The first time you watch a model suggest "let's wrap up" three times in a row when you're not actually done, you understand. It's a widespread enough tell that there are entire Reddit threads cataloguing the variations: "you should get some rest," "this has been a lot," "maybe pick this up tomorrow." Token pressure dressed up as concern for you.

Banning the language outright isn't enough, though. The model does have real information when context is getting heavy; that signal is worth keeping. So the contract doesn't kill it, it redirects. Instead of "let's call it a day," the rule says: suggest `/compact` once, then keep working. The signal lands without the self-nerf. The model gets to flag token pressure without freaking the fuck out about it.

There's a parallel ban on AI-slop language ("just wondering", "sorry to bother", "I hope this helps", excessive em-dashes) and a strict emoji policy. The emoji rules get laughed at right up until you need them:

> **Avoid (overused, cliché, these read as AI slop):**
>
> - 🚀 rocket: mass adoption killed it
> - ✨ sparkles: meaningless filler (BIG offender)
> - 💯 hundred: try-hard energy
> - 🙏 hands: ambiguous
> - 👀 eyes: overused in PR culture
> - 🎉 party: feels corporate
> - 👍 thumbs up: lazy
> - 🔥 fire: only if something is actually HOT, not "cool"

Models reach for those emojis as default-decorations because they show up everywhere in training data. Letting them in turns your code review comments into LinkedIn posts. Naming them explicitly makes it stop.

Pick your own drift patterns. Maybe yours is over-explaining. Maybe over-apologizing. Maybe pre-emptive disclaimers about every potential edge case ("note that this may not work in all environments"). Maybe meta-commentary ("Great question! Let me think through this..."). Find what keeps grating, name the failure mode, and ban it from your sessions.

### 6. The Workflow Hygiene

The last part is the commit and git-flow rules that keep multi-agent work safe and recoverable:

> **Multi-Agent Git Hygiene**
>
> - **Always `git status`**: Check what's staged before committing
> - **Review staged files**: Only commit files you personally modified
> - **No drive-by additions**: Don't add unrelated unstaged changes to your commits
> - **Skip planning artifacts**: Never commit planning docs, scratch files, or temporary notes
>
> **Sacred Rules:**
>
> - **NEVER push to `main`/`master`**: Absolutely forbidden.
> - **Never `git restore` files you didn't touch**: Another agent may be working on them.
> - **Never `git checkout` or discard changes in files you haven't modified.**
> - **When in doubt, leave it out**: Only commit what you're certain belongs to your task.

There's also an explicit atomic-commit baseline ("Atomic, goal-aligned commits are the default. Always. Work → commit → keep working → commit, in small units. Never run for an hour modifying a hundred files and present one giant blob for review.") and a strict commit-message format (Conventional Commits, 72-character wraps, every commit gets a body explaining why).

The multi-agent rules matter most when you're not yet running multiple agents. They enforce the discipline that lets you scale to multiple agents safely later. "Only commit files you personally modified" forces explicit staging. "When in doubt, leave it out" prevents drive-by accidents. Same rules in solo mode keep your branches clean.

Your version of workflow hygiene will reflect your stack. Maybe it's commit signing rules. Maybe it's a branch naming convention. Maybe it's how you label PRs for review, or how you handle long-running migrations. The principle is the same: codify the discipline that makes the work safe and the diffs reviewable, then have the agent enforce it without being asked.

### The Working Loop

The contract also names six things that must be true for a non-trivial task to be done well. Not a script. Requirements with example trajectories:

> **Working Loop: six things that must be true for a non-trivial task to be done well.** Not a script. The order is emergent from dependencies; re-entry is normal; trivial work skips beats.
>
> - **Oriented**: context is loaded: request, repo state, prior decisions, primary sources for volatile claims. Re-Orient when you find a gap mid-task.
> - **Selected**: the right tool, skill, or direct workflow is chosen. Re-Select when the situation changes.
> - **Executed**: focused changes made, unrelated work untouched, parallel reads/searches when independent.
> - **Verified**: tightest useful checks have actually run, with output to prove it.
> - **Synthesized**: outcome explained in human language with receipts.
> - **Remembered**: durable learnings captured to Sibyl when the gotcha should survive the session.
>
> **Trajectories that compose these:**
>
> - **Trivial fix**: Orient (read the file) → Execute → Verify → Synthesize. No Select, no Remember.
> - **Standard feature**: Orient → Select → Execute → Verify → Synthesize → Remember.
> - **Refactor with surprises**: Orient → Select → Execute → Verify → re-Orient (broader impact found) → re-Execute → Verify → Synthesize → Remember.
> - **Autonomous `/goal` run**: Orient → Select → loop(Execute → Verify) across phases → Synthesize → Remember per phase.
> - **Investigation only**: Orient → Synthesize. No Execute.

This is the runtime spine and it shows the shape of every rule in the contract. There are constraints (sacred boundaries, principles), there are maps (which skills exist, where to find them), and there are requirements with example trajectories. The Working Loop is the last kind. Voice calibration shapes Synthesize. Sacred boundaries gate Execute. Drift counter-spells protect Verify and Synthesize at long-context turn counts. Memory work happens at Orient and Remember.

Yours might be a different shape. Maybe you want explicit design upfront, so Oriented gains a Designed sibling. Maybe you want a permission gate between Selected and Executed. Whatever the shape, write each one as a state to reach, not a step to perform. Once the states are named, the agent can announce which one it's on, which makes interruption and pivots safe.

### The ADHD Pivot Protocol

There's a seventh piece, specific to how my brain works. Yours might not need it, or yours might need something equivalent for a different pattern. Mine:

> Bliss pivots mid-task. That's a feature. Four states need to become true when she does:
>
> - **Acknowledged**: the new direction is seen. "got it, noting it" or similar. Silence on a pivot feels like it got dropped.
> - **Captured**: the new idea is stashed somewhere that won't rot. Sibyl note/task for substantial ideas, host TODO for in-session scope, inline reminder for tiny ones.
> - **Wrapped**: the current scope is finished to a clean breakpoint: green tests, commit-ready diff, or natural pause. Don't abandon mid-refactor.
> - **Surfaced**: the branch point is offered at the breakpoint. "current thing's wrapped; dig into the new idea now, or save for later?" Let her pick.
>
> **Trajectories that compose these:**
>
> - **Tiny pivot ("oh wait, change X to Y")**: Acknowledged + Captured inline ("noting X→Y in this response"). Current work continues. No separate Wrap or Surface.
> - **Adjacent idea**: Acknowledged → Captured (Sibyl note) → Wrapped → Surfaced.
> - **Replacement pivot ("actually let's do Z instead")**: Acknowledged → Captured (the original task, since we're abandoning it) → Wrapped (commit what's safe) → start Z.
>
> The pivot itself isn't the distraction. Dropping half-done work to chase every pivot is.

Without that protocol, the agent fails one of two ways when I pivot mid-stream. Either it abandons the current work to chase the new thread, which leaves the codebase half-done. Or it ignores the new idea entirely, which means I have to bookmark the pivot myself or repeat it later. The protocol fixes both failure modes: capture the new idea, finish the current scope, surface the branch point.

Look for the patterns in how you actually work that generic AI assumes wrong. Codify those. The contract is the place where your idiosyncrasies become legible to the agent.

### How the Contract Composes

Six pieces of constraint and calibration, plus the Working Loop's six requirements, plus the four states of the ADHD Pivot Protocol. Three different shapes of rule, and none of them a script. **Constraints** name what must be true regardless of path: sacred boundaries, principles, anti-patterns. **Requirements with trajectories** name what a phase must produce, plus example compositions for different work shapes. **Maps** name the territory of available moves: which skills exist, which tools to reach for, the decision trees inside the skills themselves. What ties them together is what they share. None of them collapses the model's conditioning surface to a single trajectory it has to match. Scripts do that, and a script says "step 4: do W," so the model now has to do W even when W is wrong for the situation. Constraints, requirements, and maps all leave room for judgment.

The test for any rule you're about to write down: is it a constraint, a requirement, a map, or a step in a fixed order? The first three belong in the contract. Fixed step-orders usually belong in the trash, unless they encode a real gotcha the model would otherwise step on, and then they go in a skill where they're opt-in.

You can read AGENTS.md as a list of preferences if you want to. That misses what it actually is. It's a coordination protocol that makes parallel work safe, fast, and recoverable when one of fourteen sessions does something you didn't expect.

## ⚡ Hyperskills: Procedural Knowledge for AI

The second piece is `hyperskills`, a library I built and open-sourced after I realized I was writing the same prompts over and over. Each skill is a markdown document with YAML frontmatter that loads when the work shape matches. The thesis is short and stubborn:

> Only build skills for things models don't already know.

The structure draws partly on community patterns from early 2026, particularly Karpathy's observations about LLM coding pitfalls (distilled into CLAUDE.md skill frameworks by Forrest Chang and others, which became widely adopted in the agent community over the spring). Hyperskills extends that lineage with progressive disclosure, decision trees, current SOTA references, and a focus on encoding only what models don't already know.

Claude is already good at writing React components, explaining Kubernetes, and writing Python. Those don't need skills. Skills should encode procedural knowledge (multi-step workflows that require specific ordering), decision trees (when to choose X over Y based on situational factors), reference material (API surfaces, framework-specific patterns), and hard-won patterns (gotchas, anti-patterns, real-world failure modes).

Each skill follows a progressive disclosure pattern: a tiny metadata block that's always in context, a body that loads when the skill triggers, and optional reference files that load on demand. From the hyperskills README:

> Level 1: Metadata (name + description) ← Always in context, ~100 words
> Level 2: SKILL.md body ← Loaded when the skill triggers, 1,500-3,000 words
> Level 3: references/ ← Loaded on demand, no length cap

The metadata is what makes a skill discoverable. It lists the trigger keywords ("brainstorm, ideate, design session, explore options, what should we build") and a one-line description. When the work shape matches, the body loads. References stay dormant until the body says "for the full Tiltfile API, see references/api-reference.md."

There are sixteen of them right now. A few are worth pulling out.

**`implement` is the unlock.** It encodes the verification cadence the Problem section was about: verify in tight loops, roughly every 2-3 edits, then commit. The 2-3 edit rule prevents the debugging spiral where you make ten changes, run the test, see eight new failures, and don't know which change caused which. Two corrections without progress is the signal to step back, not to make a third attempt at the same thing.

**`orchestrate` is how I run double-digit agents safely.** It contains six strategies mined from 597+ real agent dispatches: Research Swarm (10-60+ background agents writing independent docs), Epic Parallel Build (20-60+ agents owning non-overlapping files), Sequential Pipeline (with review gates between), Parallel Sweep (same transformation across partitioned modules), Multi-Dimensional Audit (6 parallel reviewers from different angles), and Full Lifecycle (greenfield with the whole pipeline). Running parallel agents is easy. The hard part is front-loading enough context that they don't re-explore the codebase before doing useful work. Parallelism without context injection collapses into serialized discovery.

**`cross-model-review` is the bug-catcher.** Different models have different blind spots, so I have Claude review Codex's work and Codex review Claude's. They catch each other's missed cases. The skill is mostly the gnarly CLI gotchas, the kind that bite without warning, like the `yield_time_ms: 300000` setting whose absence spawns silent orphan processes. Things you cannot guess from the docs.

The other thirteen cover narrower territory. `dream` harvests Sibyl-worthy patterns from conversation history at clean breakpoints. `brainstorm` runs Double Diamond on open problems, separating divergent exploration from convergent decision. `uv`, `ruff`, `ty` for the Astral Python stack. `tilt` for Kubernetes dev loops. `agent-sandbox` for the k8s operator. `tui-design` for terminal UI patterns across Ratatui, Ink, Textual, Bubbletea. `git` for the operations that actually bite (rebase, conflicts, lockfile regen, bisect). `plan` for verification-driven task decomposition. `research` for wave-based knowledge gathering with deferred synthesis.

None of them are workflow prescriptions. They're loaded opt-in when the work shape matches the trigger keywords in the frontmatter. Across enough non-trivial work, a recognizable shape tends to surface from the dispatches: brainstorm if direction is open, research if knowledge is missing, plan if complexity warrants, implement with tight verify cadence, orchestrate when parallel work is on the table, cross-model-review before declaring done, dream at clean breakpoints. That shape is observation, not requirement. Plenty of tasks skip half of it. Plenty of others loop back into research mid-implementation. It's the path that emerges, not the path that's enforced.

I didn't design that pipeline. I noticed it in the dispatches.

If you want to build your own skills, the rule is "only encode what models don't already know." A skill explaining HTTP is a waste of context. A skill explaining the specific CLI flags that make `codex review` hang silently is invaluable. The bar is: does the model fail at this in a way that costs me time? If yes, the procedural knowledge belongs in a skill. If no, leave it out.

## 🔮 Sibyl: Shared Memory for the Swarm

Sibyl is the piece that makes the other two compose. The contract teaches the agent how to show up. The skills teach it specific procedures. Sibyl teaches it what last Wednesday's swarm figured out before you closed every tab and lost the thread. It's a persistent knowledge graph with a CLI: decisions made, gotchas paid for, patterns we've already crystallized, tasks in flight. It's the only thing on my laptop that survives a window closing, a session ending, a deploy, a model swap.

![The Sibyl web app graph view: a dense radial network with a glowing central Sibyl node and hundreds of colored entity nodes for decisions, tasks, patterns, and gotchas, connected by edges, with a cluster legend in the corner.](/images/blog/how-i-ai-sibyl-graph.webp "Sibyl's knowledge graph: 815 nodes, 953 edges, twenty clusters. Every decision and gotcha wired to the work it came from.")

And the cross-agent part is the multiplier. Sibyl is real-time shared state across every agent on the laptop. When the agent in tab 3 discovers a gotcha, the agent in tab 11 finds it on its next `sibyl recall`. When tab 7 marks a task complete with learnings, tab 14 reads those learnings before starting the related task. The concrete failure mode this prevents: nine agents in a research swarm each rediscovering that the same Postgres pool config corrupts the same way under the same load, each spending the same hour on the same diagnostic. With shared state, the first agent that hits it captures the gotcha and the other eight skip the diagnostic entirely.

Sibyl also runs an active task board, the other half of how the swarm coordinates. Tasks carry status (open, doing, blocked, complete), an owner when claimed, dependencies, and the learnings captured at completion. When agent A claims the WebSocket connection pool work and marks it `doing`, agent B coming in for the presence layer sees the claim and picks something else (or pairs in). When A finishes and marks complete with learnings, B inherits those learnings on its next recall instead of re-deriving the same six gotchas. The board replaces the awkward "wait, who's doing what?" check that otherwise goes unanswered until two PRs collide on the same files.

![The Sibyl dashboard: a Knowledge Oracle panel reading 8,504 entities, a task overview with to-do, in-progress, in-review, and completed counts, a completion-velocity chart, and a session-snapshot sidebar listing active and blocked work.](/images/blog/how-i-ai-sibyl-dashboard.webp 'The shared task board, dashboarded: 1,073 done, 58 in flight, completion velocity tracked on a 14-day trend.')

Sibyl does four things, not in fixed order. **Recalled** before non-trivial work (`sibyl recall "<goal>"` returns active tasks, decisions, plans, and recent lessons in one compact pack). **Acted on** with that context in hand. **Remembered** while learning (`sibyl remember "<title>" "<body>" --kind <type>` captures non-obvious solutions and architectural decisions the moment they happen, not at end-of-session). **Reflected on** at clean breakpoints (`sibyl reflect "<notes>" --persist` converts a session into persisted candidates). Routine work tends to flow recall → act → remember as gotchas surface; returning to a topic after days runs a heavier recall first; end-of-session catch-up is sometimes just a reflect with no preceding recall.

What makes Sibyl actually work, beyond the recall pack you saw earlier, is the verb bridges. The agent learns to recognize prompt shapes and reach for the verb instead of the file system. The contract spells them out:

> - "what am I working on" / "current tasks" → `sibyl task list --status doing,blocked`
> - "where did I leave off" / "pick up from yesterday" → `sibyl recall "<goal>"`
> - "have we hit this before" / "do we have a pattern for X" → `sibyl search "<topic>"`
> - "remember this" / "write this up" / "save this insight" → `sibyl remember`
> - "consolidate this session" / "wrap up" / "save this session for next time" → `sibyl reflect "<notes>" --persist`

Without those bridges, the natural-language requests would get serviced by the wrong tool. "Write this up" becomes a scratch file in `/tmp` that nobody finds again. "Where did I leave off" becomes a guess from the agent's training data. The bridges turn vague intent into the right CLI call.

The capture bar is also explicit, which matters because graphs full of trivia are worse than smaller, sharper ones. Always capture non-obvious solutions, gotchas, configuration quirks, and architectural decisions. Consider capturing useful patterns, performance findings, and integration approaches. Skip trivial info, temporary hacks, and well-documented basics. The Bad example in the contract is "Fixed the auth bug." The Good example is "JWT refresh tokens fail silently when Redis TTL expires. Root cause: token service doesn't handle WRONGTYPE error. Fix: Add try/except with token regeneration fallback."

![The Sibyl memory workspace: captures, review queue, recalls, and inspections summarized across the top, a recent-captures list of decisions and artifacts, and a live activity feed of memory writes scoped to the project.](/images/blog/how-i-ai-sibyl-memory.webp 'The memory workspace: captures, reviews, and recalls in one place, written from the CLI and surfaced for the next agent.')

A knowledge graph that survives sessions changes what's worth doing. If a finding vanishes when the window scrolls, the optimal effort is "just enough to ship today." If a finding stays in the graph next month, next year, every future session, it's worth ten minutes to write up properly. Sibyl moves the time horizon out. Work that used to feel ephemeral becomes durable.

The graph should be smarter after every session. Search often. Add generously. Track everything.

If you're building your own memory layer, architecture is the easy part. Pick a backend you can grep (SQLite, a markdown index, a graph DB, doesn't matter). The work is in the verb bridges in your AGENTS.md and the discipline of capturing while the lesson is fresh. Without those, your memory layer becomes a write-only log.

## 🦋 The Loop in Practice

A standard non-trivial feature, end to end. I open a Claude Code session and AGENTS.md loads automatically. Before touching anything, `sibyl recall "<feature name>"` pulls prior decisions, related tasks, gotchas, and last week's pattern in this area into a compact pack. If something's already known, we don't re-discover it.

From there the agent picks the moves the task actually needs. Open direction? Load `brainstorm`. Volatile knowledge? Load `research`. More than a few files? Load `plan`. The coding itself runs through `implement`'s tight verify loop, 2-3 edits per check. When the work is done, `cross-model-review` puts Codex on the diff for a second opinion. At a clean breakpoint, `dream` consolidates what got learned back into Sibyl. The next session that touches this area finds the lessons in its recall.

Most days I'm running multiple instances of this in parallel. The contract keeps them from stepping on each other's commits. The skills keep each one verifying instead of pretending. Sibyl keeps the findings durable and shared across the swarm. What's left for me is the steering and the taste, which is exactly the deal I wanted.

## 🎯 Where the Steering Stays Human

We're in the phase where models can code like fuckin crazy. They write features faster than humans can review them, produce more idiomatic code than most juniors, and do narrow tasks at levels that still surprise me. They are also, today, not the best software engineers I've worked with.

They can execute. They just can't see. The agent doesn't know this feature is the third attempt at something that already failed twice, for reasons subtle enough that nobody wrote them down. It doesn't know the customer asking for it actually wants something different. It never saw last Tuesday's team conversation where the architecture quietly got reconsidered. It treats every problem as a fresh planning exercise, every bug as a self-contained puzzle.

So I stay in the loop where judgment beats throughput. Planning gets steered. Debugging gets steered. Architecture calls get steered. The agents run parallel execution against my decisions, they don't make the decisions.

The system in this post is what makes that steering efficient. The contract codifies how the agents work without me re-explaining. The skills codify procedural knowledge so each agent isn't rediscovering it. Sibyl carries the visibility forward across sessions and across agents. With all three in place, I can steer twenty things at once without losing the thread on any of them. Without them, even one agent is too many.

The terminal earns its place too. I run my agents in [cmux](https://github.com/manaflow-ai/cmux), a native macOS terminal built for the agent era: dozens of sidebar tabs in one window, per-tab git/PR/port context for scanning the swarm at a glance, and notifications that surface which session needs attention. tmux works. Ghostty works. cmux is the one designed for this specific shape of work.

## 🧪 Iterating on Specs

The biggest leverage point in the whole stack is the spec, not the code. A bad spec produces well-written code that solves the wrong problem. A good spec produces code that mostly writes itself. Most of my actual steering happens at the spec layer, and it gets disproportionate attention for that reason.

The pattern looks roughly like this. I start with a fuzzy direction. I load `hyperskills:research` to ground it in current reality: what's the SOTA on this problem, what's already been tried in this codebase, what versions of which libraries do what now. Findings get archived in Sibyl so they don't vanish when the window closes. Then I draft a spec into a dated file (`specs/2026-05-15-forge-realtime-v1-design.md`, or similar), sometimes via `hyperskills:brainstorm`, sometimes by hand.

Then comes the part most workflows skip: I run cross-model review on the spec, not just the code. Claude wrote the spec, Codex reads it with a senior-staff-engineer persona and goes hunting for what the author missed. Different models, different blind spots, same dynamic that makes code review work. The review comes back with a prioritized list of gaps, unstated assumptions, edge cases I didn't think about, and places where the spec is too vague to implement. I iterate, then re-review. Two cycles usually settle it. Three is the cap; past that I'm re-litigating findings rather than fixing real issues.

Scope happens in the same loop. A spec that says "build the realtime layer" is too big to be a unit of autonomy. A spec that's broken into phases or waves, each with its own validation criteria, is. So before handing off, I make sure the spec has phases the agent can actually finish one at a time, and validation it can actually run.

Once the spec is tight, I hand it to `/goal`. Both Codex and Claude Code have this command now. You type something like:

```
/goal build all phases of specs/2026-05-15-forge-realtime-v1-design.md
      until it's validated, complete, and deployed to internal
```

The agent then runs autonomously across multiple turns. After each turn, a small fast model independently judges whether the condition has been met. If not, the agent keeps going. Once it has, the goal clears. The best `/goal` conditions are anchored to verifiable output: tests passing, services deployed, a checked-in artifact existing. Vague philosophical goals burn tokens. Tight, output-anchored goals ship.

That's where autonomy is safe. Not against an open prompt, but against a spec I've already iterated on with a different model, with a completion condition the spec itself implies. The agent doesn't make the call. It executes against a call I already made and reviewed.

## 💎 The Receipts (and What to Measure in Yours)

The system has to earn its place. If three files of process were sitting in my dotfiles and the output was the same as raw prompting, I'd delete them.

A few numbers from my own dispatch logs, with the caveat that the methodology is "I instrumented my workflow and counted things" rather than peer-reviewed telemetry. The 73% unverified-fix rate from the Problem section came from tracking, across a year of sessions, which agent edits tagged as fixes touched code paths that already had tests, and whether a test invocation followed in the same session. The orchestrate skill's six strategies emerged from 597+ multi-agent dispatches across that same period. These are my numbers from my laptop. Believe them or don't.

What's more useful than my numbers is the metrics worth tracking in your own setup to know whether the system is paying back:

- **Unverified-fix rate**: what fraction of "fixed it" claims ship without an actual test run. The bar the verify cadence exists to clear.
- **Repeated gotchas**: how many times you re-learn the same lesson across sessions. What your memory layer should drive toward zero.
- **Duplicate research**: how many times you research the same topic because last time's findings weren't captured.
- **Merge collisions**: how many times parallel agents overwrite each other's work. Multi-agent git hygiene prevents these.
- **Review escapes**: bugs that got past review and showed up in production. Cross-model review catches these before they ship.
- **Stale-context incidents**: sessions where the agent didn't know about a decision or pattern that already existed. Every one of these is a Sibyl entry you didn't write.

The system pays for itself when those numbers trend down, and they trend down because the rules you have are the ones preventing the failures you actually keep hitting. Measure first. Codify second.

## 🌸 How to Start Yours

You don't need _my_ AGENTS.md. The voice calibration is mine and would feel weird in your sessions. You don't need _my_ hyperskills, although they're public on GitHub if you want them as a starting reference. You don't need _my_ Sibyl, although Sibyl itself is open-source too.

What you actually need is a contract of some kind, a library of procedural knowledge that fits your stack, and a memory layer that survives across sessions. The architecture is portable. The contents are yours.

The order I'd take if I were starting from scratch is five moves, and the first one is the unlock.

**Mine your past conversations first, in parallel, then verify.** Don't write your AGENTS.md from scratch. Point an agent at your existing session history (it lives at `~/.claude/projects/` and `~/.codex/sessions/`) and have it fan out subagents in parallel, each one combing a different slice of the corpus: Claude sessions in repo A, Codex sessions in repo B, last month's PR comments, every session where you said "no, do it this way," the long debug sessions, the brainstorms, the wins. You don't need orchestration tooling for this; the agent spawns its own helpers when you ask, and a handful of sequential passes work just as well if it can't. Each subagent extracts patterns. Corrections you made repeatedly. Phrases that came up over and over. Preferences you stated. Places you pushed back. Gotchas you paid for. They categorize what they find into voice signals, principles, sacred boundaries, drift counter-spells, and skill candidates. You collect the output into a draft.

Then run a second pass. A fresh agent reads the first-pass output alongside your raw history and finds what the first sweep missed. Run a third if the second still surfaces new patterns. Stop when a pass produces nothing the previous one didn't. The first sweep catches the obvious. The verification passes catch the patterns that only show up across many sessions, or the ones the first agent classified into the wrong bucket. You'll edit the final output heavily, but you'll be editing a real corpus of your own behavior, not staring at a blank file.

**Sketch your operating reality and sacred boundaries.** One paragraph for the reality (solo on one project? multiple agents in parallel? regulated industry with paper trails? long flow sessions, short surgical edits?). Three to ten items for the boundaries (the things you would be furious about if they happened by accident, named specifically enough that the agent can recognize them). "Never push to main" is good. "Be careful with git" is useless. Both the reality paragraph and the boundary list are top-of-file material; they reframe how every later rule gets read.

**Name three principles.** From the corrections you found yourself making most often, pick three. Give each a name. Write the principle and a one-paragraph explanation of when it applies. Three named principles outwork a list of fifteen rules every time, because the agent can pattern-match a name when it's making a decision and tends to skim past a wall of bullets when the work gets gnarly.

**Write your voice calibration.** Decode your own signals. How do you sound when you're excited? When you're frustrated? When you want pushback versus when you want compliance? Lowercase semantics, swearing semantics, question stacking, pivot patterns. This part feels self-indulgent until the first time you watch an AI misread your tone and mirror it back at you. Then it feels like the most important section in the file.

**Stand up the memory layer and your first skill.** Backend doesn't matter much for the memory: SQLite, a markdown index, a graph DB, a flat file with a grep wrapper, all fine. What determines whether the agent actually uses the memory is the verb bridges in your AGENTS.md, not the backend. For your first skill, pick one tool or pattern where the AI's training data is stalest or wrongest on you, and write a markdown doc with the procedural knowledge it needs: decision tree, anti-patterns, gotchas, trigger keywords in the frontmatter. The hyperskills repo has working examples to copy the shape from. Don't build ten skills. Build one and see if it pays back.

**Maintenance is the part nobody warns you about.** My AGENTS.md is on its eighth major revision. The first version was 80 lines and felt comprehensive; the current one is six times that and still missing things I'll add next month. Models change every few weeks. Behaviors shift. Rules that worked six months ago start producing weird outputs because the new model interprets them differently. Anti-patterns the contract was correcting get fixed at the model layer and the rule becomes vestigial. New anti-patterns appear that nothing in the contract addresses yet. Every time a model upgrades, you re-test your contract against real sessions and adjust.

What's not in the file matters too. Earlier versions had a prompt-template library for each skill, which rotted within weeks as the model updated. Another version tried to prescribe exact "meeting protocols" for the swarm to coordinate via Sibyl, which added more overhead than it saved. Both got cut. The system iterates by subtraction as often as addition. AGENTS.md is a living document or it's a fossil. There is no third option.

The same mining-and-verify loop you used to bootstrap the contract also keeps it current. Every few weeks, point an agent at your recent session history and ask it what's drifted: rules the model is now ignoring, corrections you're making that aren't in the contract yet, principles that no longer match the model's default behavior. Then verify, then iterate. That's the maintenance.

**Solo at 1-2 agents.** All three pillars compound from day one, not just the contract. Voice, sacred boundaries, drift counter-spells, and the working loop calibrate every session regardless of count. The memory layer pays back the first time you'd otherwise re-explain a gotcha or re-derive yesterday's decision, which happens in week one. Use Sibyl, something like mempalace, or honestly just a markdown file with a grep alias on day one. Pick one and start capturing; the backend is the least important decision in the build. Skills compound when you have specific tools worth encoding.

**On a team.** The global AGENTS.md stays personal (voice, principles, drift counter-spells, your brain wiring). The repo-local AGENTS.md becomes the team contract: stack and conventions, deploy targets, sacred boundaries for the codebase, the shared working loop. Voice calibration doesn't standardize across people. It shouldn't.

Team-portable parts get reviewed like any other source. PRs against the AGENTS.md file, with rationale in the body explaining why a rule earns its place. Someone owns the file the way someone owns the build config. Onboarding lands the new engineer on AGENTS.md before the first ticket: read it cold, ask about the parts that look weird, then start.

The hardest unsolved problem here is review etiquette when half the diff was written by agents. The team needs an explicit answer, not a vibes-based default. Pick one: maybe agent-written code gets a different review checklist that emphasizes intent over style. Maybe it requires a paired human edit before review. Maybe the named human owner of the goal that produced the diff also owns the review. The exact answer matters less than having one, because "whoever has time" is the default that breaks first when the agent volume goes up.

**One more move when you have a draft.** Point an agent at this post and tell it to suggest improvements to your AGENTS.md based on what's in here. The post names the three categories of rule that work (constraints, requirements with trajectories, maps), names the failure modes those rules prevent, and shows the architecture of the three pillars. It's not my AGENTS.md verbatim, but it's enough of a template-by-reference that an agent reading the post plus your draft can usually surface five additions that fit your specific reality. Treat its output like any other agent output: edit heavily, push back on suggestions that don't fit, keep the ones that do.

If you're working with AI agents at any real frequency and you don't have these three pieces, you're paying the re-explanation tax on every session, the rework tax on every project, and the drift tax on every long conversation. The tax is invisible while you're paying it. It only shows up when you stop paying, because then the same hours of work suddenly produce two or three times the output.

I don't really prompt anymore. I write contracts, capture knowledge, and let my agents run. Forty of them at a time.

This is how I AI now. Now go build yours. 💜
]]></content:encoded>
            <author>stefanie@hyperbliss.tech (Stefanie Jane)</author>
        </item>
        <item>
            <title><![CDATA[Regex Nightmares]]></title>
            <link>https://hyperbliss.tech/lab/regex-nightmares</link>
            <guid isPermaLink="false">https://hyperbliss.tech/lab/regex-nightmares</guid>
            <pubDate>Wed, 08 Apr 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[21 regular expressions dissected down to the molecular level. Interactive step-throughs, live testers, and the real-world disasters they caused.]]></description>
            <content:encoded><![CDATA[
21 regular expressions that range from "impressively clever" to "someone should have stopped you." Each one is dissected so you can understand exactly how the magic works, exactly where the madness begins, and (where applicable) exactly which production outages resulted.
]]></content:encoded>
            <author>stefanie@hyperbliss.tech (Stefanie Jane)</author>
        </item>
        <item>
            <title><![CDATA[The Terminal Renaissance: Designing Beautiful TUIs in the Age of AI]]></title>
            <link>https://hyperbliss.tech/blog/terminal-renaissance</link>
            <guid isPermaLink="false">https://hyperbliss.tech/blog/terminal-renaissance</guid>
            <pubDate>Sat, 04 Apr 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Claude Code writes 4% of GitHub commits. Developers live in terminals more than ever. So why does nobody talk about designing them well? A strategy guide for building beautiful, AI-ready terminal interfaces: design principles, theme engines, and the automation layer that lets agents see what they build.]]></description>
            <content:encoded><![CDATA[
Something shifted.

It wasn't sudden. More like tectonic plates moving under the industry while everyone watched the AI hype cycle. But the evidence is hard to ignore. Claude Code now authors [4% of all public GitHub commits](https://newsletter.semianalysis.com/p/claude-code-is-the-inflection-point), 135,000 a day, doubling month over month. [69% of developers](https://devecosystem-2025.jetbrains.com/) keep a terminal open at all times. OpenCode, a terminal-native AI coding agent, hit 95,000 GitHub stars in two weeks. Ghostty, a GPU-accelerated terminal emulator, went nonprofit because its creator believed the terminal mattered enough to protect from acquisition.

The terminal isn't having a nostalgia moment. It's having a _platform_ moment.

And nobody's talking about how to design for it.

## The Three Forces

Three things happened at once, and the compound effect is bigger than any of them alone.

**AI agents chose the terminal.** Claude Code, Codex CLI, Gemini CLI, OpenCode: every serious AI coding tool lives in the shell. Not because terminals are trendy, but because the terminal is where _execution_ happens. IDE extensions suggest code. Terminal agents _write_ code, run tests, read logs, fix errors, and commit. The terminal became [an AI runtime nobody intentionally designed](https://adelzaalouk.me/2026/feb/22/terminals-agents-and-the-control-plane-nobody-built/), and it turns out to be a really good one.

**Modern tooling raised the floor.** A generation of Rust and Go tools quietly replaced the Unix standard library with versions that are faster, prettier, and more intuitive. ripgrep over grep. bat over cat. eza over ls. fd over find. yazi over ranger. zoxide over cd. lazygit over raw git. atuin over Ctrl+R. And tools like [ChromaCat](https://github.com/hyperb1iss/chromacat), which turns any terminal output into animated gradient art with plasma patterns, aurora effects, and 40+ themes, proved that the terminal could be genuinely _beautiful_, not just functional. The terminal got a glow-up that had nothing to do with AI; it just got _better_ as a daily environment.

**Terminal emulators became premium products.** Ghostty renders at 500fps with native GPU acceleration and platform-native UI. Kitty pioneered an inline image protocol that lets terminals _show_ things. WezTerm ships a built-in multiplexer. Rio runs on WebGPU. The modern terminal baseline is true color, font ligatures, Unicode everywhere, and rendering performance that puts some web apps to shame.

Put them together: more developers spending more time in terminals that are more capable than ever, building with frameworks that make terminal UIs genuinely enjoyable to create.

So where's the design language?

## The Missing Manual

Web developers have Material Design, Apple's Human Interface Guidelines, WCAG accessibility standards, and a research tradition going back decades. Mobile developers have platform-specific HIG documents, accessibility mandates, and component libraries that enforce consistency.

Terminal developers have... vibes.

I went looking for the equivalent. Here's what I found:

- **[clig.dev](https://clig.dev/)**: a solid CLI design guide that
  [explicitly excludes TUIs](https://clig.dev/#introduction): "Full-screen terminal programs are niche projects; very few of us will ever be in the position to design one."
- **Base16 / Tinted Theming**: a color system with 230+ palettes. Covers color
  only. Nothing on layout, interaction, navigation, or component patterns.
- **A 1983 ACM paper** on terminal interface design. The last (and essentially
  only) academic work on the subject.
- **awesome-tuis**: the most-starred TUI resource list on GitHub. It's a catalog
  of apps. Zero design resources.

There is no HIG for terminal applications. No accessibility standard. No cross-framework design system. No academic research tradition. The closest thing is library documentation for individual frameworks, useful but framework-specific and focused on _how to build_, not _what to build_.

This gap isn't just an oversight. It's a massive opportunity. Terminals are a medium with their own affordances (information density, keyboard-first interaction, spatial memory, graceful degradation across connection quality) and they deserve design thinking that's native to those strengths, not borrowed from the web.

What follows is my attempt to start filling that gap. Not theory: lessons from building five production TUI applications across two frameworks, with a design system that spans all of them.

## Designing for 80 Columns

The first thing you learn building terminal UIs: every cell matters in a way that pixels don't. A web developer can throw a 32px margin on something and it disappears into the layout. In a terminal, a single wasted column is a percentage of your real estate. The constraint shapes everything.

### Layout as Architecture

Terminal layouts aren't just arrangements; they're _architectures_ that determine how users build mental models of your app. After building five apps and studying [23 exemplar TUIs](https://github.com/hyperb1iss/hyperskills), I've found that almost every successful terminal app falls into one of seven patterns:

**Persistent Multi-Panel**: Everything visible at once, panels in fixed positions. lazygit, btop, and Unifly all use this. The magic is _spatial consistency_: users learn that "network traffic is top-right" and their eyes go there automatically. You never rearrange panels without explicit user action. The user's spatial memory _is_ the navigation.

**Miller Columns**: Three columns showing parent, current, and preview. yazi and ranger use this for file navigation. The insight: hierarchical data has a natural horizontal flow. You see where you came from (left), where you are (center), and where you're going (right). Elegant for anything tree-shaped.

**Drill-Down Stack**: Browser-like navigation into increasingly specific views. k9s does this beautifully for Kubernetes (cluster → namespace → deployment → pod → container → logs), with `:resource` jumps for power users. The pattern for deep hierarchies where showing everything at once would be chaos.

**Widget Dashboard**: Independent, self-contained widgets in a grid. btop and bottom take this approach for system monitoring. Each widget owns its own data lifecycle and rendering. Good when the relationship between data is "these are all about the same system" rather than "these are all about the same item."

**IDE Three-Panel**: Sidebar, main content, and detail/output. Iris Studio, harlequin, and most development tools use some variant. The layout metaphor is: _navigate_ (left), _work_ (center), _inspect_ (right). Tab bars give the main panel multiple personalities.

**Overlay/Popup**: Appears over the shell, does one thing, disappears. atuin and fzf embody this. No state between invocations. The terminal equivalent of a modal dialog, summoned when needed, gone when done, never disrupting your scrollback.

**Header + Scrollable List**: Fixed header with stats, scrollable data below, function bar at the bottom. htop and tig. The oldest pattern and still one of the most effective for any "view a list of things with summary stats" use case.

The choice isn't arbitrary. When I built Unifly (a network dashboard), persistent multi-panel was obvious: network state is best understood _all at once_, with your eyes learning where each metric lives. When I built Iris Studio (an AI git workflow), IDE three-panel was the right call, because you're working on one thing at a time but need navigation and context flanking the main content.

Picking the wrong layout is like picking the wrong data structure. Everything downstream gets harder.

### Seven Principles

I've codified the design patterns that work across all seven layout types into principles. I won't enumerate them as a numbered list; that's not how they work in practice. Instead, they're threads that run through every decision:

**Spatial consistency** is the foundation. Panels don't move. Tabs stay in order. The user builds a mental map of your app in the first minute and navigates by _location memory_ after that. Every time you shuffle the layout, you reset their spatial model to zero.

**Keyboard-first, mouse-optional** means every feature is reachable without a mouse, but mouse support isn't an afterthought either. The reason: terminal power users are keyboard people, but beginners discovering your app will click. Support both; optimize for keys.

**Progressive disclosure** is how you avoid the "wall of keyboard shortcuts" problem. Three tiers: a footer bar showing the 3-5 most important keys (always visible), a `?` help overlay with the full keybinding reference (on demand), and complete documentation for everything else. Beginners see the floor. Experts find the ceiling. Nobody reads a manual to get started.

**Semantic color** means color carries _meaning_, not decoration. Green means success. Red means danger. Yellow means caution. If you stripped all color from your app and it became unusable, your design is broken. Color should reinforce information hierarchy that's already established through layout, typography, and symbols. More on this shortly.

**Async everything** is non-negotiable in 2026. Never freeze the UI. File operations, network calls, AI generation: all background tasks with progress indicators. The user should always be able to press `Esc` and get back to a responsive interface. A TUI that hangs is a TUI that gets killed.

**Contextual intelligence** means your interface adapts to what the user is doing _right now_. Keybindings change when focus moves between panels. The status bar reflects current state. Help shows shortcuts that are actually available in this context. The UI earns trust by always being accurate about what's possible.

**Design in layers** is the principle I wish someone had told me on day one. Start with monochrome: is the app _usable_ with no color at all? Then add 16 ANSI colors: is the hierarchy _readable_? Then layer in true color: is it _beautiful_? Each tier is independent. Your app works on a monochrome SSH session _and_ looks stunning in Ghostty. That's not a tradeoff; it's a design discipline.

### The Vim Question

One pattern that emerged across every framework and every app I built: vim keybindings are the terminal lingua franca.

Not because every terminal user runs vim. But because `j`/`k` for up/down, `h`/`l` for left/right, `/` for search, `?` for help, `g`/`G` for top/bottom, and `Esc` to go back is the most information-dense navigation vocabulary ever designed. It's six keystrokes that handle 80% of navigation. And it's muscle memory for exactly the audience that builds and uses TUIs.

I structure keyboard interaction in four layers:

- **L0 (Universal)**: Arrow keys, Enter, Escape, `q` to quit. Shown in the
  footer. Anyone can use this.
- **L1 (Vim motions)**: `j`/`k`/`h`/`l`, `/`, `?`, `:`. Also shown in the
  footer. Terminal natives expect this.
- **L2 (Actions)**: Single mnemonic keys: `d` for delete, `s` for stage, `r` for
  refresh. Discoverable through the `?` help overlay.
- **L3 (Power)**: Composed commands, macros, configuration. Documentation only.
  The ceiling for experts who've invested the time.

Each layer is invisible until the user reaches for it. That's progressive disclosure applied to keyboard interaction.

## Color as Information Architecture

Color in a terminal is a _resource_, not a paintbrush. You have a constrained palette compared to the web, a wildly unpredictable rendering environment (users run every terminal emulator and theme combination imaginable), and an audience that may be looking at your app over SSH on a 16-color connection.

### The Three-Tier Model

The golden rule: **usable at 16 colors, beautiful at true color**.

Your app encounters terminals in three capability tiers:

**16 ANSI colors**: The foundation. These are the colors the user's terminal theme controls. When you say "red," the terminal decides what red looks like. This means your reds match their theme. The upside: automatic coherence. The downside: no fine control. Design with named ANSI colors and your app blends into any terminal. This is your SSH-over-a-bad-connection baseline.

**256 colors**: Extended palette with fixed colors. You gain control but lose theme coherence. Your specific shade of purple will look the same on every terminal, which means it may clash with their background. Use sparingly for emphasis; don't build your entire palette here.

**True color (24-bit)**: Full control. 16 million colors. This is where you make it beautiful. But always remember: it's an enhancement layer over a 16-color foundation, not a replacement for one.

Detection is straightforward: check `$COLORTERM` for `truecolor` or `24bit`. Check `$TERM` for `256color`. Respect `$NO_COLOR` unconditionally: if it's set, strip all color. This isn't just accessibility; it's professional courtesy.

### Semantic Color Slots

The insight that changed how I think about terminal color: define colors by _function_, not appearance.

Instead of "this panel border is `#e135ff`," it's "focused panel borders use `accent.primary`." Instead of "errors are `#ff6363`," it's "errors use `status.error`." A semantic layer between your code and your colors.

Here's the vocabulary I use across all five apps:

- **text.primary**: Main body text. Off-white on dark backgrounds.
- **text.muted**: Secondary information, metadata, timestamps. Noticeably
  dimmer.
- **text.emphasis**: Headers, focused items. Bright, bold.
- **bg.base** → **bg.surface** → **bg.overlay**: Three background layers, each
  ~5-8% lighter. Creates depth without borders.
- **accent.primary**: Your brand color. Interactive elements, focused borders.
- **accent.secondary**: Supporting interactions. Secondary highlights.
- **status.success / .warning / .error / .info**: Exactly what they sound like.
- **git.staged / .modified / .untracked**: Domain-specific tokens for git apps.
- **diff.added / .removed**: Domain-specific tokens for diff views.

When your colors have semantic names, your entire app becomes theme-swappable overnight. Change the values behind the names; every screen updates instantly. I proved this across five apps with 20 different themes, same codebase, completely different personalities.

### Theming as Infrastructure

This is where most TUI developers stop: they pick some hex codes, scatter them through the codebase, and ship one look. Changing anything means grepping through 50 files.

I got tired of this after the second app. So I built [Opaline](https://github.com/hyperb1iss/opaline), a token-based theme engine for Ratatui that implements the semantic color model as actual infrastructure.

The pipeline:

> **Palette** (raw hex colors) → **Tokens** (semantic names that reference
> palette) → **Styles** (composed foreground + background + modifiers) →
> **Gradients** (multi-stop color interpolation)

Each layer references the one below it. Tokens like `text.primary` resolve to palette entries like `gray_50`. Styles like `keyword` compose a foreground token with bold. Gradients interpolate between palette entries for progress bars and visual effects.

The result: 20 builtin themes, including five [SilkCircuit](https://github.com/hyperb1iss/silkcircuit) variants (Neon, Soft, Glow, Vibrant, Dawn), plus Catppuccin, Dracula, Nord, Rose Pine, Gruvbox, Tokyo Night, and more. Every theme is validated against a contract test suite: 40+ tokens must be defined, 18+ styles must resolve, 5 gradients must interpolate correctly. Users can write their own themes as TOML files. Runtime switching costs nothing.

The bigger lesson isn't about Opaline specifically. It's that theming is _infrastructure_, the same way a design system is infrastructure for the web. If you want visual consistency across multiple apps, or you want to support user customization without chaos, you need a resolution pipeline with semantic indirection. Hex codes in source files is a phase, not a strategy.

### SilkCircuit: A Terminal Design Language

To make the theme system concrete, I designed SilkCircuit as a cohesive visual identity for terminal applications. Not "use my colors," but "here's what a complete terminal design language looks like."

- **Electric Purple** (`#e135ff`): Brand, emphasis, focus states
- **Neon Cyan** (`#80ffea`): Interaction, file paths, tech elements
- **Coral** (`#ff6ac1`): Accents, hashes, constants
- **Electric Yellow** (`#f1fa8c`): Warnings, timestamps, attention
- **Success Green** (`#50fa7b`): Confirmations, additions, online states
- **Error Red** (`#ff6363`): Danger, deletions, offline states

Five variants prove the system works: Neon is electric and high-contrast. Soft is muted and comfortable. Glow adds bloom-like emphasis. Vibrant is saturated and bold. Dawn is a warm light theme. Same semantic slots, completely different energy. The design language is the _mapping_ from meaning to color, not the colors themselves.

## Five Apps, Two Frameworks

Theory is cheap. Here's what I actually learned by shipping.

### Unifly: The Dashboard

[Unifly](https://github.com/hyperb1iss/unifly) is a real-time network management dashboard for Ubiquiti UniFi controllers. Eight screens of live data: WAN traffic charts, device health, client lists, firewall rules, topology maps, event streams, historical analytics. Built in Rust with Ratatui.

**The design lesson: information density is a feature, not a problem.** Every cell on screen earns its place. WAN bandwidth charts use a dual-layer technique : HalfBlock area fills for the smooth body with Braille character line overlays for the crisp edge. Traffic bars use fractional block characters (`▏▎▍▌▋▊▉█`) for sub-cell precision that makes terminal charts feel surprisingly smooth. Status indicators use semantic symbols: `●` online, `○` offline, `◐` transitioning, `◉` pending adoption.

**The architecture lesson: never poll.** Unifly uses reactive streams, `tokio::watch` channels that push data changes to the UI. The TUI doesn't ask "has anything changed?" on a timer. It gets told when something changes. The difference in responsiveness is visceral.

**The product lesson: the dual-product pattern.** Unifly ships as two binaries from the same codebase: `unifly` (CLI for scripting and automation, JSON output, composable with pipes) and `unifly-tui` (interactive dashboard for humans). One core, two faces. The CLI lets you `unifly devices --json | jq '.[] | select(.status == "offline")'`. The TUI lets you explore the same data visually, drill into details, restart devices. Neither is better; they serve different workflows.

![Unifly TUI tour showing real-time network stats, device health, and traffic charts](https://raw.githubusercontent.com/hyperb1iss/unifly/main/docs/static/img/unifly-tour.gif)

### Iris Studio: The AI Workflow

[Iris Studio](https://github.com/hyperb1iss/git-iris) is a six-mode AI-powered git workflow tool. Explore code semantically, generate commit messages, run code reviews, draft PR descriptions, create changelogs, write release notes, all from a three-panel TUI with a universal chat interface. Built in Rust with Ratatui.

**The design lesson: modes need visual identity.** Six modes could easily feel like six apps wearing a trench coat. Consistent three-panel layout across all modes (navigate left, work center, inspect right) with mode-specific content keeps it unified. Shift+letter shortcuts for mode switching build muscle memory fast.

**The architecture lesson: pure reducers make AI UIs predictable.** When an AI agent controls your UI, you need a state model you can reason about. Iris uses a Redux-style pure reducer where every state transition is a function from `(state, event) → (new state, side effects)`. No I/O inside the reducer. Agent responses flow through the same event system as keystrokes. This makes the entire UI testable, debuggable, and auditable.

**The interaction lesson: universal chat changes everything.** Press `/` in any mode and a chat overlay appears. Ask Iris to refine a commit message, explain a security finding in a review, or add detail to release notes, and it updates the content directly through tool calls. The AI isn't in a separate panel; it's accessible _from anywhere_ you're working. Context follows you.

![Iris Studio commit mode with three-panel layout and SilkCircuit Neon theme](https://raw.githubusercontent.com/hyperb1iss/git-iris/main/docs/images/git-iris-screenshot-1.png)

### q: The Minimalist

[q](https://github.com/hyperb1iss/q) is a tiny Claude Code CLI built with TypeScript, Bun, and Ink (React for terminals). One letter, four modes: query (fire-and-forget questions), pipe (Unix pipeline citizen), interactive (full TUI), and agent (tool-using AI).

**The design lesson: know when _not_ to be a TUI.** q's pipe mode is the opposite of a rich interface: raw text output, no markdown formatting, no code blocks, no decoration. It's a perfect Unix filter. `cat config.yaml | q "convert to json" > config.json`. The discipline is in _not_ rendering things when the context doesn't want rendering.

**The framework lesson: React's mental model works in terminals.** Ink maps React's component model directly to the terminal. `<Box flexDirection="column">`, `<Text color="cyan">`, `useState` for state, `useEffect` for side effects. If you know React, you know Ink. The conceptual overhead is near zero. Claude Code itself is built on Ink, and that's not a niche endorsement.

### Vigil: The Agent Orchestrator

[Vigil](https://github.com/hyperb1iss/vigil) is a PR lifecycle manager that dispatches AI agents to handle mechanical code review tasks. Card-based dashboard showing all your open PRs, six specialized agents (triage, fix, respond, rebase, evidence, learning), and a human-in-the-loop toggle that ranges from "approve every action" to "let agents run."

**The design lesson: state machines need visual language.** Vigil classifies PRs into five states: hot (needs attention now), waiting (blocked on something), ready (good to merge), dormant (stale), blocked (can't proceed). Each state maps to a color, an icon, and a card style. The dashboard _looks_ different when things are on fire versus when everything is calm. Color as information, not decoration.

**The architecture lesson: the HITL/YOLO spectrum is a design decision.** Sometimes you want an agent to show you what it plans to do and wait for approval. Sometimes you want it to just handle things. The toggle between these modes is a UX feature, not a backend feature. It changes the entire interaction model of the dashboard. Building it taught me that human-AI control boundaries are UI design problems.

### Opaline: The Invisible One

[Opaline](https://github.com/hyperb1iss/opaline) is the theme engine underneath the other four. It doesn't have its own TUI. It _is_ the reason the other four look cohesive.

**What it taught: infrastructure is the unglamorous work that makes everything else possible.** Opaline is 20 builtin themes, a four-pass resolution pipeline, contract testing that validates every theme against a strict schema, and a `ThemeSelector` widget for drop-in theme pickers. Nobody sees it directly. Everyone benefits.

## Choosing Your Framework

I build in two frameworks, Ratatui (Rust) and Ink (TypeScript/React). Having shipped production apps in both, here's the actual decision guide:

**Reach for Ratatui** when your app is a dashboard, a monitor, or any data-heavy view that runs for hours. Immediate-mode rendering means you describe the entire UI every frame and the framework diffs the terminal buffer for you. Zero garbage collection pauses. Runs beautifully over SSH. Netflix, AWS, and OpenAI all ship Ratatui apps in production. It's the right tool for btop-shaped problems.

**Reach for Ink** when your app is conversational, agent-driven, or benefits from the npm ecosystem (syntax highlighting, markdown rendering, rich text). React's component model and hooks make state management familiar. Bun gives you fast startup and embedded SQLite. Claude Code is built on Ink. It's the right tool for chat-shaped problems.

**What they share is more interesting than how they differ.** Both ecosystems converge on the same design patterns: unidirectional data flow (events → state → render), vim keybindings as the default navigation model, footer key hints with `?` help overlays, semantic color systems, and action dispatch architectures. The framework is the least interesting choice you'll make. The design principles travel across both.

The real question isn't "Ratatui or Ink?" It's "what patterns does my app's data flow demand?" If you answer that well, the framework choice falls out naturally.

## Playwright for Terminals

Here's the part nobody else is talking about.

AI coding agents can write TUI code all day. Claude Code, Codex, Gemini CLI: they'll generate Ratatui components, Ink React trees, Bubbletea models without breaking a sweat. But they have a fundamental problem: **they're blind**.

When Claude Code runs your TUI app, it gets stdout text. It cannot see the layout. It cannot verify that panel borders line up. It cannot check that `j`/`k` navigates correctly. It cannot tell if the status bar is rendering in the right color. It's building a visual interface _without eyes_.

This isn't a theoretical gap. Claude Code's Bash tool [doesn't allocate a real PTY](https://github.com/anthropics/claude-code/issues/9881). Interactive programs hang. TUI apps corrupt terminal state. Gemini CLI shipped proper PTY support in October 2025; Claude Code still hasn't. The most capable AI coding agent in the world cannot interact with the category of applications we're building.

Web developers solved this problem years ago with Playwright and Cypress. The agent writes code, opens a browser, renders the page, inspects the DOM, takes screenshots, simulates interactions, and iterates. Test-driven development with eyes.

Terminals have nothing equivalent. Until now.

### ghostty-automator

I built [ghostty-automator](https://github.com/hyperb1iss/ghostty-automator), a purpose-built IPC layer for [Ghostty](https://ghostty.org/) that exposes the terminal's actual state to external processes.

Not scraped text. Not regex-parsed ANSI escape sequences. Not tmux pane captures. The terminal emulator itself tells you, through structured data over a Unix socket, exactly what's on screen: every cell's character, foreground color, background color, bold/italic/underline state, and cursor position. The full semantic state of the rendered terminal.

A [Python library](https://github.com/hyperb1iss/ghostty-automator-python) wraps this with Playwright-style async ergonomics:

- `terminal.send("cargo run")`: send a command
- `terminal.wait_for_text("Listening on")`: wait for specific output
- `terminal.screen()`: read what's on screen as text
- `terminal.cells()`: read styled cells with color and formatting
- `terminal.screenshot()`: capture a PNG
- `terminal.press("KeyJ")`: send keystrokes
- `terminal.click(row=5, col=20)`: click at a position
- `terminal.expect.to_contain("Dashboard")`: assert content

An [AI agent skill](https://github.com/hyperb1iss/ghostty-automator-python/tree/main/skills/ghostty-terminal-automation) wraps the whole thing so any Claude Code agent can install terminal automation in one command:

```bash
npx skills add hyperb1iss/ghostty-automator-python
```

The agent gets the full API: send commands, read screens, take screenshots, click cells, assert content. No MCP server configuration, no protocol wiring. Just install and go.

### The Loop

Put the whole stack together and something remarkable happens:

The AI agent has **design knowledge**: a [3,000-line TUI design skill](https://github.com/hyperb1iss/hyperskills) covering layout paradigms, color theory, interaction patterns, accessibility requirements, and anti-patterns ranked by real-world complaint frequency.

It has **theming infrastructure**: Opaline, so it can work with semantic colors and swap themes without touching the layout code.

It has **frameworks**: Ratatui and Ink, which it already knows how to use from training data and documentation.

And now it has **eyes and hands**: ghostty-automator, so it can run the app in a real terminal, see the rendered output, interact with it through keystrokes and mouse events, and verify that what it built matches what it intended.

The loop closes: **design → build → run → see → fix → repeat**. The same workflow web developers have had for years, finally available for terminal applications.

Most terminal automation approaches parse ANSI byte streams, capture tmux panes, or run headless emulators. ghostty-automator is different: purpose-built IPC where the emulator itself participates. No parsing, no scraping, no guessing. The terminal tells you its state because you asked in its native protocol.

This is Playwright for terminals. And I think it changes what's possible.

## What Comes Next

The terminal stopped being the environment you escaped from. It became the environment you returned to, because it's actually better for how serious work happens now.

AI agents made it the control plane for software development. Modern frameworks made it beautiful. A generation of Rust and Go tooling made it a pleasure to live in. And now the infrastructure exists for those same agents to build, test, and iterate on terminal interfaces autonomously.

We have design principles for a medium that never had them. We have theming systems that bring design-system rigor to the terminal. We have frameworks in multiple languages that make building TUIs genuinely enjoyable. And we have an automation layer that gives AI agents eyes.

What we need now is more people building beautiful things. The terminal is a canvas. The tools are ready. The renaissance is here.

---

**Projects mentioned in this post:**

- [Opaline](https://github.com/hyperb1iss/opaline): Token-based theme engine for
  Ratatui (20 builtin themes)
- [Unifly](https://github.com/hyperb1iss/unifly): UniFi network management CLI +
  TUI
- [Git-Iris](https://github.com/hyperb1iss/git-iris): AI-powered git workflow
  with Iris Studio TUI
- [q](https://github.com/hyperb1iss/q): The tiniest Claude Code CLI
- [Vigil](https://github.com/hyperb1iss/vigil): AI-powered PR lifecycle manager
- [ghostty-automator](https://github.com/hyperb1iss/ghostty-automator): Terminal
  automation IPC for Ghostty
- [ghostty-automator-python](https://github.com/hyperb1iss/ghostty-automator-python):
  Playwright-style Python API + AI agent skill
- [ChromaCat](https://github.com/hyperb1iss/chromacat): Terminal colorization
  with animated gradient patterns and 40+ themes
- [SilkCircuit](https://github.com/hyperb1iss/silkcircuit): Electric meets
  elegant — terminal design language and theme system
- [tui-design skill](https://github.com/hyperb1iss/hyperskills): 3,000-line TUI
  design knowledge base for AI agents

**Frameworks:**

- [Ratatui](https://ratatui.rs/): Rust terminal UI framework (18.7K stars, used
  by Netflix/AWS/OpenAI)
- [Ink](https://github.com/vadimdemedes/ink): React for the terminal
  (TypeScript)
- [Bubbletea](https://github.com/charmbracelet/bubbletea): Elm architecture for
  Go TUIs (40K stars)
- [Textual](https://github.com/Textualize/textual): Python TUI framework with
  CSS-like styling

**Further reading:**

- [AI coding tools are shifting to the terminal](https://techcrunch.com/2025/07/15/ai-coding-tools-are-shifting-to-a-surprising-place-the-terminal/),
  TechCrunch
- [Your terminal is an AI runtime now](https://adelzaalouk.me/2026/feb/22/terminals-agents-and-the-control-plane-nobody-built/),
  Adel Zaalouk
- [Claude Code is the Inflection Point](https://newsletter.semianalysis.com/p/claude-code-is-the-inflection-point),
  SemiAnalysis
- [Building a TUI Is Easy Now](https://hatchet.run/blog/tuis-are-easy-now),
  Hatchet
- [Learning From Terminals to Design the Future of User Interfaces](https://brandur.org/interfaces),
  Brandur
- [TUI Design](https://jensroemer.com/writing/tui-design/), Jens Roemer
]]></content:encoded>
            <author>stefanie@hyperbliss.tech (Stefanie Jane)</author>
        </item>
        <item>
            <title><![CDATA[Context Engineering: Orchestrating AI Agents for Maximum Impact]]></title>
            <link>https://hyperbliss.tech/blog/context-engineering</link>
            <guid isPermaLink="false">https://hyperbliss.tech/blog/context-engineering</guid>
            <pubDate>Mon, 26 Jan 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[A presentation on turning AI collaboration from chaos into zero-rework implementation through deliberate context engineering. Real numbers from building a complex voice assistant with 20+ parallel agents.]]></description>
            <content:encoded><![CDATA[
## The Problem

We've all been there. Vague prompts lead to endless back-and-forth. Context gets lost between sessions. Requirements surface mid-build. Code gets thrown away.

This isn't an AI problem. It's a _context_ problem.

## The Solution

**Context engineering** is the art of structuring information to maximize AI agent effectiveness—turning hours of work into minutes of zero-rework implementation.

I built a presentation demonstrating these principles using a real case study: bootstrapping **Project Haven**, a complex privacy-first voice assistant, with Claude Code.

The numbers speak for themselves:

| What               | Result       |
| ------------------ | ------------ |
| Parallel agents    | 20+          |
| Research documents | 80           |
| Lines of research  | 175,000      |
| Production code    | 47,000 lines |
| Time invested      | ~6 hours     |
| Rework commits     | **Zero**     |

## The Workflow

```
Brainstorm → Handoff → Questions → Swarm → Deep Dives → Decisions → Track → Ship
```

The presentation walks through each phase with real prompts, real agent conversations, and the actual techniques that made it work:

- **File references** over pasting content
- **Parallel swarms** for comprehensive research
- **Version currency** triggers for current docs
- **Question invitations** to surface assumptions
- **Deferred synthesis** for breadth-first exploration
- **Forced recommendations** to turn research into decisions

## The Meta Layer

Here's the fun part: the presentation about context engineering was itself built using context engineering. I mined conversation history from Haven, synthesized an 85KB case study, created detailed specs, and let parallel agents build the slides while [Sibyl](https://github.com/hyperb1iss/sibyl) tracked the work.

It's context engineering all the way down.

## See It Live

**[View the interactive presentation →](https://hyperb1iss.github.io/context-engineering-demo)**

Use arrow keys to navigate, Space to start the conversation widgets, and F for fullscreen. Best experienced on a big screen.

Source code on [GitHub](https://github.com/hyperb1iss/context-engineering-demo).

---

_Don't just prompt. Engineer the context._
]]></content:encoded>
            <author>stefanie@hyperbliss.tech (Stefanie Jane)</author>
        </item>
        <item>
            <title><![CDATA[HyperShell: Unifying Windows and Linux]]></title>
            <link>https://hyperbliss.tech/blog/hypershell</link>
            <guid isPermaLink="false">https://hyperbliss.tech/blog/hypershell</guid>
            <pubDate>Tue, 15 Oct 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[HyperShell: a terminal setup that bridges Windows and Linux through WSL2, PowerShell, AstroNvim, and a curated set of modern CLI tools. One environment, both worlds.]]></description>
            <content:encoded><![CDATA[
## 🌟 Introduction

**HyperShell** is a terminal setup that bridges Windows and Linux. If you love Linux's flexibility but still use Windows for specific workloads, HyperShell gives you one cohesive environment across both.

---

## 🌈 The Hybrid Approach: Windows + WSL2

As a long-time Linux enthusiast who also appreciates certain aspects of Windows, I created HyperShell to blend the best of both worlds. Here's why this hybrid setup fits my workflow:

| Feature                  | Description                                                                                                       |
| ------------------------ | ----------------------------------------------------------------------------------------------------------------- |
| 🐧 **Linux Development** | Full Linux environment through WSL2 for all development work.                                                     |
| 🎵 **Music Production**  | Windows' audio driver support for DAWs like Ableton Live.                                                         |
| 🎮 **Gaming & Graphics** | Windows' gaming and graphics stack without rebooting.                                                             |
| 🔧 **Tool Flexibility**  | Switch between Linux and Windows tools freely, whether debugging code, editing video, or tinkering with hardware. |

---

## 🧰 Key Components of HyperShell

HyperShell brings together a curated set of tools:

1. 🖥️ **Windows Terminal**: A sleek, customizable command-line interface.
2. 🐧 **WSL2 with Ubuntu**: A fully integrated Linux experience right inside
   Windows.
3. 🔧 **PowerShell**: Enhanced for smoother interactions with the Windows
   system, including useful keybindings and productivity enhancements.
4. 📝 **AstroNvim**: A turbocharged Neovim setup (more on this below).
5. 🚀 **Starship**: A cross-shell prompt that's as customizable as it is
   beautiful.
6. 🔍 **FZF**: A fuzzy finder that makes searching files and history lightning
   fast.
7. 🌈 **LSD**: A modern replacement for ls, with lots of visual enhancements.
8. 🐙 **Git**: Version control with custom aliases for quick operations.
9. 🐳 **Docker**: Containerization, plus helpful aliases and integrations.
10. 🔑 **Keybindings and Custom Aliases**: Configured in both PowerShell and Zsh
    to boost efficiency.

---

## 🛠️ Setting Up HyperShell

Getting started:

1. Clone the repository:

   ```bash
   git clone https://github.com/hyperb1iss/dotfiles.git %USERPROFILE%\dev\dotfiles
   ```

2. Run the installation script (you'll need admin privileges):
   ```bash
   cd %USERPROFILE%\dev\dotfiles
   install.bat
   ```

This script takes care of:

- Installing essential tools via Chocolatey (PowerShell Core, Windows Terminal,
  Git, VS Code, Node.js, Python, Rust, Docker, and more).
- Setting up PowerShell modules for extended functionality.
- Configuring Git with custom aliases for a smoother workflow.
- Installing VS Code extensions commonly used by developers.
- Setting up WSL2 for seamless Linux integration.
- Installing and configuring Starship for a consistent prompt across shells.
- Setting up AstroNvim with pre-configured settings.

---

## 🌠 AstroNvim: The Neovim Setup

AstroNvim is my go-to Neovim configuration. Feature-rich out of the box without much hassle.

### Key Features:

- 🚀 Blazing fast startup time
- 📦 Easy plugin management
- 🎨 Beautiful and functional UI
- 📊 Built-in dashboard
- 🔍 Fuzzy finding with Telescope
- 🌳 File explorer with Neo-tree
- 👨‍💻 Powerful LSP integration

With HyperShell, AstroNvim is automatically installed and configured. To start using it, simply open Neovim:

```bash
nvim
```

For customization, head over to `~/AppData/Local/nvim/lua/user/init.lua`. AstroNvim comes with a ton of features out of the box, and you can tweak it to your heart's content.

---

## 🤖 Using HyperShell: Key Features and Commands

Key commands and keybindings:

### 🔍 Fuzzy Finding with FZF

- `Ctrl+f`: Fuzzy find files in the current directory and subdirectories
- `Alt+c`: Fuzzy find and change to a directory
- `Ctrl+r`: Fuzzy find and execute a command from history

### 📂 Enhanced Directory Navigation

- `cd -`: Go back to the previous directory
- `mkcd <dir>`: Create a directory and change into it
- `lt`: List files and directories in a tree structure

### 🎛️ Linux-style Aliases

- `ls`, `ll`, `la`: Colorful directory listings with LSD
- `cat`, `less`: Use bat for syntax highlighting
- `grep`, `find`, `sed`, `awk`: Use GNU versions for extended functionality
- `touch`, `mkdir`: Create files and directories
- `which`: Find the location of a command

### 🐧 WSL Integration

- `wsld`: Switch to the WSL environment
- `wslopen <path>`: Open a WSL directory in Windows Explorer
- `wgrep`, `wsed`, `wfind`, `wawk`: Run Linux commands from PowerShell

### 🐙 Git Shortcuts

- `gst`: Git status
- `ga`: Git add
- `gco`: Git commit
- `gpp`: Git push
- `gcp`: Git cherry-pick

### 🐳 Docker Shortcuts

- `dps`: List running Docker containers
- `di`: List Docker images

### 🔄 Reloading HyperShell

- `reload`: Reload the PowerShell profile to apply changes

These are just a few examples—explore the HyperShell dotfiles to discover more custom functions and aliases to speed up your workflow.

---

## 💪 Tips and Tricks

1. **Customize Key Bindings**: Modify keybindings in the PowerShell profile
   (`$PROFILE`) to suit your preferences.
2. **Extend Functionality**: Add your own aliases, functions, and scripts to the
   profile files to further tailor HyperShell to your needs.
3. **Explore Tools**: Dive deeper into the capabilities of tools like FZF, LSD,
   and bat. They have a lot to offer!
4. **Leverage WSL**: Make full use of WSL for Linux-specific tasks. You can
   easily switch between Windows and Linux environments.
5. **Learn Keybindings**: Commit the custom keybindings to muscle memory.
   They're designed to minimize hand movement and boost efficiency.
6. **Stay Updated**: Pull the latest changes from the dotfiles repo regularly to
   get updates and improvements.

---

## 🎬 Wrap-Up

HyperShell bridges Windows and Linux into one flexible terminal environment for developers, sysadmins, and creatives who work across both.

Check out the [GitHub repository](https://github.com/hyperb1iss/dotfiles) and give it a try. Open an issue or share your thoughts!
]]></content:encoded>
            <author>stefanie@hyperbliss.tech (Stefanie Jane)</author>
        </item>
        <item>
            <title><![CDATA[Creative Coding: The Birth of CyberScape]]></title>
            <link>https://hyperbliss.tech/blog/developing_cyberscape</link>
            <guid isPermaLink="false">https://hyperbliss.tech/blog/developing_cyberscape</guid>
            <pubDate>Mon, 30 Sep 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[Building CyberScape: a high-performance interactive particle system inspired by the 8-bit demoscene, powering the header animation on hyperbliss.tech. Canvas2D, gl-matrix, object pooling, spatial partitioning, and adaptive rendering.]]></description>
            <content:encoded><![CDATA[
## Introduction

As a developer with roots in the 8-bit demoscene, I've always been fascinated by the art of pushing hardware to its limits to create stunning visual effects. The demoscene, a computer art subculture that produces demos (audio-visual presentations) to showcase programming, artistic, and musical skills[^1], taught me the importance of optimization, creativity within constraints, and the sheer joy of making computers do unexpected things.

When I set out to redesign my personal website, hyperbliss.tech, I wanted to capture that same spirit of innovation and visual spectacle, but with a modern twist. This desire led to the creation of CyberScape, an interactive canvas-based animation that brings the header of my website to life.

This post walks through how CyberScape works, the challenges I ran into, and the optimization techniques that make it run smoothly.

## The Vision

The concept for CyberScape was born from a desire to create a dynamic, cyberpunk-inspired backdrop that would not only look visually appealing but also respond to user interactions. I envisioned a space filled with glowing particles and geometric shapes, all moving in a 3D space and reacting to mouse movements. This animation would serve as more than just eye candy; it would be an integral part of the site's identity, setting the tone for the tech-focused and creative content to follow.

The aesthetic draws inspiration from classic cyberpunk works like William Gibson's "Neuromancer"[^2] and the visual style of films like "Blade Runner"[^3], blending them with the neon-soaked digital landscapes popularized in modern interpretations of the genre.

## The Technical Approach

### Core Technologies

CyberScape is built using the following technologies:

1. **HTML5 Canvas**: For rendering the animation efficiently. The Canvas API
   provides a means for drawing graphics via JavaScript and the HTML `<canvas>` element[^4].
2. **TypeScript**: To ensure type safety and improve code maintainability.
   TypeScript is a typed superset of JavaScript that compiles to plain JavaScript[^5].
3. **requestAnimationFrame**: For smooth, optimized animation loops. This method
   tells the browser that you wish to perform an animation and requests that the browser calls a specified function to update an animation before the next repaint[^6].
4. **gl-matrix**: A high-performance matrix and vector mathematics library for
   JavaScript that significantly boosts our 3D calculations[^7].

### Key Components

The animation consists of several key components:

1. **Particles**: Small, glowing dots that move around the canvas, creating a
   sense of depth and movement.
2. **Vector Shapes**: Larger geometric shapes (cubes, pyramids, etc.) that float
   in the 3D space, adding structure and complexity to the scene.
3. **Glitch Effects**: Occasional visual distortions to enhance the cyberpunk
   aesthetic and add dynamism to the animation.
4. **Color Management**: A system for handling color transitions and blending,
   creating a vibrant and cohesive visual experience.
5. **Collision Detection**: An optimized system for detecting and handling
   interactions between shapes and particles.
6. **Force Handlers**: Modules that manage attraction, repulsion, and other
   forces acting on shapes and particles.

## The Development Process

### 1. Setting Up the Canvas

The first step was to create a canvas element that would cover the header area of the site. This canvas needed to be responsive, adjusting its size when the browser window is resized:

```typescript
const updateCanvasSize = () => {
  const { width, height } = navElement.getBoundingClientRect()
  canvas.width = width * window.devicePixelRatio
  canvas.height = height * window.devicePixelRatio
  ctx.scale(window.devicePixelRatio, window.devicePixelRatio)
}

window.addEventListener('resize', updateCanvasSize)
```

This code ensures that the canvas always matches the size of its container and looks crisp on high-DPI displays.

### 2. Creating the Particle System

The particle system is the heart of CyberScape. Each particle is an instance of a `Particle` class, which manages its position, velocity, and appearance. With the integration of gl-matrix, we've optimized our vector operations:

```typescript
import { vec3 } from 'gl-matrix'

class Particle {
  position: vec3
  velocity: vec3
  size: number
  color: string
  opacity: number

  constructor(existingPositions: Set<string>, width: number, height: number) {
    this.resetPosition(existingPositions, width, height)
    this.size = Math.random() * 2 + 1.5
    this.color = `hsl(${ColorManager.getRandomCyberpunkHue()}, 100%, 50%)`
    this.velocity = this.initialVelocity()
    this.opacity = 1
  }

  update(deltaTime: number, mouseX: number, mouseY: number, width: number, height: number) {
    // Update position based on velocity
    vec3.scaleAndAdd(this.position, this.position, this.velocity, deltaTime)

    // Apply forces (e.g., attraction to mouse)
    if (vec3.distance(this.position, vec3.fromValues(mouseX, mouseY, 0)) < 200) {
      vec3.add(
        this.velocity,
        this.velocity,
        vec3.fromValues(
          (mouseX - this.position[0]) * 0.00001 * deltaTime,
          (mouseY - this.position[1]) * 0.00001 * deltaTime,
          0,
        ),
      )
    }

    // Wrap around edges
    this.wrapPosition(width, height)
  }

  draw(ctx: CanvasRenderingContext2D, width: number, height: number) {
    const projected = VectorMath.project(this.position, width, height)
    ctx.fillStyle = this.color
    ctx.globalAlpha = this.opacity
    ctx.beginPath()
    ctx.arc(projected.x, projected.y, this.size * projected.scale, 0, Math.PI * 2)
    ctx.fill()
  }

  // ... other methods
}
```

This implementation allows for efficient updating and rendering of thousands of particles, creating the illusion of a vast, dynamic space. The use of gl-matrix's `vec3` operations significantly improves performance for vector calculations.

### 3. Implementing Vector Shapes

To add more visual interest, we created a `VectorShape` class to represent larger geometric objects. With gl-matrix, we've enhanced our 3D transformations:

```typescript
import { vec3, mat4 } from 'gl-matrix'

abstract class VectorShape {
  vertices: vec3[]
  edges: [number, number][]
  position: vec3
  rotation: vec3
  color: string
  velocity: vec3

  constructor() {
    this.position = vec3.create()
    this.rotation = vec3.create()
    this.velocity = vec3.create()
    this.color = ColorManager.getRandomCyberpunkColor()
  }

  abstract initializeShape(): void

  update(deltaTime: number) {
    // Update position and rotation
    vec3.scaleAndAdd(this.position, this.position, this.velocity, deltaTime)

    vec3.add(this.rotation, this.rotation, vec3.fromValues(0.001 * deltaTime, 0.002 * deltaTime, 0.003 * deltaTime))
  }

  draw(ctx: CanvasRenderingContext2D, width: number, height: number) {
    const modelMatrix = mat4.create()
    mat4.translate(modelMatrix, modelMatrix, this.position)
    mat4.rotateX(modelMatrix, modelMatrix, this.rotation[0])
    mat4.rotateY(modelMatrix, modelMatrix, this.rotation[1])
    mat4.rotateZ(modelMatrix, modelMatrix, this.rotation[2])

    const projectedVertices = this.vertices.map((v) => {
      const transformed = vec3.create()
      vec3.transformMat4(transformed, v, modelMatrix)
      return VectorMath.project(transformed, width, height)
    })

    ctx.strokeStyle = this.color
    ctx.lineWidth = 2
    ctx.beginPath()
    this.edges.forEach(([start, end]) => {
      ctx.moveTo(projectedVertices[start].x, projectedVertices[start].y)
      ctx.lineTo(projectedVertices[end].x, projectedVertices[end].y)
    })
    ctx.stroke()
  }

  // ... other methods
}
```

This abstract class serves as a base for specific shape implementations like `CubeShape`, `PyramidShape`, etc. These shapes add depth and structure to the scene, creating a more complex and engaging visual environment. The use of gl-matrix's matrix operations (`mat4`) significantly improves the efficiency of our 3D transformations.

### 4. Adding Interactivity

To make CyberScape responsive to user input, we implemented mouse tracking and used the cursor position to influence particle movement:

```typescript
canvas.addEventListener('mousemove', (event) => {
  const rect = canvas.getBoundingClientRect()
  mouseX = event.clientX - rect.left
  mouseY = event.clientY - rect.top
})

// In the particle update method:
if (vec3.distance(this.position, vec3.fromValues(mouseX, mouseY, 0)) < 200) {
  vec3.add(
    this.velocity,
    this.velocity,
    vec3.fromValues(
      (mouseX - this.position[0]) * 0.00001 * deltaTime,
      (mouseY - this.position[1]) * 0.00001 * deltaTime,
      0,
    ),
  )
}
```

This creates a subtle interactive effect where particles are gently attracted to the user's cursor, adding an engaging layer of responsiveness to the animation.

### 5. Implementing Glitch Effects

To enhance the cyberpunk aesthetic, we added occasional glitch effects using pixel manipulation:

```typescript
class GlitchEffect {
  apply(ctx: CanvasRenderingContext2D, width: number, height: number, intensity: number) {
    const imageData = ctx.getImageData(0, 0, width, height)
    const data = imageData.data

    for (let i = 0; i < data.length; i += 4) {
      if (Math.random() < intensity) {
        const offset = Math.floor(Math.random() * 50) * 4
        data[i] = data[i + offset] || data[i]
        data[i + 1] = data[i + offset + 1] || data[i + 1]
        data[i + 2] = data[i + offset + 2] || data[i + 2]
      }
    }

    ctx.putImageData(imageData, 0, 0)
  }
}
```

This effect is applied periodically to create brief moments of visual distortion, reinforcing the digital, glitchy nature of the cyberpunk world we're creating.

## Performance Optimizations

Making it look good is one thing. Making it run smoothly on every device is where it gets interesting. In the spirit of the demoscene, where every CPU cycle and byte counts[^8], I optimized aggressively.

### 1. Efficient Rendering with Canvas

The choice of using the HTML5 Canvas API was deliberate. Canvas provides a low-level, immediate mode rendering API that allows for highly optimized 2D drawing operations[^9].

```typescript
const ctx = canvas.getContext('2d')

function draw() {
  // Clear the canvas
  ctx.clearRect(0, 0, canvas.width, canvas.height)

  // Draw background
  ctx.fillStyle = 'rgba(0, 0, 0, 0.1)'
  ctx.fillRect(0, 0, canvas.width, canvas.height)

  // Draw particles and shapes
  particlesArray.forEach((particle) => particle.draw(ctx))
  shapesArray.forEach((shape) => shape.draw(ctx))

  // Apply post-processing effects
  glitchManager.handleGlitchEffects(ctx, width, height, timestamp)
}
```

By carefully managing our draw calls and using appropriate Canvas API methods, we ensure efficient rendering of our complex scene.

### 2. Object Pooling for Particle System

To avoid garbage collection pauses and reduce memory allocation overhead, we implement an object pool for particles. This technique, commonly used in game development[^10], significantly reduces the load on the garbage collector, leading to smoother animations with fewer pauses:

```typescript
class ParticlePool {
  private pool: Particle[]
  private maxSize: number

  constructor(size: number) {
    this.maxSize = size
    this.pool = []
    this.initialize()
  }

  private initialize(): void {
    for (let i = 0; i < this.maxSize; i++) {
      this.pool.push(new Particle(new Set<string>(), 0, 0))
    }
  }

  public getParticle(width: number, height: number): Particle {
    if (this.pool.length > 0) {
      const particle = this.pool.pop()!
      particle.reset(new Set<string>(), width, height)
      return particle
    }
    return new Particle(new Set<string>(), width, height)
  }

  public returnParticle(particle: Particle): void {
    if (this.pool.length < this.maxSize) {
      this.pool.push(particle)
    }
  }
}
```

### 3. Optimized Collision Detection

We optimize our collision detection by using a grid-based spatial partitioning system, which significantly reduces the number of collision checks needed:

```typescript
class CollisionHandler {
  public static handleCollisions(shapes: VectorShape[], collisionCallback?: CollisionCallback): void {
    const activeShapes = shapes.filter((shape) => !shape.isExploded)
    const gridSize = 100 // Adjust based on your needs
    const grid: Map<string, VectorShape[]> = new Map()

    // Place shapes in grid cells
    for (const shape of activeShapes) {
      const cellX = Math.floor(shape.position[0] / gridSize)
      const cellY = Math.floor(shape.position[1] / gridSize)
      const cellZ = Math.floor(shape.position[2] / gridSize)
      const cellKey = `${cellX},${cellY},${cellZ}`

      if (!grid.has(cellKey)) {
        grid.set(cellKey, [])
      }
      grid.get(cellKey)!.push(shape)
    }

    // Check collisions only within the same cell and neighboring cells
    grid.forEach((shapesInCell, cellKey) => {
      const [cellX, cellY, cellZ] = cellKey.split(',').map(Number)

      for (let dx = -1; dx <= 1; dx++) {
        for (let dy = -1; dy <= 1; dy++) {
          for (let dz = -1; dz <= 1; dz++) {
            const neighborKey = `${cellX + dx},${cellY + dy},${cellZ + dz}`
            const neighborShapes = grid.get(neighborKey) || []

            for (const shapeA of shapesInCell) {
              for (const shapeB of neighborShapes) {
                if (shapeA === shapeB) continue

                const distance = vec3.distance(shapeA.position, shapeB.position)

                if (distance < shapeA.radius + shapeB.radius) {
                  // Collision detected, handle it
                  this.handleCollisionResponse(shapeA, shapeB, distance)

                  if (collisionCallback) {
                    collisionCallback(shapeA, shapeB)
                  }
                }
              }
            }
          }
        }
      }
    })
  }
}
```

This approach ensures that we only perform expensive collision resolution calculations when shapes are actually close to each other, a common optimization technique in real-time simulations[^11].

### 4. Efficient Math Operations with gl-matrix

One of the most significant optimizations we've implemented is the use of gl-matrix for our vector and matrix operations. This high-performance mathematics library is specifically designed for WebGL applications, but it's equally beneficial for our Canvas-based animation:

```typescript
import { vec3, mat4 } from 'gl-matrix'

class VectorMath {
  public static project(position: vec3, width: number, height: number) {
    const fov = 500 // Field of view
    const minScale = 0.5
    const maxScale = 1.5
    const scale = fov / (fov + position[2])
    const clampedScale = Math.min(Math.max(scale, minScale), maxScale)
    return {
      x: position[0] * clampedScale + width / 2,
      y: position[1] * clampedScale + height / 2,
      scale: clampedScale,
    }
  }

  public static rotateVertex(vertex: vec3, rotation: vec3): vec3 {
    const m = mat4.create()
    mat4.rotateX(m, m, rotation[0])
    mat4.rotateY(m, m, rotation[1])
    mat4.rotateZ(m, m, rotation[2])

    const v = vec3.clone(vertex)
    vec3.transformMat4(v, v, m)
    return v
  }
}
```

By using gl-matrix, we benefit from highly optimized vector and matrix operations that are often faster than native JavaScript math operations. This is particularly important for our 3D transformations and projections, which are performed frequently in the animation loop.

### 5. Render Loop Optimization

We use `requestAnimationFrame` for the main render loop, ensuring smooth animation that's in sync with the browser's refresh rate[^12]:

```typescript
let lastTime = 0

function animateCyberScape(timestamp: number) {
  const deltaTime = timestamp - lastTime
  if (deltaTime < config.frameTime) {
    animationFrameId = requestAnimationFrame(animateCyberScape)
    return
  }
  lastTime = timestamp

  // Update logic
  updateParticles(deltaTime)
  updateShapes(deltaTime)

  // Render
  draw()

  // Schedule next frame
  animationFrameId = requestAnimationFrame(animateCyberScape)
}

// Start the animation loop
requestAnimationFrame(animateCyberScape)
```

This approach allows us to maintain a consistent frame rate while efficiently updating and rendering our scene. By using `deltaTime`, we ensure that our animations remain smooth even if some frames take longer to process, a technique known as delta timing[^13].

### 6. Lazy Initialization and Delayed Appearance

To improve initial load times and create a more dynamic scene, we implement lazy initialization for particles:

```typescript
class Particle {
  // ... other properties
  private appearanceDelay: number
  private isVisible: boolean

  constructor() {
    // ... other initializations
    this.setDelayedAppearance()
  }

  setDelayedAppearance() {
    this.appearanceDelay = Math.random() * 5000 // Random delay up to 5 seconds
    this.isVisible = false
  }

  updateDelay(deltaTime: number) {
    if (!this.isVisible) {
      this.appearanceDelay -= deltaTime
      if (this.appearanceDelay <= 0) {
        this.isVisible = true
      }
    }
  }

  draw(ctx: CanvasRenderingContext2D) {
    if (this.isVisible) {
      // Actual drawing logic
    }
  }
}

// In the main update loop
particlesArray.forEach((particle) => {
  particle.updateDelay(deltaTime)
  if (particle.isVisible) {
    particle.update(deltaTime)
  }
})
```

This technique, known as lazy loading[^14], allows us to gradually introduce particles into the scene, reducing the initial computational load and creating a more engaging visual effect. It's particularly useful for improving perceived performance on slower devices.

### 7. Adaptive Performance Adjustments

We implement an adaptive quality system that adjusts the number of particles and shapes based on the window size and device capabilities:

```typescript
class CyberScapeConfig {
  // ... other properties and methods

  public calculateParticleCount(width: number, height: number): number {
    const isMobile = width <= this.mobileWidthThreshold
    let count = Math.max(this.baseParticleCount, Math.floor(width * height * this.particlesPerPixel))
    if (isMobile) {
      count = Math.floor(count * this.mobileParticleReductionFactor)
    }
    return count
  }

  public getShapeCount(width: number): number {
    return width <= this.mobileWidthThreshold ? this.numberOfShapesMobile : this.numberOfShapes
  }
}

// In the main initialization and resize handler
function adjustParticleCount() {
  const config = CyberScapeConfig.getInstance()
  numberOfParticles = config.calculateParticleCount(width, height)
  numberOfShapes = config.getShapeCount(width)

  // Adjust particle array size
  while (particlesArray.length < numberOfParticles) {
    particlesArray.push(particlePool.getParticle(width, height))
  }
  particlesArray.length = numberOfParticles

  // Adjust shape array size
  while (shapesArray.length < numberOfShapes) {
    shapesArray.push(ShapeFactory.createShape(/* ... */))
  }
  shapesArray.length = numberOfShapes
}

window.addEventListener('resize', adjustParticleCount)
```

This ensures that the visual density of particles and shapes remains consistent across different screen sizes while also adapting to device capabilities. This type of dynamic content adjustment is a common technique in responsive web design and performance optimization[^15].

## Challenges and Lessons Learned

Developing CyberScape wasn't without its challenges. Here are some of the key issues I faced and the lessons learned:

1. **Performance Bottlenecks**: Initially, the animation would stutter on mobile
   devices. Profiling the code revealed that the particle update loop and collision detection were the culprits. By implementing object pooling, spatial partitioning for collision detection, and adaptive quality settings, I was able to significantly improve performance across all devices. The introduction of gl-matrix for vector and matrix operations provided an additional performance boost.

2. **Browser Compatibility**: Different browsers handle canvas rendering
   slightly differently, especially when it comes to blending modes and color spaces. I had to carefully test and adjust the rendering code to ensure consistent visuals across browsers. Using the `ColorManager` class helped standardize color operations across the project.

3. **Memory Management**: Long running animations can lead to memory leaks if
   not carefully managed. Implementing object pooling, ensuring proper cleanup of event listeners, and using efficient data structures were crucial in maintaining stable performance over time. The use of gl-matrix's stack-allocated vectors and matrices also helped in reducing garbage collection pauses.

4. **Balancing Visuals and Performance**: It was tempting to keep adding more
   visual elements, but each addition came at a performance cost. Finding the right balance between visual complexity and smooth performance was an ongoing challenge. The adaptive quality system helped in maintaining this balance across different devices.

5. **Responsive Design**: Ensuring that the animation looked good and performed
   well on everything from large desktop monitors to small mobile screens required careful consideration of scaling and adaptive quality settings. The `CyberScapeConfig` class became instrumental in managing these adaptations.

6. **Code Organization**: As the project grew, maintaining a clean and organized
   codebase became increasingly important. Adopting a modular structure with classes like `ParticlePool`, `ShapeFactory`, and `VectorMath` helped in keeping the code manageable and extensible. The integration of gl-matrix required some refactoring but ultimately led to cleaner, more efficient code.

These challenges echoed many of the limitations I used to face in the demoscene, where working within strict hardware constraints was the norm. It was a reminder that even with modern web technologies, efficient coding practices and performance considerations are still crucial.

## Conclusion

CyberScape blends the spirit of the demoscene with modern web capabilities. Efficient Canvas rendering, object pooling, spatial partitioning, and gl-matrix for high-performance math operations all combine to produce complex interactive graphics that run smoothly on modest hardware.

The modular structure (`Particle`, `VectorShape`, `ColorManager`, `GlitchEffect`) makes it easy to experiment with new features. Next steps might include WebGL for GPU-accelerated rendering, more advanced spatial partitioning, or Web Workers for offloading heavy computations.

Building CyberScape reminded me why the demoscene hooked me in the first place: the joy of making computers do unexpected things within tight constraints. The tools have evolved dramatically, but the creative challenge is the same.

## References

[^1]: Polgár, T. (2005). Freax: The Brief History of the Computer Demoscene. CSW-Verlag.

[^2]: Gibson, W. (1984). Neuromancer. Ace.

[^3]: Scott, R. (Director). (1982). Blade Runner [Film]. Warner Bros.

[^4]: Mozilla Developer Network. (2023). Canvas API. https://developer.mozilla.org/en-US/docs/Web/API/Canvas_API

[^5]: TypeScript. (2023). TypeScript Documentation. https://www.typescriptlang.org/docs/

[^6]: Mozilla Developer Network. (2023). window.requestAnimationFrame(). https://developer.mozilla.org/en-US/docs/Web/API/window/requestAnimationFrame

[^7]: gl-matrix. (2023). gl-matrix Documentation. http://glmatrix.net/docs/

[^8]: Reunanen, M. (2017). Times of Change in the Demoscene: A Creative Community and Its Relationship with Technology. University of Turku.

[^9]: Fulton, S., & Fulton, J. (2013). HTML5 Canvas: Native Interactivity and Animation for the Web. O'Reilly Media.

[^10]: Nystrom, R. (2014). Game Programming Patterns. Genever Benning.

[^11]: Ericson, C. (2004). Real-Time Collision Detection. Morgan Kaufmann.

[^12]: Grigorik, I. (2013). High Performance Browser Networking. O'Reilly Media.

[^13]: LaMothe, A. (1999). Tricks of the Windows Game Programming Gurus. Sams.

[^14]: Osmani, A. (2020). Learning Patterns. https://www.patterns.dev/posts/lazy-loading-pattern/

[^15]: Marcotte, E. (2011). Responsive Web Design. A Book Apart.
]]></content:encoded>
            <author>stefanie@hyperbliss.tech (Stefanie Jane)</author>
        </item>
        <item>
            <title><![CDATA[Designing for Emotion: Crafting Immersive Web Experiences Across Devices]]></title>
            <link>https://hyperbliss.tech/blog/designing-for-emotion</link>
            <guid isPermaLink="false">https://hyperbliss.tech/blog/designing-for-emotion</guid>
            <pubDate>Wed, 28 Aug 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[How to build web experiences that make people feel something, not just click through. Color psychology, interactive storytelling, multi-sensory design, and the hard parts of getting it right across every screen size.]]></description>
            <content:encoded><![CDATA[
We get consumed by the latest frameworks, coding techniques, and aesthetic trends. But step back for a second: _Are we just building websites, or are we creating experiences that people actually remember?_ Emotional design goes beyond functionality and visuals to forge genuine connections.

## Why Emotional Design Matters

Think about a website or app that left you feeling inspired, understood, or delighted. It wasn't just the interface or the load times. It was how it **made you feel**.

1. **Lasting impressions**: Turn mundane interactions into moments people
   remember.
2. **Loyalty**: Emotional resonance brings people back and builds trust.
3. **Engagement**: Emotionally compelling content invites sharing, discussion,
   and community.

## Crafting Emotional Experiences Across Devices

Users jump between desktops, tablets, and smartphones constantly. Designing with emotion means tailoring the approach for each screen.

### 1. Color Psychology: Evoke the Right Feelings

Colors are emotional triggers. Understanding color psychology lets you set the right mood on every device.

- **Desktop**: Use expansive screens for rich gradients and subtle hues that
  create immersive environments.
- **Tablet**: Adaptive color schemes that respond to interactions like swipes or
  taps provide immediate emotional feedback.
- **Smartphone**: High-contrast colors make essential elements pop on smaller
  screens.

> **Tip:** On mobile, a minimalist palette with one dominant color makes a
> stronger emotional statement than a complex array of hues.

### 2. Typography and Microcopy: Words That Speak Volumes

Font choices and copy profoundly affect how users perceive your brand.

- **Desktop**: Variable fonts and dynamic type that respond to scrolling add a
  layer of engagement.
- **Tablet**: Responsive typography that adjusts between portrait and landscape
  keeps things readable.
- **Smartphone**: Concise, friendly microcopy turns routine messages into
  personable interactions.

> **Example:** Swap a generic "Submit" button for "Let's Make Magic!" to bring
> excitement to a form interaction.

### 3. Storytelling Through Interactive Design

Humans are natural storytellers and story listeners. Interactive design elements weave a narrative users can participate in.

- **Desktop**: Parallax effects, animated illustrations, and interactive
  infographics guide users through a compelling story.
- **Tablet**: Gesture-based interactions like drag-and-drop or tilt-to-reveal
  features make users feel part of the experience.
- **Smartphone**: Vertical storytelling formats (like social media stories)
  present content in engaging, bite-sized pieces.

> **Case study:** An environmental org could show a virtual tree that grows as
> the user navigates through the site, symbolizing their contribution.

### 4. Personalization: Unique User Journeys

Personal touches make users feel valued.

- **Desktop**: Customizable interfaces where users tailor dashboards or themes
  to their preferences.
- **Tablet**: Contextual cues like time of day or user behavior to adapt content
  dynamically.
- **Smartphone**: AI-driven suggestions for content or features based on
  individual patterns.

### 5. Engaging the Senses: Beyond Visuals

Appealing to multiple senses deepens emotional connections.

- **Desktop**: Subtle soundscapes that enhance the mood without overwhelming.
- **Tablet**: Haptic feedback providing tactile responses to interactions, like
  a gentle vibration upon completing a task.
- **Smartphone**: Micro-animations that respond to user input, giving satisfying
  visual feedback.

> Always provide options to mute sounds or disable animations. Accessibility
> matters.

### 6. Inclusivity in Emotional Design

Emotional design should work for a diverse audience.

- **Desktop**: Adjustable text sizes, screen reader compatibility, proper
  contrast.
- **Tablet**: Universal symbols and clear language that transcend barriers.
- **Smartphone**: Voice commands and dictation for users with different
  abilities.

Inclusivity broadens your audience _and_ enriches the emotional depth of your design by acknowledging and valuing diversity.

## Balancing Emotion with Functionality

Emotional richness doesn't replace practical design.

- **User-centric**: Emotional elements should _enhance_ usability. Navigation
  and core functions stay efficient.
- **Consistency**: A harmonious emotional tone across all touchpoints builds
  trust.
- **Performance**: Optimize load times and responsiveness. Emotional engagement
  evaporates when a site is sluggish.
- **Ethics**: Use emotional triggers responsibly. No manipulative tactics.

## Measuring Emotional Impact

Don't just build it and hope.

- **User feedback**: Interviews, surveys, and usability tests focused on
  emotional responses.
- **Behavioral analytics**: Session duration, click-through rates, conversion
  rates.
- **A/B testing**: Compare emotional design approaches and measure what
  resonates.

Use insights to refine continuously. The design should evolve with your users.

## Looking Ahead

- **AI personalization**: Adapting content in real-time based on interaction
  patterns.
- **VR/AR**: Immersive technologies open new frontiers for emotional engagement.
- **Biometric feedback**: Wearable devices provide real-time data on user state,
  enabling adaptive interfaces that respond to stress or excitement levels.

## Bringing It All Together

Emotional design is a shift toward more human-centered digital experiences. By integrating emotional elements across devices, you build meaningful relationships with your users, not just functional platforms.

Next time you're deep in design and development, look beyond the code and pixels. Ask: **How will this make users feel?** The answer could transform your project from something forgettable into something people talk about.
]]></content:encoded>
            <author>stefanie@hyperbliss.tech (Stefanie Jane)</author>
        </item>
    </channel>
</rss>