Skip to content

About

Self-hostable memory for AI agents. A neural surprise gate decides what to keep for zero LLM tokens, and nothing leaves your machine. Memory and a Gemma 4 model run on one AMD MI300X, ROCm end to end. Free, open source for devs, sovereign for enterprise.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

flashbulb-amd

Open, self-hostable memory for AI agents — a memory that learns, at inference, what's worth keeping and decides it for zero LLM tokens, with nothing ever leaving your machine. The memory and the model run on one AMD Instinct MI300X.

AMD Developer Hackathon: ACT II · Track 3 (Unicorn). Live demo: https://somtri--flashbulb-amd-web.modal.run

The problem — two taxes, not one

Every shipped agent-memory system (Mem0, Zep, Letta) pays two taxes on the write path:

  1. A token meter. Each turn, they call an LLM to extract facts and decide add/update/noop. That per-turn tax grows with usage — roughly $5k–$30k/month at 10M turns — and dominates production cost.
  2. A data-egress you can't turn off. That write-path call ships your agent's memory to a cloud API. For a bank, hospital, or government agency, that isn't a cost line — it's a compliance dead-end. The buyers who most need durable agent memory legally can't let it leave their boundary.

flashbulb-amd removes both.

The idea

A ~2.8M-param Titans-style neural memory is trained at inference time. Each observation produces a learned surprise signal — how poorly current memory predicts it:

  • high surprise → novel → store it (a "flashbulb" memory) and learn it
  • low surprise → redundant → skip it

Surprise is a ~5 ms tensor op — not an LLM call. Zero tokens — and it's a local network, so nothing is sent anywhere. Forgetting falls out of the memory's weight decay: a contradicting update decays the stale fact.

It is deliberately hybrid — the neural memory is only a salience/retention filter; verbatim recall comes from an explicit SQLite + embeddings store of the records that passed the gate. (Published work shows raw neural memories memorize but fail free-form retrieval; this avoids that.)

Measured results (on the MI300X, reproducible — see benchmarks/)

Metric Result
Write-path tokens, 30 turns (same model) 0 vs 2,055 (llm_extract baseline)
Decision latency ~5 ms vs ~400 ms LLM call — ~37× faster, and free
Gate accuracy (at deployed τ=0.40) 96.7% (14/14 novel, 2/2 updates, 9/10 dups)
Edge cases 6/6 pass (empty, unicode, injection, rapid-dup)

Why AMD — the enabler, not the decoration

The surprise gate, the memory store, and a Gemma 4 (26B-A4B) LLM are co-located on a single MI300X (192 GB HBM3), ROCm end-to-end. This is what makes zero-egress self-hosting real: memory and model on one accelerator — zero external inference, no second API vendor, no network hop. Non-NVIDIA is a feature here — hardware sovereignty and supply diversification are exactly what regulated and government buyers want. The gate is "free" because the memory network is a FLOP rounding-error next to the LLM, and the MI300X has the headroom to hold both. Real custom test-time-training compute on ROCm, not an API decorated as "AMD."

Who it's for — and how it's a business (open core)

The same property — it runs entirely on your own hardware — serves two buyers, which is why this is one company and not two:

  • Developers (free, MIT). Self-host anywhere: a laptop, your own GPU. Drives adoption, credibility, and community. IC-facing features are free.
  • Sovereign Enterprise (paid). Regulated and air-gapped orgs get compliance (audit logs, RBAC, data-residency), a managed AMD appliance, and support/SLAs. Exec-facing features — compliance, control, support — are paid.

Land bottom-up with developers; the enterprise deal follows the champion. The incumbents are architecturally cloud-API-dependent, so their "self-hosted" story is incomplete — memory still phones home on writes. flashbulb-amd is the only stack that's self-hostable end to end.

Architecture

Browser ─HTTP─► FastAPI (container on the MI300X)
                   │  per turn
                   ├─ MemorySidecar (memory/sidecar.py)   ◄─ runs on ROCm
                   │    embed → NeuralMemory surprise → gate (τ) → SQLite store + forgetting
                   └─ Gemma 4 via an OpenAI-compatible shim (llm_client.py) ◄─ also on the MI300X

The memory core (memory/sidecar.py) imports no provider; all LLM calls go through one OpenAI-SDK shim, so the model is swappable without touching the memory.

Run it

Local dev (CPU, Python 3.12):

uv venv --python 3.12 .venv
uv pip install --python .venv/Scripts/python.exe torch --index-url https://download.pytorch.org/whl/cpu
uv pip install --python .venv/Scripts/python.exe -r requirements.txt
.venv/Scripts/python.exe -m pytest        # 11/11
.venv/Scripts/python.exe -m uvicorn app:app --port 8137   # dashboard at :8137 (mock LLM)

Set .env (see .env.example) to point llm_client at any OpenAI-compatible endpoint to go live. On the MI300X, the memory and Gemma 4 (served via vLLM/transformers) run on-GPU (DEVICE=cuda).

Honest positioning

Token-free write gating isn't first-to-market (SAGE, arXiv:2605.30711, June 2026), and on short conversations our learned gate only ties a simpler cosine-density baseline. Our contribution is a learned/trainable gate (Titans gradient-surprise) that is self-hostable end to end alongside the generating LLM on AMD silicon — the only stack that runs memory and model token-free and zero-egress on one MI300X. The learned memory's edge is online adaptation over longer horizons (future work, with adaptive-τ and contradiction-aware forgetting).

Tech

AMD Instinct MI300X · ROCm · AMD Developer Cloud · titans-pytorch · PyTorch · Gemma 4 · sentence-transformers · FastAPI · SQLite · Docker

License

MIT — see LICENSE.

The memory core released here is significantly updated for a later submission to the Qwen Cloud hackathon; this repo is its first public, AMD-native release. Lineage disclosed, not hidden.

About

Self-hostable memory for AI agents. A neural surprise gate decides what to keep for zero LLM tokens, and nothing leaves your machine. Memory and a Gemma 4 model run on one AMD MI300X, ROCm end to end. Free, open source for devs, sovereign for enterprise.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages