Open, self-hostable memory for AI agents — a memory that learns, at inference, what's worth keeping and decides it for zero LLM tokens, with nothing ever leaving your machine. The memory and the model run on one AMD Instinct MI300X.
AMD Developer Hackathon: ACT II · Track 3 (Unicorn). Live demo: https://somtri--flashbulb-amd-web.modal.run
Every shipped agent-memory system (Mem0, Zep, Letta) pays two taxes on the write path:
- A token meter. Each turn, they call an LLM to extract facts and decide add/update/noop. That per-turn tax grows with usage — roughly $5k–$30k/month at 10M turns — and dominates production cost.
- A data-egress you can't turn off. That write-path call ships your agent's memory to a cloud API. For a bank, hospital, or government agency, that isn't a cost line — it's a compliance dead-end. The buyers who most need durable agent memory legally can't let it leave their boundary.
flashbulb-amd removes both.
A ~2.8M-param Titans-style neural memory is trained at inference time. Each observation produces a learned surprise signal — how poorly current memory predicts it:
- high surprise → novel → store it (a "flashbulb" memory) and learn it
- low surprise → redundant → skip it
Surprise is a ~5 ms tensor op — not an LLM call. Zero tokens — and it's a local network, so nothing is sent anywhere. Forgetting falls out of the memory's weight decay: a contradicting update decays the stale fact.
It is deliberately hybrid — the neural memory is only a salience/retention filter; verbatim recall comes from an explicit SQLite + embeddings store of the records that passed the gate. (Published work shows raw neural memories memorize but fail free-form retrieval; this avoids that.)
Measured results (on the MI300X, reproducible — see benchmarks/)
| Metric | Result |
|---|---|
| Write-path tokens, 30 turns (same model) | 0 vs 2,055 (llm_extract baseline) |
| Decision latency | ~5 ms vs ~400 ms LLM call — ~37× faster, and free |
| Gate accuracy (at deployed τ=0.40) | 96.7% (14/14 novel, 2/2 updates, 9/10 dups) |
| Edge cases | 6/6 pass (empty, unicode, injection, rapid-dup) |
The surprise gate, the memory store, and a Gemma 4 (26B-A4B) LLM are co-located on a single MI300X (192 GB HBM3), ROCm end-to-end. This is what makes zero-egress self-hosting real: memory and model on one accelerator — zero external inference, no second API vendor, no network hop. Non-NVIDIA is a feature here — hardware sovereignty and supply diversification are exactly what regulated and government buyers want. The gate is "free" because the memory network is a FLOP rounding-error next to the LLM, and the MI300X has the headroom to hold both. Real custom test-time-training compute on ROCm, not an API decorated as "AMD."
The same property — it runs entirely on your own hardware — serves two buyers, which is why this is one company and not two:
- Developers (free, MIT). Self-host anywhere: a laptop, your own GPU. Drives adoption, credibility, and community. IC-facing features are free.
- Sovereign Enterprise (paid). Regulated and air-gapped orgs get compliance (audit logs, RBAC, data-residency), a managed AMD appliance, and support/SLAs. Exec-facing features — compliance, control, support — are paid.
Land bottom-up with developers; the enterprise deal follows the champion. The incumbents are architecturally cloud-API-dependent, so their "self-hosted" story is incomplete — memory still phones home on writes. flashbulb-amd is the only stack that's self-hostable end to end.
Browser ─HTTP─► FastAPI (container on the MI300X)
│ per turn
├─ MemorySidecar (memory/sidecar.py) ◄─ runs on ROCm
│ embed → NeuralMemory surprise → gate (τ) → SQLite store + forgetting
└─ Gemma 4 via an OpenAI-compatible shim (llm_client.py) ◄─ also on the MI300X
The memory core (memory/sidecar.py) imports no provider; all LLM calls go through one OpenAI-SDK
shim, so the model is swappable without touching the memory.
Local dev (CPU, Python 3.12):
uv venv --python 3.12 .venv
uv pip install --python .venv/Scripts/python.exe torch --index-url https://download.pytorch.org/whl/cpu
uv pip install --python .venv/Scripts/python.exe -r requirements.txt
.venv/Scripts/python.exe -m pytest # 11/11
.venv/Scripts/python.exe -m uvicorn app:app --port 8137 # dashboard at :8137 (mock LLM)Set .env (see .env.example) to point llm_client at any OpenAI-compatible endpoint to go live.
On the MI300X, the memory and Gemma 4 (served via vLLM/transformers) run on-GPU (DEVICE=cuda).
Token-free write gating isn't first-to-market (SAGE, arXiv:2605.30711, June 2026), and on short conversations our learned gate only ties a simpler cosine-density baseline. Our contribution is a learned/trainable gate (Titans gradient-surprise) that is self-hostable end to end alongside the generating LLM on AMD silicon — the only stack that runs memory and model token-free and zero-egress on one MI300X. The learned memory's edge is online adaptation over longer horizons (future work, with adaptive-τ and contradiction-aware forgetting).
AMD Instinct MI300X · ROCm · AMD Developer Cloud · titans-pytorch · PyTorch · Gemma 4 · sentence-transformers · FastAPI · SQLite · Docker
MIT — see LICENSE.
The memory core released here is significantly updated for a later submission to the Qwen Cloud hackathon; this repo is its first public, AMD-native release. Lineage disclosed, not hidden.