Skip to content

Add a deployment sizing calculator under tools/sizing (port of #1553 to main) - #1680

Merged
malatewang merged 1 commit into
MemMachine:mainfrom
wanghy73:port-main/sizing-calculator-1553
Sep 18, 2026
Merged

malatewang merged 1 commit into
MemMachine:mainfrom
wanghy73:port-main/sizing-calculator-1553

Conversation

@wanghy73

@wanghy73 wanghy73 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Port of #1553 from speedkick to main.

Cherry-pick of ac353c0ab7de88c105b055635cb7ad773afff680, the squash commit #1553 merged as on
speedkick. It applied to main with no conflicts, and the diff is
byte-identical to the one reviewed on #1553 (checked with git patch-id), so the
review there still stands.

Re-checked on this branch, based on main:

  • The calculator's own suite: 476 tests, all pass (python -m unittest from
    tools/sizing, standard library only).
  • ruff check and ruff format --check clean repo-wide. Note [tool.ruff] exclude
    in pyproject.toml lists tools, so CI does not lint this directory.
  • Server unit suite unchanged from main (1869 passed): this PR is additive and
    touches nothing under packages/.

The body below is #1553's, unchanged.


What this is

A sizing calculator for MemMachine deployments, under tools/sizing/. Give it a design peak in
operations per second and it returns the hardware a deployment needs: API servers, vector-store
machines and their memory size, a PostgreSQL server, embedding and agent-model GPU cards, hot vector
memory and disk for a chosen retention period, the max_connections PostgreSQL will need, the
network peaks, and how many callers that capacity holds.

Why it exists

Capacity questions about MemMachine were being answered in prose, and prose cannot be re-run. When
someone asks "what if we keep episodes for a year instead of ninety days", or "what if a tenth of
searches use agent mode instead of one in a hundred", the honest answer is a different hardware
order, not a paragraph. This puts the model in one file so the answer can be recomputed instead of
re-argued, and so every assumption behind it is visible.

Three properties matter more than the arithmetic:

  • Every number is labelled measured, derived, estimate or assumption. A reader can see at a
    glance which figures rest on a benchmark and which rest on a guess that has never been tested.
  • Every input is a knob, on the command line and on the web form, so nobody has to edit code to
    ask a different question.
  • It states what it does not know. The embedding-card rate has never been benchmarked. The
    traffic mix is a planning assumption nobody has measured. The tool says so where it uses them,
    rather than presenting a confident total.

Usage

It needs no dependencies — Python standard library only.

cd tools/sizing
uv run --no-project python memmachine_sizing.py --help
Subcommand What it does
tier pilot|target|scale the full report for one of three built-in design peaks (20, 100, 1,000 ops/s)
calc --ops N the same report for any design peak
users --humans N --automated M converts a caller population into the capacity it demands
validate prints every published figure for all three tiers and writes them as JSON
serve a local web form and a JSON endpoint offering the same inputs

Re-price an order against a different traffic mix:

uv run --no-project python memmachine_sizing.py tier target --agent 2 --plain 53

Size for a mixed population, where each kind of caller sends a different mix:

uv run --no-project python memmachine_sizing.py users --humans 2500 --automated 75 \
  --human-mix 90/9/1 --automated-mix 10/20/70

Two things that sound alike and are not

The README leads with this, because confusing them sizes a deployment wrongly.

Agent-mode search is a property of a request — the agent_mode flag on a MemMachine search.
A search sent with it fans out into a multi-hop retrieval of roughly 22 embedding calls, 22 vector
searches, 44 database reads and one or two language-model calls. It is the third share of the
traffic mix, set with --agent. It is a cost multiplier: how expensive one request is.

An automated client is a property of a caller — a program sending requests in a loop rather than
a person typing. One in a five-second tool loop is assumed to send about 0.4 operations per second,
where a human chat session sends 0.011 to 0.028. It is a population count on users, set with
--automated. It is a rate multiplier: how fast requests arrive.

They are independent, so each population carries its own traffic mix and the calculator blends them.

Where the numbers come from

The one measured anchor is 180 searches per second per 16-vCPU AMD EPYC server at 8 worker
processes, from a benchmark on 30 August 2026 with a real OpenAI embedder, a 12,000-episode corpus
and the rrf-hybrid reranker enabled — so the reranker's cost is already inside every derived server
count. Everything else is derived from it, estimated, or an assumption, and the README lists all of
them with their defaults.

What it deliberately does not do

It does not price high-availability additions — a second copy of every vector, a PostgreSQL standby
or a second gateway are a separate decision, and it says so rather than guessing.

Testing

382 tests, standard-library unittest, run with uv run --no-project python -m unittest from
tools/sizing. They cover the arithmetic at exact machine boundaries, every command-line
subcommand, the web form driven end to end, and rejection of bad input on every path. ruff check tools/sizing is clean.

@edwinyyyu edwinyyyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not human-reviewable.

The calculator turns a design peak in operations per second, plus a traffic
mix of adds, plain searches and agent-mode searches, into a hardware order:
API servers, vector-store machines and their RAM size, a PostgreSQL server,
embedding and agent-model GPU cards, hot vector RAM and disk for a chosen
retention period, the max_connections PostgreSQL will need, north-south and
east-west network peaks, and how many callers that capacity holds. It also
works backwards, turning a caller population into the capacity it demands.

Every number it prints is labelled measured, derived, estimate or assumption,
and every input is a command-line flag and a box on the web form, so a reader
changes any of them without editing code.

Two things that sound alike are kept apart throughout. Agent-mode search is a
property of one REQUEST - the agent_mode flag, which fans one search out into
about 22 - and it is a cost multiplier. An automated client is a property of a
CALLER - a program sending requests in a loop - and it is a rate multiplier.
Each caller population carries its own traffic mix and the two are blended,
because a person can send agent-mode searches and a program need not.

The unit everywhere is a session sending requests at the same moment, whoever
is behind it: one developer driving a ten-user load test is ten sessions. A
user count is not a session count, so converting one to the other takes two
figures - the share of users active at the busiest moment, and the sessions
one active user holds. Both ship as example defaults, labelled as such, with
a warning naming which were defaulted and what to replace them with. The
share is 10 per 100 people: a convention, not a measurement, plausibly 5 to
20, and every machine count moves with it.

The two per-caller rates carry published sources. A human chat session at
0.011 to 0.028 operations per second is the median to about the 90th
percentile of a session's busiest five minutes, measured across 55,295
sessions in the BurstGPT dataset (arXiv:2401.17644). An automated client at
0.4 is TraceLab's measured 5.0-second median step across roughly 4,300
production agent sessions (arXiv:2606.30560). Both reports also carry the
figures those sources revealed and no count uses: a heavy session at 0.06,
and an automated client at 0.07 sustained, six times below its burst because
an agent is idle most of the wall-clock time.

Sizing for the worst sustained five minutes follows ITU-T E.500, which
requires read-out periods greater than five minutes so that resources are not
dimensioned for infrequent small-interval peaks.

It deliberately does not price high-availability additions: a second copy of
every vector, a PostgreSQL standby or a second gateway are a separate
decision.

476 tests, standard-library unittest, covering the arithmetic at exact
machine boundaries, every subcommand, the web form driven end to end, and
rejection of bad input on every path. ruff check is clean.

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng
(cherry picked from commit 4aeb020)
Signed-off-by: Haiyan Wang <[email protected]>
@wanghy73
wanghy73 force-pushed the port-main/sizing-calculator-1553 branch from 5e474f1 to 3a7d47c Compare September 18, 2026 00:36
@malatewang
malatewang merged commit 4ad28de into MemMachine:main Sep 18, 2026
44 checks passed
This was referenced Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants