Skip to content

Add a deployment sizing calculator under tools/sizing - #1553

Merged
leomem merged 1 commit into
speedkickfrom
add-sizing-calculator
Aug 30, 2026
Merged

leomem merged 1 commit into
speedkickfrom
add-sizing-calculator

Conversation

@leomem

@leomem leomem commented Aug 30, 2026

Copy link
Copy Markdown
Collaborator

What this is

A sizing calculator for MemMachine deployments, under tools/sizing/. Give it a design peak in
operations per second and it returns the hardware a deployment needs: API servers, vector-store
machines and their memory size, a PostgreSQL server, embedding and agent-model GPU cards, hot vector
memory and disk for a chosen retention period, the max_connections PostgreSQL will need, the
network peaks, and how many callers that capacity holds.

Why it exists

Capacity questions about MemMachine were being answered in prose, and prose cannot be re-run. When
someone asks "what if we keep episodes for a year instead of ninety days", or "what if a tenth of
searches use agent mode instead of one in a hundred", the honest answer is a different hardware
order, not a paragraph. This puts the model in one file so the answer can be recomputed instead of
re-argued, and so every assumption behind it is visible.

Three properties matter more than the arithmetic:

  • Every number is labelled measured, derived, estimate or assumption. A reader can see at a
    glance which figures rest on a benchmark and which rest on a guess that has never been tested.
  • Every input is a knob, on the command line and on the web form, so nobody has to edit code to
    ask a different question.
  • It states what it does not know. The embedding-card rate has never been benchmarked. The
    traffic mix is a planning assumption nobody has measured. The tool says so where it uses them,
    rather than presenting a confident total.

Usage

It needs no dependencies — Python standard library only.

cd tools/sizing
uv run --no-project python memmachine_sizing.py --help
Subcommand What it does
tier pilot|target|scale the full report for one of three built-in design peaks (20, 100, 1,000 ops/s)
calc --ops N the same report for any design peak
users --humans N --automated M converts a caller population into the capacity it demands
validate prints every published figure for all three tiers and writes them as JSON
serve a local web form and a JSON endpoint offering the same inputs

Re-price an order against a different traffic mix:

uv run --no-project python memmachine_sizing.py tier target --agent 2 --plain 53

Size for a mixed population, where each kind of caller sends a different mix:

uv run --no-project python memmachine_sizing.py users --humans 2500 --automated 75 \
  --human-mix 90/9/1 --automated-mix 10/20/70

Two things that sound alike and are not

The README leads with this, because confusing them sizes a deployment wrongly.

Agent-mode search is a property of a request — the agent_mode flag on a MemMachine search.
A search sent with it fans out into a multi-hop retrieval of roughly 22 embedding calls, 22 vector
searches, 44 database reads and one or two language-model calls. It is the third share of the
traffic mix, set with --agent. It is a cost multiplier: how expensive one request is.

An automated client is a property of a caller — a program sending requests in a loop rather than
a person typing. One in a five-second tool loop is assumed to send about 0.4 operations per second,
where a human chat session sends 0.011 to 0.028. It is a population count on users, set with
--automated. It is a rate multiplier: how fast requests arrive.

They are independent, so each population carries its own traffic mix and the calculator blends them.

Where the numbers come from

The one measured anchor is 180 searches per second per 16-vCPU AMD EPYC server at 8 worker
processes, from a benchmark on 30 August 2026 with a real OpenAI embedder, a 12,000-episode corpus
and the rrf-hybrid reranker enabled — so the reranker's cost is already inside every derived server
count. Everything else is derived from it, estimated, or an assumption, and the README lists all of
them with their defaults.

What it deliberately does not do

It does not price high-availability additions — a second copy of every vector, a PostgreSQL standby
or a second gateway are a separate decision, and it says so rather than guessing.

Testing

382 tests, standard-library unittest, run with uv run --no-project python -m unittest from
tools/sizing. They cover the arithmetic at exact machine boundaries, every command-line
subcommand, the web form driven end to end, and rejection of bad input on every path. ruff check tools/sizing is clean.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng

The calculator turns a design peak in operations per second, plus a traffic
mix of adds, plain searches and agent-mode searches, into a hardware order:
API servers, vector-store machines and their RAM size, a PostgreSQL server,
embedding and agent-model GPU cards, hot vector RAM and disk for a chosen
retention period, the max_connections PostgreSQL will need, north-south and
east-west network peaks, and how many human chat sessions or automated clients
the capacity holds. It also converts a caller population back into the capacity
it demands, and prints a sensitivity table showing what the agent-mode search
rate costs in servers.

Two ideas that sound alike are kept apart throughout, because a reader who
confuses them sizes the deployment wrongly. Agent-mode search is a property of
a REQUEST: the agent_mode flag on a MemMachine search, which fans one search
out into about 22. It keeps the product's own name and the flag --agent, and
it is the third share of the traffic mix. An automated client is a property of
a CALLER: a program that sends requests in a loop, about 0.4 operations a
second where a human chat session sends 0.011 to 0.028. It is a population
count on the users subcommand, and its flag is --automated. One says how
expensive a request is, the other how fast a caller sends requests. The old
flag --agents, which was one letter from --agent and meant the caller, is
refused by name with a message that says which flag to use and why they are
different; the web address setting "agents" is answered the same way rather
than with a suggestion of "agent", which would be the wrong half of the
confusion.

Because a caller of either kind can send requests of either kind, each
population on the users subcommand carries its own traffic mix: --human-mix
and --automated-mix, each three numbers as adds/plain/agent-mode written with
"/" or with ",". Both default to the model's own default mix, so every existing
answer stays where it was. The subcommand reports the operations per second
each population demands, both mixes and the blended mix across the whole
population, the smallest tier that holds the demand, and the machines that
demand needs. The blended mix is the two mixes averaged, each weighted by the
operations its population demands at the busy end of the human rate, and it is
what sizes the machines - the program's default mix is not used once a
population is given. So "5,000 people who rarely use multi-hop search, plus 200
automated clients that use it constantly" is now a question the calculator can
answer, and it answers it with more hardware than the same operations per
second at the default mix would buy.

It deliberately does not price high-availability additions: a second copy of
every vector, a PostgreSQL standby or a second gateway are a separate
decision. Nor does it choose a machine class for the PostgreSQL and
vector-store machines. Only the API server has one that was benchmarked, so
the report gives the vector-store RAM the model itself chose and says the rest
is undecided rather than repeating the API server's vCPU and RAM for machines
nobody has sized.

Every number the program prints is labelled measured, derived, estimate or
assumption, and every input the model uses -- both the command-line flags and
the named constants in the file -- is listed with its default and its label in
tools/sizing/README.md. A measured label names the run it came from by its
date, its configuration, or both. Every input that tier, calc and users take as
a flag is also a box on the web form, so a reader can change any of them
without editing code; the RAM per vector-store machine is chosen automatically
when that box is left empty, and the two caller-population counts are empty
until somebody asks that question.

Every other box on the form has to be answered. An empty box on a form
submission is a blank answer, not a request for the default, so the page says
which box is blank instead of inventing a number and printing a hardware plan
for it. A parameter left out of a /api/calc URL altogether is different and
still takes the default, exactly as a command-line flag that is left off does.
A value that is not a number is reported with the name of the box it came
from and what was typed into it, and the message is announced to a screen
reader and points at the box that has to change. A setting the calculator does
not know is refused rather than ignored, and the suggested spelling ignores
case, so OPS and Ops are both answered with ops.

Nothing the reader typed is quietly changed. Vector dimensions must be a whole
number, and every input that scales the arithmetic carries an upper bound far
above any real deployment, so a fraction or a number too large to size is
refused by name rather than being cut down or overflowing to infinity part-way
through. Those bounds change no machine count. An input has one name across
every message about it -- the box labelled "Bytes per number" is called bytes
per number whether what was typed in it is zero or too large -- and the value
is quoted as it was given, so a refused 1,000,000,001 does not print as the
1,000,000,000 it is said to exceed. The report echoes its inputs the same way,
so half a byte per number reads as 0.5 above the tables that were sized for
half a byte, and the answer can be reproduced from the inputs printed with it.

Machine counts always round up, and any work at all costs at least one
machine. That covers the vector store as well as the API servers: a traffic
mix with no adds in it stores no bytes, and so does a retention of zero days,
but both still search vectors, so the order is never fewer than one
vector-store machine. A population of nobody demands no operations, and the
users report says there is nothing to size instead of naming a tier for it.

The report says only what the model knows. Headroom is measured against the
busier of the two network directions, which is what the row is labelled. The
sensitivity table marks the run's own row, because a rate of 2.04 and a rate
of 2.0 both print as 2.0 and are two different sizings.

The tool is a single file that imports only the standard library, so it runs
with `uv run --no-project python memmachine_sizing.py`. It offers five
subcommands (tier, calc, users, validate, serve), including a self-contained
web form and JSON endpoint. The page carries no external files, wide tables
scroll inside themselves so a narrow screen never drags the page sideways, and
/favicon.ico answers 204 so a page load leaves nothing in the browser console.
The test suite covers the model, the command line, the web form and its error
handling, and the standard-library-only constraint.

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng
@leomem
leomem merged commit ac353c0 into speedkick Aug 30, 2026
26 of 42 checks passed
@leomem
leomem deleted the add-sizing-calculator branch August 30, 2026 20:25
leomem added a commit that referenced this pull request Aug 30, 2026
The calculator turns a design peak in operations per second, plus a traffic
mix of adds, plain searches and agent-mode searches, into a hardware order:
API servers, vector-store machines and their RAM size, a PostgreSQL server,
embedding and agent-model GPU cards, hot vector RAM and disk for a chosen
retention period, the max_connections PostgreSQL will need, north-south and
east-west network peaks, and how many callers that capacity holds. It also
works backwards, turning a caller population into the capacity it demands.

Every number it prints is labelled measured, derived, estimate or assumption,
and every input is a command-line flag and a box on the web form, so a reader
changes any of them without editing code.

Two things that sound alike are kept apart throughout. Agent-mode search is a
property of one REQUEST - the agent_mode flag, which fans one search out into
about 22 - and it is a cost multiplier. An automated client is a property of a
CALLER - a program sending requests in a loop - and it is a rate multiplier.
Each caller population carries its own traffic mix and the two are blended,
because a person can send agent-mode searches and a program need not.

The unit everywhere is a session sending requests at the same moment, whoever
is behind it: one developer driving a ten-user load test is ten sessions. A
user count is not a session count, so converting one to the other takes two
figures - the share of users active at the busiest moment, and the sessions
one active user holds. Both ship as example defaults, labelled as such, with
a warning naming which were defaulted and what to replace them with. The
share is 10 per 100 people: a convention, not a measurement, plausibly 5 to
20, and every machine count moves with it.

The two per-caller rates carry published sources. A human chat session at
0.011 to 0.028 operations per second is the median to about the 90th
percentile of a session's busiest five minutes, measured across 55,295
sessions in the BurstGPT dataset (arXiv:2401.17644). An automated client at
0.4 is TraceLab's measured 5.0-second median step across roughly 4,300
production agent sessions (arXiv:2606.30560). Both reports also carry the
figures those sources revealed and no count uses: a heavy session at 0.06,
and an automated client at 0.07 sustained, six times below its burst because
an agent is idle most of the wall-clock time.

Sizing for the worst sustained five minutes follows ITU-T E.500, which
requires read-out periods greater than five minutes so that resources are not
dimensioned for infrequent small-interval peaks.

It deliberately does not price high-availability additions: a second copy of
every vector, a PostgreSQL standby or a second gateway are a separate
decision.

476 tests, standard-library unittest, covering the arithmetic at exact
machine boundaries, every subcommand, the web form driven end to end, and
rejection of bad input on every path. ruff check is clean.

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng
wanghy73 pushed a commit to wanghy73/MemMachine that referenced this pull request Sep 18, 2026
The calculator turns a design peak in operations per second, plus a traffic
mix of adds, plain searches and agent-mode searches, into a hardware order:
API servers, vector-store machines and their RAM size, a PostgreSQL server,
embedding and agent-model GPU cards, hot vector RAM and disk for a chosen
retention period, the max_connections PostgreSQL will need, north-south and
east-west network peaks, and how many callers that capacity holds. It also
works backwards, turning a caller population into the capacity it demands.

Every number it prints is labelled measured, derived, estimate or assumption,
and every input is a command-line flag and a box on the web form, so a reader
changes any of them without editing code.

Two things that sound alike are kept apart throughout. Agent-mode search is a
property of one REQUEST - the agent_mode flag, which fans one search out into
about 22 - and it is a cost multiplier. An automated client is a property of a
CALLER - a program sending requests in a loop - and it is a rate multiplier.
Each caller population carries its own traffic mix and the two are blended,
because a person can send agent-mode searches and a program need not.

The unit everywhere is a session sending requests at the same moment, whoever
is behind it: one developer driving a ten-user load test is ten sessions. A
user count is not a session count, so converting one to the other takes two
figures - the share of users active at the busiest moment, and the sessions
one active user holds. Both ship as example defaults, labelled as such, with
a warning naming which were defaulted and what to replace them with. The
share is 10 per 100 people: a convention, not a measurement, plausibly 5 to
20, and every machine count moves with it.

The two per-caller rates carry published sources. A human chat session at
0.011 to 0.028 operations per second is the median to about the 90th
percentile of a session's busiest five minutes, measured across 55,295
sessions in the BurstGPT dataset (arXiv:2401.17644). An automated client at
0.4 is TraceLab's measured 5.0-second median step across roughly 4,300
production agent sessions (arXiv:2606.30560). Both reports also carry the
figures those sources revealed and no count uses: a heavy session at 0.06,
and an automated client at 0.07 sustained, six times below its burst because
an agent is idle most of the wall-clock time.

Sizing for the worst sustained five minutes follows ITU-T E.500, which
requires read-out periods greater than five minutes so that resources are not
dimensioned for infrequent small-interval peaks.

It deliberately does not price high-availability additions: a second copy of
every vector, a PostgreSQL standby or a second gateway are a separate
decision.

476 tests, standard-library unittest, covering the arithmetic at exact
machine boundaries, every subcommand, the web form driven end to end, and
rejection of bad input on every path. ruff check is clean.

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng
(cherry picked from commit 4aeb020)
Signed-off-by: Haiyan Wang <[email protected]>
malatewang pushed a commit that referenced this pull request Sep 18, 2026
…to main) (#1680)

Add a deployment sizing calculator under tools/sizing (#1553)

The calculator turns a design peak in operations per second, plus a traffic
mix of adds, plain searches and agent-mode searches, into a hardware order:
API servers, vector-store machines and their RAM size, a PostgreSQL server,
embedding and agent-model GPU cards, hot vector RAM and disk for a chosen
retention period, the max_connections PostgreSQL will need, north-south and
east-west network peaks, and how many callers that capacity holds. It also
works backwards, turning a caller population into the capacity it demands.

Every number it prints is labelled measured, derived, estimate or assumption,
and every input is a command-line flag and a box on the web form, so a reader
changes any of them without editing code.

Two things that sound alike are kept apart throughout. Agent-mode search is a
property of one REQUEST - the agent_mode flag, which fans one search out into
about 22 - and it is a cost multiplier. An automated client is a property of a
CALLER - a program sending requests in a loop - and it is a rate multiplier.
Each caller population carries its own traffic mix and the two are blended,
because a person can send agent-mode searches and a program need not.

The unit everywhere is a session sending requests at the same moment, whoever
is behind it: one developer driving a ten-user load test is ten sessions. A
user count is not a session count, so converting one to the other takes two
figures - the share of users active at the busiest moment, and the sessions
one active user holds. Both ship as example defaults, labelled as such, with
a warning naming which were defaulted and what to replace them with. The
share is 10 per 100 people: a convention, not a measurement, plausibly 5 to
20, and every machine count moves with it.

The two per-caller rates carry published sources. A human chat session at
0.011 to 0.028 operations per second is the median to about the 90th
percentile of a session's busiest five minutes, measured across 55,295
sessions in the BurstGPT dataset (arXiv:2401.17644). An automated client at
0.4 is TraceLab's measured 5.0-second median step across roughly 4,300
production agent sessions (arXiv:2606.30560). Both reports also carry the
figures those sources revealed and no count uses: a heavy session at 0.06,
and an automated client at 0.07 sustained, six times below its burst because
an agent is idle most of the wall-clock time.

Sizing for the worst sustained five minutes follows ITU-T E.500, which
requires read-out periods greater than five minutes so that resources are not
dimensioned for infrequent small-interval peaks.

It deliberately does not price high-availability additions: a second copy of
every vector, a PostgreSQL standby or a second gateway are a separate
decision.

476 tests, standard-library unittest, covering the arithmetic at exact
machine boundaries, every subcommand, the web form driven end to end, and
rejection of bad input on every path. ruff check is clean.


Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng
(cherry picked from commit 4aeb020)

Signed-off-by: Haiyan Wang <[email protected]>
Co-authored-by: Leo Qiu <[email protected]>
Co-authored-by: Claude Fable 5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant