Repository navigation
Add a deployment sizing calculator under tools/sizing - #1553
Merged
Merged
Conversation
The calculator turns a design peak in operations per second, plus a traffic mix of adds, plain searches and agent-mode searches, into a hardware order: API servers, vector-store machines and their RAM size, a PostgreSQL server, embedding and agent-model GPU cards, hot vector RAM and disk for a chosen retention period, the max_connections PostgreSQL will need, north-south and east-west network peaks, and how many human chat sessions or automated clients the capacity holds. It also converts a caller population back into the capacity it demands, and prints a sensitivity table showing what the agent-mode search rate costs in servers. Two ideas that sound alike are kept apart throughout, because a reader who confuses them sizes the deployment wrongly. Agent-mode search is a property of a REQUEST: the agent_mode flag on a MemMachine search, which fans one search out into about 22. It keeps the product's own name and the flag --agent, and it is the third share of the traffic mix. An automated client is a property of a CALLER: a program that sends requests in a loop, about 0.4 operations a second where a human chat session sends 0.011 to 0.028. It is a population count on the users subcommand, and its flag is --automated. One says how expensive a request is, the other how fast a caller sends requests. The old flag --agents, which was one letter from --agent and meant the caller, is refused by name with a message that says which flag to use and why they are different; the web address setting "agents" is answered the same way rather than with a suggestion of "agent", which would be the wrong half of the confusion. Because a caller of either kind can send requests of either kind, each population on the users subcommand carries its own traffic mix: --human-mix and --automated-mix, each three numbers as adds/plain/agent-mode written with "/" or with ",". Both default to the model's own default mix, so every existing answer stays where it was. The subcommand reports the operations per second each population demands, both mixes and the blended mix across the whole population, the smallest tier that holds the demand, and the machines that demand needs. The blended mix is the two mixes averaged, each weighted by the operations its population demands at the busy end of the human rate, and it is what sizes the machines - the program's default mix is not used once a population is given. So "5,000 people who rarely use multi-hop search, plus 200 automated clients that use it constantly" is now a question the calculator can answer, and it answers it with more hardware than the same operations per second at the default mix would buy. It deliberately does not price high-availability additions: a second copy of every vector, a PostgreSQL standby or a second gateway are a separate decision. Nor does it choose a machine class for the PostgreSQL and vector-store machines. Only the API server has one that was benchmarked, so the report gives the vector-store RAM the model itself chose and says the rest is undecided rather than repeating the API server's vCPU and RAM for machines nobody has sized. Every number the program prints is labelled measured, derived, estimate or assumption, and every input the model uses -- both the command-line flags and the named constants in the file -- is listed with its default and its label in tools/sizing/README.md. A measured label names the run it came from by its date, its configuration, or both. Every input that tier, calc and users take as a flag is also a box on the web form, so a reader can change any of them without editing code; the RAM per vector-store machine is chosen automatically when that box is left empty, and the two caller-population counts are empty until somebody asks that question. Every other box on the form has to be answered. An empty box on a form submission is a blank answer, not a request for the default, so the page says which box is blank instead of inventing a number and printing a hardware plan for it. A parameter left out of a /api/calc URL altogether is different and still takes the default, exactly as a command-line flag that is left off does. A value that is not a number is reported with the name of the box it came from and what was typed into it, and the message is announced to a screen reader and points at the box that has to change. A setting the calculator does not know is refused rather than ignored, and the suggested spelling ignores case, so OPS and Ops are both answered with ops. Nothing the reader typed is quietly changed. Vector dimensions must be a whole number, and every input that scales the arithmetic carries an upper bound far above any real deployment, so a fraction or a number too large to size is refused by name rather than being cut down or overflowing to infinity part-way through. Those bounds change no machine count. An input has one name across every message about it -- the box labelled "Bytes per number" is called bytes per number whether what was typed in it is zero or too large -- and the value is quoted as it was given, so a refused 1,000,000,001 does not print as the 1,000,000,000 it is said to exceed. The report echoes its inputs the same way, so half a byte per number reads as 0.5 above the tables that were sized for half a byte, and the answer can be reproduced from the inputs printed with it. Machine counts always round up, and any work at all costs at least one machine. That covers the vector store as well as the API servers: a traffic mix with no adds in it stores no bytes, and so does a retention of zero days, but both still search vectors, so the order is never fewer than one vector-store machine. A population of nobody demands no operations, and the users report says there is nothing to size instead of naming a tier for it. The report says only what the model knows. Headroom is measured against the busier of the two network directions, which is what the row is labelled. The sensitivity table marks the run's own row, because a rate of 2.04 and a rate of 2.0 both print as 2.0 and are two different sizings. The tool is a single file that imports only the standard library, so it runs with `uv run --no-project python memmachine_sizing.py`. It offers five subcommands (tier, calc, users, validate, serve), including a self-contained web form and JSON endpoint. The page carries no external files, wide tables scroll inside themselves so a narrow screen never drags the page sideways, and /favicon.ico answers 204 so a page load leaves nothing in the browser console. The test suite covers the model, the command line, the web form and its error handling, and the standard-library-only constraint. Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng
leomem
added a commit
that referenced
this pull request
Aug 30, 2026
The calculator turns a design peak in operations per second, plus a traffic mix of adds, plain searches and agent-mode searches, into a hardware order: API servers, vector-store machines and their RAM size, a PostgreSQL server, embedding and agent-model GPU cards, hot vector RAM and disk for a chosen retention period, the max_connections PostgreSQL will need, north-south and east-west network peaks, and how many callers that capacity holds. It also works backwards, turning a caller population into the capacity it demands. Every number it prints is labelled measured, derived, estimate or assumption, and every input is a command-line flag and a box on the web form, so a reader changes any of them without editing code. Two things that sound alike are kept apart throughout. Agent-mode search is a property of one REQUEST - the agent_mode flag, which fans one search out into about 22 - and it is a cost multiplier. An automated client is a property of a CALLER - a program sending requests in a loop - and it is a rate multiplier. Each caller population carries its own traffic mix and the two are blended, because a person can send agent-mode searches and a program need not. The unit everywhere is a session sending requests at the same moment, whoever is behind it: one developer driving a ten-user load test is ten sessions. A user count is not a session count, so converting one to the other takes two figures - the share of users active at the busiest moment, and the sessions one active user holds. Both ship as example defaults, labelled as such, with a warning naming which were defaulted and what to replace them with. The share is 10 per 100 people: a convention, not a measurement, plausibly 5 to 20, and every machine count moves with it. The two per-caller rates carry published sources. A human chat session at 0.011 to 0.028 operations per second is the median to about the 90th percentile of a session's busiest five minutes, measured across 55,295 sessions in the BurstGPT dataset (arXiv:2401.17644). An automated client at 0.4 is TraceLab's measured 5.0-second median step across roughly 4,300 production agent sessions (arXiv:2606.30560). Both reports also carry the figures those sources revealed and no count uses: a heavy session at 0.06, and an automated client at 0.07 sustained, six times below its burst because an agent is idle most of the wall-clock time. Sizing for the worst sustained five minutes follows ITU-T E.500, which requires read-out periods greater than five minutes so that resources are not dimensioned for infrequent small-interval peaks. It deliberately does not price high-availability additions: a second copy of every vector, a PostgreSQL standby or a second gateway are a separate decision. 476 tests, standard-library unittest, covering the arithmetic at exact machine boundaries, every subcommand, the web form driven end to end, and rejection of bad input on every path. ruff check is clean. Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng
wanghy73
pushed a commit
to wanghy73/MemMachine
that referenced
this pull request
Sep 18, 2026
The calculator turns a design peak in operations per second, plus a traffic mix of adds, plain searches and agent-mode searches, into a hardware order: API servers, vector-store machines and their RAM size, a PostgreSQL server, embedding and agent-model GPU cards, hot vector RAM and disk for a chosen retention period, the max_connections PostgreSQL will need, north-south and east-west network peaks, and how many callers that capacity holds. It also works backwards, turning a caller population into the capacity it demands. Every number it prints is labelled measured, derived, estimate or assumption, and every input is a command-line flag and a box on the web form, so a reader changes any of them without editing code. Two things that sound alike are kept apart throughout. Agent-mode search is a property of one REQUEST - the agent_mode flag, which fans one search out into about 22 - and it is a cost multiplier. An automated client is a property of a CALLER - a program sending requests in a loop - and it is a rate multiplier. Each caller population carries its own traffic mix and the two are blended, because a person can send agent-mode searches and a program need not. The unit everywhere is a session sending requests at the same moment, whoever is behind it: one developer driving a ten-user load test is ten sessions. A user count is not a session count, so converting one to the other takes two figures - the share of users active at the busiest moment, and the sessions one active user holds. Both ship as example defaults, labelled as such, with a warning naming which were defaulted and what to replace them with. The share is 10 per 100 people: a convention, not a measurement, plausibly 5 to 20, and every machine count moves with it. The two per-caller rates carry published sources. A human chat session at 0.011 to 0.028 operations per second is the median to about the 90th percentile of a session's busiest five minutes, measured across 55,295 sessions in the BurstGPT dataset (arXiv:2401.17644). An automated client at 0.4 is TraceLab's measured 5.0-second median step across roughly 4,300 production agent sessions (arXiv:2606.30560). Both reports also carry the figures those sources revealed and no count uses: a heavy session at 0.06, and an automated client at 0.07 sustained, six times below its burst because an agent is idle most of the wall-clock time. Sizing for the worst sustained five minutes follows ITU-T E.500, which requires read-out periods greater than five minutes so that resources are not dimensioned for infrequent small-interval peaks. It deliberately does not price high-availability additions: a second copy of every vector, a PostgreSQL standby or a second gateway are a separate decision. 476 tests, standard-library unittest, covering the arithmetic at exact machine boundaries, every subcommand, the web form driven end to end, and rejection of bad input on every path. ruff check is clean. Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng (cherry picked from commit 4aeb020) Signed-off-by: Haiyan Wang <[email protected]>
malatewang
pushed a commit
that referenced
this pull request
Sep 18, 2026
…to main) (#1680) Add a deployment sizing calculator under tools/sizing (#1553) The calculator turns a design peak in operations per second, plus a traffic mix of adds, plain searches and agent-mode searches, into a hardware order: API servers, vector-store machines and their RAM size, a PostgreSQL server, embedding and agent-model GPU cards, hot vector RAM and disk for a chosen retention period, the max_connections PostgreSQL will need, north-south and east-west network peaks, and how many callers that capacity holds. It also works backwards, turning a caller population into the capacity it demands. Every number it prints is labelled measured, derived, estimate or assumption, and every input is a command-line flag and a box on the web form, so a reader changes any of them without editing code. Two things that sound alike are kept apart throughout. Agent-mode search is a property of one REQUEST - the agent_mode flag, which fans one search out into about 22 - and it is a cost multiplier. An automated client is a property of a CALLER - a program sending requests in a loop - and it is a rate multiplier. Each caller population carries its own traffic mix and the two are blended, because a person can send agent-mode searches and a program need not. The unit everywhere is a session sending requests at the same moment, whoever is behind it: one developer driving a ten-user load test is ten sessions. A user count is not a session count, so converting one to the other takes two figures - the share of users active at the busiest moment, and the sessions one active user holds. Both ship as example defaults, labelled as such, with a warning naming which were defaulted and what to replace them with. The share is 10 per 100 people: a convention, not a measurement, plausibly 5 to 20, and every machine count moves with it. The two per-caller rates carry published sources. A human chat session at 0.011 to 0.028 operations per second is the median to about the 90th percentile of a session's busiest five minutes, measured across 55,295 sessions in the BurstGPT dataset (arXiv:2401.17644). An automated client at 0.4 is TraceLab's measured 5.0-second median step across roughly 4,300 production agent sessions (arXiv:2606.30560). Both reports also carry the figures those sources revealed and no count uses: a heavy session at 0.06, and an automated client at 0.07 sustained, six times below its burst because an agent is idle most of the wall-clock time. Sizing for the worst sustained five minutes follows ITU-T E.500, which requires read-out periods greater than five minutes so that resources are not dimensioned for infrequent small-interval peaks. It deliberately does not price high-availability additions: a second copy of every vector, a PostgreSQL standby or a second gateway are a separate decision. 476 tests, standard-library unittest, covering the arithmetic at exact machine boundaries, every subcommand, the web form driven end to end, and rejection of bad input on every path. ruff check is clean. Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng (cherry picked from commit 4aeb020) Signed-off-by: Haiyan Wang <[email protected]> Co-authored-by: Leo Qiu <[email protected]> Co-authored-by: Claude Fable 5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
A sizing calculator for MemMachine deployments, under
tools/sizing/. Give it a design peak inoperations per second and it returns the hardware a deployment needs: API servers, vector-store
machines and their memory size, a PostgreSQL server, embedding and agent-model GPU cards, hot vector
memory and disk for a chosen retention period, the
max_connectionsPostgreSQL will need, thenetwork peaks, and how many callers that capacity holds.
Why it exists
Capacity questions about MemMachine were being answered in prose, and prose cannot be re-run. When
someone asks "what if we keep episodes for a year instead of ninety days", or "what if a tenth of
searches use agent mode instead of one in a hundred", the honest answer is a different hardware
order, not a paragraph. This puts the model in one file so the answer can be recomputed instead of
re-argued, and so every assumption behind it is visible.
Three properties matter more than the arithmetic:
glance which figures rest on a benchmark and which rest on a guess that has never been tested.
ask a different question.
traffic mix is a planning assumption nobody has measured. The tool says so where it uses them,
rather than presenting a confident total.
Usage
It needs no dependencies — Python standard library only.
tier pilot|target|scalecalc --ops Nusers --humans N --automated MvalidateserveRe-price an order against a different traffic mix:
Size for a mixed population, where each kind of caller sends a different mix:
Two things that sound alike and are not
The README leads with this, because confusing them sizes a deployment wrongly.
Agent-mode search is a property of a request — the
agent_modeflag on a MemMachine search.A search sent with it fans out into a multi-hop retrieval of roughly 22 embedding calls, 22 vector
searches, 44 database reads and one or two language-model calls. It is the third share of the
traffic mix, set with
--agent. It is a cost multiplier: how expensive one request is.An automated client is a property of a caller — a program sending requests in a loop rather than
a person typing. One in a five-second tool loop is assumed to send about 0.4 operations per second,
where a human chat session sends 0.011 to 0.028. It is a population count on
users, set with--automated. It is a rate multiplier: how fast requests arrive.They are independent, so each population carries its own traffic mix and the calculator blends them.
Where the numbers come from
The one measured anchor is 180 searches per second per 16-vCPU AMD EPYC server at 8 worker
processes, from a benchmark on 30 August 2026 with a real OpenAI embedder, a 12,000-episode corpus
and the rrf-hybrid reranker enabled — so the reranker's cost is already inside every derived server
count. Everything else is derived from it, estimated, or an assumption, and the README lists all of
them with their defaults.
What it deliberately does not do
It does not price high-availability additions — a second copy of every vector, a PostgreSQL standby
or a second gateway are a separate decision, and it says so rather than guessing.
Testing
382 tests, standard-library
unittest, run withuv run --no-project python -m unittestfromtools/sizing. They cover the arithmetic at exact machine boundaries, every command-linesubcommand, the web form driven end to end, and rejection of bad input on every path.
ruff check tools/sizingis clean.🤖 Generated with Claude Code
https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng