Repository navigation
Add a deployment sizing calculator under tools/sizing (port of #1553 to main) - #1680
Merged
malatewang merged 1 commit intoSep 18, 2026
Merged
Conversation
malatewang
approved these changes
Sep 17, 2026
The calculator turns a design peak in operations per second, plus a traffic mix of adds, plain searches and agent-mode searches, into a hardware order: API servers, vector-store machines and their RAM size, a PostgreSQL server, embedding and agent-model GPU cards, hot vector RAM and disk for a chosen retention period, the max_connections PostgreSQL will need, north-south and east-west network peaks, and how many callers that capacity holds. It also works backwards, turning a caller population into the capacity it demands. Every number it prints is labelled measured, derived, estimate or assumption, and every input is a command-line flag and a box on the web form, so a reader changes any of them without editing code. Two things that sound alike are kept apart throughout. Agent-mode search is a property of one REQUEST - the agent_mode flag, which fans one search out into about 22 - and it is a cost multiplier. An automated client is a property of a CALLER - a program sending requests in a loop - and it is a rate multiplier. Each caller population carries its own traffic mix and the two are blended, because a person can send agent-mode searches and a program need not. The unit everywhere is a session sending requests at the same moment, whoever is behind it: one developer driving a ten-user load test is ten sessions. A user count is not a session count, so converting one to the other takes two figures - the share of users active at the busiest moment, and the sessions one active user holds. Both ship as example defaults, labelled as such, with a warning naming which were defaulted and what to replace them with. The share is 10 per 100 people: a convention, not a measurement, plausibly 5 to 20, and every machine count moves with it. The two per-caller rates carry published sources. A human chat session at 0.011 to 0.028 operations per second is the median to about the 90th percentile of a session's busiest five minutes, measured across 55,295 sessions in the BurstGPT dataset (arXiv:2401.17644). An automated client at 0.4 is TraceLab's measured 5.0-second median step across roughly 4,300 production agent sessions (arXiv:2606.30560). Both reports also carry the figures those sources revealed and no count uses: a heavy session at 0.06, and an automated client at 0.07 sustained, six times below its burst because an agent is idle most of the wall-clock time. Sizing for the worst sustained five minutes follows ITU-T E.500, which requires read-out periods greater than five minutes so that resources are not dimensioned for infrequent small-interval peaks. It deliberately does not price high-availability additions: a second copy of every vector, a PostgreSQL standby or a second gateway are a separate decision. 476 tests, standard-library unittest, covering the arithmetic at exact machine boundaries, every subcommand, the web form driven end to end, and rejection of bad input on every path. ruff check is clean. Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01DiqwYpXVGDet5NiHjdU7Ng (cherry picked from commit 4aeb020) Signed-off-by: Haiyan Wang <[email protected]>
wanghy73
force-pushed
the
port-main/sizing-calculator-1553
branch
from
September 18, 2026 00:36
5e474f1 to
3a7d47c
Compare
This was referenced Sep 18, 2026
Merged
Merged
[session storage 2/2] Remove open-or-create from both stores, and close from the segment store
#1625
Draft
Closed
Draft
[qdrant options] Let a deployment tune a Qdrant collection's HNSW, optimizers and quantization
#1618
Draft
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
A sizing calculator for MemMachine deployments, under
tools/sizing/. Give it a design peak inoperations per second and it returns the hardware a deployment needs: API servers, vector-store
machines and their memory size, a PostgreSQL server, embedding and agent-model GPU cards, hot vector
memory and disk for a chosen retention period, the
max_connectionsPostgreSQL will need, thenetwork peaks, and how many callers that capacity holds.
Why it exists
Capacity questions about MemMachine were being answered in prose, and prose cannot be re-run. When
someone asks "what if we keep episodes for a year instead of ninety days", or "what if a tenth of
searches use agent mode instead of one in a hundred", the honest answer is a different hardware
order, not a paragraph. This puts the model in one file so the answer can be recomputed instead of
re-argued, and so every assumption behind it is visible.
Three properties matter more than the arithmetic:
glance which figures rest on a benchmark and which rest on a guess that has never been tested.
ask a different question.
traffic mix is a planning assumption nobody has measured. The tool says so where it uses them,
rather than presenting a confident total.
Usage
It needs no dependencies — Python standard library only.
tier pilot|target|scalecalc --ops Nusers --humans N --automated MvalidateserveRe-price an order against a different traffic mix:
Size for a mixed population, where each kind of caller sends a different mix:
Two things that sound alike and are not
The README leads with this, because confusing them sizes a deployment wrongly.
Agent-mode search is a property of a request — the
agent_modeflag on a MemMachine search.A search sent with it fans out into a multi-hop retrieval of roughly 22 embedding calls, 22 vector
searches, 44 database reads and one or two language-model calls. It is the third share of the
traffic mix, set with
--agent. It is a cost multiplier: how expensive one request is.An automated client is a property of a caller — a program sending requests in a loop rather than
a person typing. One in a five-second tool loop is assumed to send about 0.4 operations per second,
where a human chat session sends 0.011 to 0.028. It is a population count on
users, set with--automated. It is a rate multiplier: how fast requests arrive.They are independent, so each population carries its own traffic mix and the calculator blends them.
Where the numbers come from
The one measured anchor is 180 searches per second per 16-vCPU AMD EPYC server at 8 worker
processes, from a benchmark on 30 August 2026 with a real OpenAI embedder, a 12,000-episode corpus
and the rrf-hybrid reranker enabled — so the reranker's cost is already inside every derived server
count. Everything else is derived from it, estimated, or an assumption, and the README lists all of
them with their defaults.
What it deliberately does not do
It does not price high-availability additions — a second copy of every vector, a PostgreSQL standby
or a second gateway are a separate decision, and it says so rather than guessing.
Testing
382 tests, standard-library
unittest, run withuv run --no-project python -m unittestfromtools/sizing. They cover the arithmetic at exact machine boundaries, every command-linesubcommand, the web form driven end to end, and rejection of bad input on every path.
ruff check tools/sizingis clean.