Repository navigation
Conversation
Signed-off-by: Edwin Yu <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 16, 2026
Improve embedding fidelity Signed-off-by: Edwin Yu <[email protected]>
This was referenced Sep 16, 2026
Merged
Closed
Merged
[session storage 2/2] Remove open-or-create from both stores, and close from the segment store
#1625
Draft
Closed
Draft
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 17, 2026
Improve embedding fidelity Signed-off-by: Edwin Yu <[email protected]>
malatewang
added a commit
that referenced
this pull request
Sep 18, 2026
* Improve embedding fidelity (speedkick) (#1587) Improve embedding fidelity Signed-off-by: Edwin Yu <[email protected]> * Clamp max_input_length to the per-request cap in the OpenAI embedder A configured max_input_length above max_total_input_length_per_request (75,000 code points) can never be honored: no chunk larger than the cap fits in any request. Before this change such a limit reached cluster_texts with an oversized chunk and raised ValueError (HTTP 500), the #1298 failure for that configuration. Greedy chunking widened which input lengths hit it: at max_input_length=80_000, balanced chunking crashed on 76,000-80,000 and 160,000-char inputs; greedy crashed on everything above 75,000. The regression test configures max_input_length=80_000 with a 100,000-char input; it fails before the clamp with the cluster_texts ValueError and passes after with two requests (75,000 + 25,000 chars). Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> --------- Signed-off-by: Edwin Yu <[email protected]> Co-authored-by: Shu Wang <[email protected]> Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Copy of #1449 onto
speedkick. Cherry-picked cleanly; no conflicts and no changes were needed. Targeted suiteserver_tests/memmachine_server/common/embedderpasses (5 passed).Purpose of the change
Chunk-merged embeddings (texts longer than the model input limit are split, embedded per chunk, and averaged) reproduce the whole-text embedding more faithfully with greedy max-size chunks and a length-weighted average than with the current balanced chunks and unweighted mean. Measured on three encoders; numbers below.
OpenAI servers no longer error when special tokens are given to text-embedding-3-small, so the special-token sanitizer is removed (probe script and results attached as a PR comment below).
Description
chunk_text) instead of balanced chunks (chunk_text_balanced) in the OpenAI, SentenceTransformer, and Amazon Bedrock embedders.np.averageinstead of unweightednp.mean.Measurements
Method: 10 hand-written texts in distinct registers (fiction, technical docs, chat transcript, news, legal, academic, recipe, business memo, product review, encyclopedia), each short enough to embed whole — the whole-text embedding is the observable ground truth. Each text is chunked at limits 800 / 1600 / 3200 chars under each policy, chunks embedded, merged, and compared to the whole-text vector by cosine similarity. Single-chunk cells excluded; 28 cells per configuration per model. Probe script and raw per-cell results are attached as PR comments below (deliberately not committed to the branch).
This PR's configuration (greedy + length-weighted) vs current behavior (balanced + unweighted), paired over the same 28 cells:
Full split × weighting grid (mean cosine to whole-text embedding, higher = more faithful):
Two structural checks that this is signal, not noise:
Mechanism: to the extent an encoder mean-pools token states, the whole-text vector is approximately a length-weighted average over the text, which a merge approximates best when weights match chunk lengths and most characters sit in maximal-context chunks. The result holds on the OpenAI models even though their pooling is undocumented.
Caveat: per-cell variance is wider on the OpenAI models than on gemma (worst cell −0.019, best +0.048), so this is a consistent aggregate win rather than a uniform per-text one.
Type of change
How Has This Been Tested?
Test Results: Fidelity tables above; probe scripts and raw results attached as PR comments below; embedder unit tests pass (
server_tests/memmachine_server/common/embedder/).Checklist
Maintainer Checklist
🤖 Generated with Claude Code
https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL