Repository navigation
Conversation
edwinyyyu
force-pushed
the
feat/turbovec-engine-speedkick
branch
from
September 8, 2026 18:59
77ae415 to
0b4d558
Compare
edwinyyyu
force-pushed
the
feat/turbovec-engine-speedkick
branch
8 times, most recently
from
September 9, 2026 18:50
f27ab80 to
7719081
Compare
turbovec keeps TurboQuant-compressed vectors in RAM, so an index is a fraction of the size of the f32 engines' and a full scan has a much smaller scaling constant than sqlite-vec's on-disk one. Search is approximate as a result: scores land near the exact value rather than on it, and the tests assert ranking and membership instead of magnitudes. Two consequences of storing only compressed vectors are worth naming. `get_vectors` raises `NotImplementedError`, since the originals are not recoverable -- which makes this engine usable by EventMemory but not by semantic memory, whose feature updates read stored embeddings back. And removal is exact and cheap: turbovec drops the id from its map rather than tombstoning, so a deleted key cannot resurface in results. `save` publishes through the shared atomic-write helper, like the other engines, so an interrupted save leaves the previously published index intact rather than a truncated file the store would treat as a hard load error. `load` clears any temp file a previous save left behind. Signed-off-by: Edwin Yu <[email protected]> Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
1.0.0 is upstream's first stable release, and what it commits to is the on-disk format: v7 is the only container turbovec reads or writes, and a file written by this release stays readable by later ones. The extra floors there rather than at 0.7.0. v7 is also what makes `sync` possible, and that changes how this engine saves. `save` no longer routes through the shared `atomic_index_write` helper. turbovec publishes the index itself: a checkpoint appends what changed since the last one rather than restating the whole index, and commits it durably -- a crash at any byte leaves the previous commit intact, and unlike the helper's rename the publication survives a power failure. Wrapping that would restate the index on every checkpoint and publish it through the weaker of the two protocols. `load` drops `clear_stale_index_temp` with it, since the engine no longer writes the `<path>.tmp` sibling it cleared. The tests follow. Save leaving no temp file is now an assertion about the whole directory, since turbovec names its own temp; the stale-temp test goes with the protocol it tested; and a new test pins what a later checkpoint carries, reading a bulk removal and an append back through a fresh engine. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Signed-off-by: Edwin Yu <[email protected]>
`SearchMatch` documents a cosine score as a cosine similarity in [-1, 1], and `SQLiteVectorStoreCollection` compares `score_threshold` against it directly. turbovec's score is a quantized inner product, so it lands a hair outside that range: at 8 dimensions and the default bit width, 40 of 50 self-matches score above 1.0 (max 1.0066), and at bit width 2 they exceed it at every dimension measured (1.0038 at 768, 1.0025 at 1536). Publishing 1.0066 as a cosine similarity is a broken promise, however small, so the hair is clamped where the match is built. A dot product is unbounded by contract and passes through untouched. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Signed-off-by: Edwin Yu <[email protected]>
turbovec indexes a dimensionality that is a positive multiple of 8, so the engine refused a width every other engine in the tree accepts -- 300, the width of a spaCy or word2vec vector, among them -- with turbovec's own error rather than a store-level one. The width is now rounded up and vectors are written into the leading columns of a zeroed buffer. Padding is exact for both metrics this engine serves: a zero coordinate adds nothing to an inner product and nothing to an L2 norm, so the padded index answers as the unpadded one would. The same write is what rejects a wrong-width vector, which previously reached turbovec as a shape it would report in its own terms. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Signed-off-by: Edwin Yu <[email protected]>
`prepare()` warms turbovec's per-index caches and its lazy id-to-slot map so the first search, contains or remove after a write does not pay a one-time cost. Against turbovec 0.8.0 that cost was the whole SIMD-blocked code layout, rebuilt per add and proportional to the index rather than the batch -- 199 ms at 100k vectors, which the first unlucky search paid if this call did not. Upstream made that re-convergence incremental in 1.0.0, and the call now costs 12-20 us and buys nothing measurable: at 100k a search after a one-vector add reads 0.52 ms with it and 0.51 ms without, and in a cold process the first search reads 0.10 ms against 0.09 ms. The remaining warm it offers -- the id-to-slot map after a load, worth 0.5 ms off the first remove at 100k -- is not on this path, and the map materializes safely under concurrent readers regardless. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Signed-off-by: Edwin Yu <[email protected]>
`VectorStoreCollection.get` had one production caller: the semantic storage read a feature's stored embedding back so it could write that same embedding again alongside fresh properties, because `upsert` demands a whole record. `return_vector=True` was passed at exactly that one call site, and `VectorSearchEngine.get_vectors` existed to serve it. Replace it with `set_properties`, which says what the caller wanted: replace a record's properties and leave its vector where it is. On SQLite and sqlite-vec that is a plain UPDATE on the records table that never touches the index. On Qdrant it is an `overwrite_payload` selected by filter rather than by id -- an id list raises on an id the collection does not hold, and would reach a point another logical collection owns in the same native one. Milvus has no partial update, so the read-back survives there, inside the one backend that needs it. With no caller left, `get`, `return_vector` and `get_vectors` all go. turbovec could not implement `get_vectors` at all -- it stores only TurboQuant-compressed vectors -- so this also retires the one contract method a conforming engine was allowed to raise on. Two consequences worth naming: - A feature's embedding is a function of its `value` alone, on add and on update alike, so keeping the stored vector when `value` did not change is the behavior that was already there. What changes is that it is no longer expressed as a read-modify-write through the index. - `get` was the only read that could see a record whose vector the index lost to a reverted publication. `query` goes through the index, so such a record is now invisible through the contract; the row survives, and the durability test reads the row directly to say so. Tests read records back through `query` instead, via a `_fetch_records` helper per suite. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL
edwinyyyu
force-pushed
the
feat/turbovec-engine-speedkick
branch
from
September 9, 2026 18:53
7719081 to
c1b3008
Compare
This was referenced Sep 9, 2026
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Copy of #1499 onto
speedkick, stacked on #1588 (thespeedkickcopy of #1460) — its five commits are the first five here, so review only the last six. Cherry-picked cleanly; no conflicts.Beyond the copy, this carries the change asked for on the port:
getis gone from the vector store contract andget_vectorsfrom the engine.Taking vector read-back out of the contract
VectorStoreCollection.gethad exactly one production caller.VectorStoreSemanticStorage.update_featureread a feature's stored embedding back so it could write that same embedding again alongside fresh properties, becauseupsertdemands a whole record and the vector record's properties mirror the relational row.return_vector=Truewas passed at that one call site and nowhere else, andVectorSearchEngine.get_vectorsexisted to serve it.It is replaced by
set_properties, which says what the caller actually wanted — replace a record's properties, leave its vector alone:UPDATEon the records table. Properties live in SQL and vectors live in the index, so this never touches the index — no pending operation, no index save, nothing for a crash to lose.overwrite_payloadselected by filter, not by id. An id list raises on an id the collection does not hold, and would reach a point another logical collection owns in the same native collection; the filter form is scoped to the partition and matches nothing when the record is absent.With no caller left,
get,return_vectorandget_vectorsall go. Worth noting for this PR in particular: turbovec could not implementget_vectorsat all — it stores only TurboQuant-compressed vectors, so the method it contributed wasraise NotImplementedError. Removing it retires the one contract method a conforming engine was allowed to refuse.On the semantic memory question: where does the
update_featureembedding come from?SemanticMemory.add_featurecomputes it asembedder.ingest_embed([value])andSemanticMemory.update_featurerecomputes it the same way, but only whenvalueis not None. So the embedding is a function ofvaluealone, on both paths — reusing the stored vector whenvaluedid not change was already correct, and the other two backends (sqlalchemy_pgvector_semantic,neo4j_semantic_storage) express it as a partialUPDATEthat simply leaves the embedding column alone.What looked weird was not the reuse, it was that this one backend expressed it as a read-modify-write through the search index — the least reliable place in the store to route a value that never needed to move. That is what is gone.
Two consequences worth calling out
A record whose vector the index lost is now invisible through the contract.
getwas the only read that could still see it;querygoes through the index. The row survives, andtest_a_reverted_publication_costs_search_not_the_record_rowreads the row directly to say so. This also makes [sqlite store fixes 1/7] Publish vector index files atomically (but not durably) #1460's last commit ("Report a lost embedding as lost, not as a missing feature") moot — the code it added is deleted here, since there is no longer a read-back that can encounter the case.getwas load-bearing for tests, not for the server. About 40 test call sites used it as their read-by-UUID observation point. They now read throughqueryvia a_fetch_recordshelper per suite, which lists the collection with an arbitrary probe vector and no score threshold. If that trade is not wanted, the alternative is to keep a properties-onlygetin the contract —get_vectors,return_vectorand the read-modify-write all still go, and the tests keep their affordance.Verification
pytest packages/server/server_tests: 2013 passed, 3 skipped. The one failure,test_get_version, is a git-describe version-string artifact of the local checkout and fails identically on the untouched base branch (0.3.9.post2.dev10+g4105a720f).ty check --project packages/server: 17 diagnostics, the same 17 as the base branch.ruff check/ruff format --check/uv lock --check: clean.Below: the original description of #1499.
Purpose of the change
SQLiteVectorStorebacked by turbovec is smaller and faster thanSQLiteVecVectorStorebacked by sqlite-vec. Both are linear full scans, butturbovec keeps TurboQuant-compressed vectors in RAM, so its scaling constant is
much smaller and an index is a fraction of the size of the f32 engines'.
Description
Adds
TurboVecVectorSearchEngine, aVectorSearchEnginebacked by turbovec,behind a new
turbovecoptional extra floored at 1.0.0. turbovec indexes adimensionality that is a multiple of 8, so any other width is zero-padded up to
one -- exact for both metrics here, since a zero coordinate adds nothing to an
inner product or to an L2 norm -- and every width the other engines accept
works.
Publication is turbovec's own. 1.0.0 is upstream's first stable release, and
what it commits to is the on-disk format: v7 is the only container turbovec
reads or writes, and a file written by 1.0.0 stays readable by later releases.
v7 is also what makes
syncpossible, andsavecalls it -- a checkpointappends what changed since the last one rather than restating the whole index,
and commits it durably, so a crash at any byte leaves the previous commit
intact and the publication survives a power failure. The engine therefore does
not route through the shared
atomic_index_writehelper that the hnswlib andusearch engines use on #1460: wrapping
syncwould restate the index on everycheckpoint and publish it through the weaker of the two protocols, a rename the
helper's own docstring declines to make durable.
That is what supersedes #1448, which added the same engine against turbovec
0.x, whose
writehad no publication protocol of its own and wrote straight tothe final path -- an interrupted save left a truncated file where the store
expects a loadable index, and because
index_savedmakes a published index adurable contract, that is a hard
IndexLoadErrorrather than a silent rebuild.Known limitations, both consequences of storing only compressed vectors:
get_vectorsraisesNotImplementedError, since the original vectors are notrecoverable. That makes the engine usable by EventMemory but not by semantic
memory, whose feature updates read stored embeddings back.
it, so the tests assert ranking and membership rather than magnitudes. The
quantized inner product also overshoots the cosine range -- 40 of 50
self-matches at 8 dimensions, and every dimension measured at bit width 2 --
so
SearchMatchscores are clamped to[-1, 1]under a cosine metric, whichis what its docstring promises and what
score_thresholdcompares against.A dot product is unbounded by contract and is not clamped.
Removal is a strong point by comparison: turbovec drops the id from its map
rather than tombstoning it, so deletions stay cheap and a removed key cannot
resurface in results.
Type of change
How Has This Been Tested?
test_turbovec_engine.pycovers construction, add, remove, cosine and dotsearch, filtered search,
get_vectorsraising, score range (clamped undercosine, untouched under dot), unaligned widths (search, save/load, and a
wrong-width vector refused), and persistence -- a save/load
round-trip, a load replacing a live index, a save leaving nothing but the index
behind, and what a later checkpoint carries, reading a bulk removal and an
append back through a fresh engine. The module is
importorskip-guarded, so thesuite still runs without the extra installed.
Test Results:
uv run pytest packages/server/server_tests/memmachine_server/common/vector_store-> 346 passed.
ruff checkandruff format --checkclean. Thetyjobs arered on
mainitself (nebulagraph_python.clientlost the members the graphstore imports), fixed by #1519 rather than here; nothing
tyreports is inthis diff.
Checklist
Maintainer Checklist
🤖 Generated with Claude Code
https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL