[Bug][v2]: Neo4j node labels #633
Description
Activity
- changed the title
[-][Bug][v2]: Neo4j node levels[/-][+][Bug][v2]: Neo4j node labels[/+]on Nov 27, 2025 Hi @edwinyyyu , could you please help to explain it here.
Neo4j labels can only contain certain characters or requiring escaping.
The upper layer passes in something like Episode_testuser/44d1e85c-7553-4a7b-b2fd-41f5957337b9. Here, the session id is testuser/44d1e85c-7553-4a7b-b2fd-41f5957337b9.
The special characters get replaced by uXXXX, and SANITIZED_ is prepended.My question is why is the label encoding so much information? I'd expect this to be
Episodeand not a node label for each session...Neo4j vector indexes can only be created for a label-property pair. If a single label is used for all episodes, then data in one session will affect the search performance for all other sessions. Using a separate label for each session allows creating a separate vector index for each session.
ah ok - that makes sense. Even though it will create a lot of labels...
I will close this now. Feel free to reopen if you have any follow-ups.
I am going to have a slot at an Neo4j conference, thus I'd like to open this up again to understand it better. :)
@edwinyyyu
When using Neo4j assemantic_memoryI can see them being build up in Neo4j.
Relationships
Embeddings, features and facts are linked (it seems) via IDs - kinda SQL style. Why are they not linked with a relationship?
Traversing these relationships should be much cheaper than doing multipleMATCHstatements toJOINnodes...MultiLabel
For features you use two labels, one generic one and one specific one.
{ "identity": 19, "labels": [ "Feature", "FeatureSet_mem_u5f_session_u5f_org_u5f_v1_u2f_project_u5f_v1" ],That makes a search across all nodes of the same type easier I suppose. Can this be done for other nodes as well?
On another note; Why are embeddings for the same phrase different? running a string through an embedder should yield the same result, right?
The graph above has the same text for two different orgs/projects have slightly different embeddings. Is this used to segment off different sessions?@o-love should know more about semantic memory.
Embeddings are not entirely deterministic. The cosine similarity should still be something like >0.99 so it doesn't matter in practice.
Hey @ChristianKniep
Semantic Memory's main implementation is in Postgres. With Neo4j being added as a way to minimize MemMachine's total dependencies for lightweight users.
I've also never really touched Neo4j or any graph database apart from adding the support for Semantic Memory.So all that probably contributes to the Neo4j usage for semantic memory looking more SQL like.
Regarding the use of multiple labels. I use that for different types of search. With set id specific queries using the label with the set id and non set id specific queries using the general label.
Regarding the different embeddings. I agree with what Edwin mentioned.
@o-love Do you plan to streamline the usage of Neo4j towards best-practice? I might help (even though I need to wrap my head around the python code for that).
For the presentation at the Neo4j event, it would be awesome if I can show a more connected graph.@ChristianKniep I have no plans or thoughts on modifying the Neo4j for semantic.
Describe the bug
Within v2 the labels of nodes encode the individual episodes within the label of sorts.
I'd say that won't scale nicely as the index neo4j needs to keep to get startet is going to grow wich each episode.
Is that a deliberate choice or just an initial artefact of pushing towards v2?
Steps to reproduce
compse file
Expected behavior
I'd expect to have a node label
Episodeand properties to encode the specifics.Environment
compose
Additional context
No response