Skip to content

fix(event-backend): make expand_context return timeline-neighbor episodes (speedkick) - #1547

Merged
edwinyyyu merged 4 commits into
MemMachine:speedkickfrom
edwinyyyu:fix/event-backend-expand-context-speedkick
Aug 28, 2026
Merged

edwinyyyu merged 4 commits into
MemMachine:speedkickfrom
edwinyyyu:fix/event-backend-expand-context-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Copy of #1541 based on speedkick instead of main — the main commits cherry-pick cleanly. See #1540 for the bug report; #1541 carries the closing reference.

Problem

On the event backend, expand_context was silently inert on the search path: EventMemory.query fetched and materialized the expanded segment windows (a LATERAL query per direction), but LongTermMemory._search_scored_event read only the seed segment's _episode_uid and score from each window, and the response schema has no context field. Responses were byte-identical for expand_context 0 and 5, while every request paid for the expansion.

Fix

Bring the event backend to parity with the declarative backend's context handling (DeclarativeMemory._unify_scored_anchored_episode_contexts):

  • Each scored window contributes the episodes its segments belong to — chronological within the window, the seed's episode as the nucleus. (Every segment carries its _episode_uid, so segment-space windows map directly onto episode-space context; with the passthrough segmenter this is 1:1 timeline neighbors.)
  • Windows are unified best-score-first with the same fill algorithm: taken whole while they fit within num_episodes_limit, then filled by weighted index-proximity to the nucleus (forward recall preferred over backward, same weighting as declarative) until the limit is met. An episode keeps the score of the first window that contributed it.
  • The unified context is returned chronologically, matching the declarative backend's ordering contract for expanded results.
  • expand_context is clamped to num_episodes_limit - 1 (declarative parity).

expand_context == 0 behavior is unchanged — score-ordered seed episodes exactly as before, so existing callers see no difference unless they pass the parameter. Reranked configurations gain the same folding on top of reranker-scored windows (which already consumed the segments for scoring).

Follow-up self-review (third and fourth commits)

Two things the first two commits got wrong:

  • expand_context could go negative. The quota clamp min(max(0, expand_context), num_episodes_limit - 1) is -1 when num_episodes_limit == 0, which SearchMemoriesSpec.top_k allows (no lower bound). EventMemory._query then derives max_backward_segments = -1 and asks the segment store for a negative window — undefined by the SegmentStorePartition contract; the SQLAlchemy store happens to short-circuit on <= 0, the in-memory store computes an empty slice and drops the seed. The floor is now applied last.
  • Neither end-to-end test could tell the fix from its absence — both passed unmodified against the pre-fix long_term_memory.py. FakeEmbedder maps text to [len(text), -len(text)], so under cosine every document scores exactly 1.0 against every query: all seven timeline episodes are seeds of equal rank, ties keep insertion order (chronological), and num_episodes_limit=7 returns all seven with or without expansion. The "expansion adds episodes" assertion compared a limit-2 search against a limit-7 one, so the limit alone explained the difference.

The expansion tests now give each episode its own similarity from an explicit search rank, with the match's four timeline neighbours ranked last. That lets them assert the contract rather than an outcome — an exact episode list or exact score values would be pinned to the fixture's ranking and to the expand_context // 3 split, neither of which the fix claims:

  • No correct top-k can return those neighbours, and any nonzero window around the match reaches at least one of them whatever the backward/forward split, so "expansion returned a neighbour the search itself would not" survives a change of ranking or of the split. Chronological order and the episode limit are asserted alongside it.
  • The clamp is asserted on the call made to the segment store (0 <= backward + forward <= limit - 1, over several limit/expand_context pairs) rather than on which episodes come back.
  • expand_context == 0 is asserted as "matches only, best score first", without naming them.

Exact lists and score values stay in the unit tests, which own the fill algorithm and the score-retention rule. All three expansion tests fail against the code they cover.

Tests

  • End-to-end through the in-memory event-backend wiring: expansion returns neighbor episodes beyond the plain matches, contiguous on the timeline, chronologically ordered, and never exceeding num_episodes_limit.

  • Unit tests for the window→episode-uid extraction (dedup, nucleus identification) and the unification algorithm (whole-context fit, overflow proximity with forward preference, first-window score retention).

  • The expansion tests fail against the pre-fix long_term_memory.py; the unit tests cover the window→episode-uid extraction and the unification algorithm directly.

ruff (pinned 0.15.14) format + check clean; ty check packages/server reports no diagnostics from the touched files. Episodic + server test suites pass on this branch (528 passed, 1 skipped).

🤖 Generated with Claude Code

https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

edwinyyyu and others added 2 commits August 28, 2026 15:32
…odes

Fixes MemMachine#1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>
Co-Authored-By: Claude Fable 5 <[email protected]>
edwinyyyu and others added 2 commits August 28, 2026 16:02
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp
The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp
@edwinyyyu
edwinyyyu merged commit 7a66113 into MemMachine:speedkick Aug 28, 2026
35 of 39 checks passed
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 2, 2026
Two conflicts, both resolved by keeping each side:

- `database_manager.py`: speedkick (MemMachine#1552) factored the engine keywords into
  `_sql_engine_kwargs` and added asyncpg's command/connect timeouts, exactly
  where this branch adds `enable_sqlite_foreign_keys` and its inline keyword
  building. Kept speedkick's helper as the keyword source and this branch's
  pragma hook on the constructed engine.
- `test_event_backend_wiring.py`: `import logging` (this branch) against
  `import math` (speedkick, from MemMachine#1547). Kept both.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp
This was referenced Sep 16, 2026
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 17, 2026
…odes (speedkick) (MemMachine#1547)

* fix(event-backend): make expand_context return timeline-neighbor episodes

Fixes MemMachine#1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: ruff format

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(event-backend): clamp expand_context above zero, and make the
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

* test(event-backend): assert the expansion contract, not a ranking

The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

---------

Co-authored-by: Claude Fable 5 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 17, 2026
…odes (speedkick) (MemMachine#1547)

* fix(event-backend): make expand_context return timeline-neighbor episodes

Fixes MemMachine#1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: ruff format

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(event-backend): clamp expand_context above zero, and make the
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

* test(event-backend): assert the expansion contract, not a ranking

The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

---------

Co-authored-by: Claude Fable 5 <[email protected]>

@edwinyyyu edwinyyyu left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Found while auditing the main version, #1541 (line comment).

assert len(expanded) <= 3


async def test_expand_context_zero_returns_matches_in_score_order(

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

test_expand_context_zero_returns_matches_in_score_order cannot fail against the pre-fix code: at expand_context=0 the new clamp is the identity and the if expand_context > 0: branch is not taken, and the source change is purely additive, so pre- and post-fix execute the same statements. The other two expansion tests do discriminate; the description's "all three expansion tests fail against the code they cover" is two. Same on #1541.

edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
…odes (speedkick) (MemMachine#1547)

* fix(event-backend): make expand_context return timeline-neighbor episodes

Fixes MemMachine#1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: ruff format

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(event-backend): clamp expand_context above zero, and make the
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

* test(event-backend): assert the expansion contract, not a ranking

The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

---------

Co-authored-by: Claude Fable 5 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
…odes (speedkick) (MemMachine#1547)

* fix(event-backend): make expand_context return timeline-neighbor episodes

Fixes MemMachine#1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: ruff format

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(event-backend): clamp expand_context above zero, and make the
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

* test(event-backend): assert the expansion contract, not a ranking

The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

---------

Co-authored-by: Claude Fable 5 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
…odes (speedkick) (MemMachine#1547)

* fix(event-backend): make expand_context return timeline-neighbor episodes

Fixes MemMachine#1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: ruff format

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(event-backend): clamp expand_context above zero, and make the
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

* test(event-backend): assert the expansion contract, not a ranking

The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

---------

Co-authored-by: Claude Fable 5 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…odes (speedkick) (MemMachine#1547)

* fix(event-backend): make expand_context return timeline-neighbor episodes

Fixes MemMachine#1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: ruff format

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(event-backend): clamp expand_context above zero, and make the
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

* test(event-backend): assert the expansion contract, not a ranking

The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

---------

Co-authored-by: Claude Fable 5 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…odes (speedkick) (MemMachine#1547)

* fix(event-backend): make expand_context return timeline-neighbor episodes

Fixes MemMachine#1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: ruff format

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(event-backend): clamp expand_context above zero, and make the
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

* test(event-backend): assert the expansion contract, not a ranking

The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

---------

Co-authored-by: Claude Fable 5 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 25, 2026
…odes (speedkick) (MemMachine#1547)

* fix(event-backend): make expand_context return timeline-neighbor episodes

Fixes MemMachine#1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: ruff format

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(event-backend): clamp expand_context above zero, and make the
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

* test(event-backend): assert the expansion contract, not a ranking

The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

---------

Co-authored-by: Claude Fable 5 <[email protected]>
edwinyyyu added a commit that referenced this pull request Sep 25, 2026
…odes (#1541)

* fix(event-backend): make expand_context return timeline-neighbor episodes (speedkick) (#1547)

* fix(event-backend): make expand_context return timeline-neighbor episodes

Fixes #1540.

On the event backend, expand_context was silently inert: EventMemory
fetched and materialized the expanded segment windows, but
LongTermMemory._search_scored_event read only the seed segment's
_episode_uid and score from each window, and the response schema has no
context field - so responses were byte-identical for expand_context 0
and 5 while every request paid the LATERAL fetch.

The declarative backend, by contrast, folds neighbor episodes into the
returned list (_unify_scored_anchored_episode_contexts). This brings
the event backend to parity:

- Each scored window now contributes the episodes its segments belong
  to (chronological within the window, the seed's episode as nucleus).
- Windows are unified best-score-first with the same fill algorithm as
  the declarative backend: taken whole while they fit within
  num_episodes_limit, then filled by weighted index-proximity to the
  nucleus (forward recall preferred) until the limit is met; an episode
  keeps the score of the first window that contributed it.
- The unified context is returned chronologically, matching the
  declarative backend's ordering contract for expanded results.
- expand_context is clamped to num_episodes_limit - 1 (declarative
  parity).

expand_context == 0 behavior is unchanged (score-ordered seeds, exactly
as before). Reranked configurations gain the same folding on top of
reranker-scored windows.

Tests: end-to-end via the in-memory event-backend wiring (neighbors
returned, chronological order, limit respected) and unit tests for the
window-to-episode-uid extraction and the unification algorithm
(whole-context fit, overflow proximity with forward preference,
first-window score retention).

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: ruff format

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix(event-backend): clamp expand_context above zero, and make the
expansion tests actually discriminate

Self-review of the two commits above turned up one defect and one hole.

Defect: the quota clamp `min(max(0, expand_context), num_episodes_limit - 1)`
goes negative when `num_episodes_limit == 0` -- reachable, since
`SearchMemoriesSpec.top_k` carries no lower bound. `EventMemory._query`
then derives `max_backward_segments = -1 // 3 = -1` and hands the segment
store a negative window, which the SegmentStorePartition contract does not
define: the SQLAlchemy store happens to short-circuit on `<= 0`, the
in-memory store computes an empty slice and drops the seed. Apply the floor
last so the clamp can only ever produce a non-negative window.

Hole: neither end-to-end test could tell the fix from its absence -- both
pass unmodified against the pre-fix `long_term_memory.py`. `FakeEmbedder`
maps text to `[len(text), -len(text)]`, so under cosine every document
scores exactly 1.0 against every query; all seven timeline episodes become
seeds of equal rank, ties keep insertion order (which is chronological),
and `num_episodes_limit=7` returns all seven with or without expansion. The
"expansion adds episodes" assertion compared a limit-2 search against a
limit-7 one, so the limit alone explained the difference.

Embed on a keyword instead: only `tl-3` matches the query, so `tl-4` and
`tl-5` -- which score zero -- can reach the result only through the
expansion. The tests now pin the exact window (`[tl-3, tl-4, tl-5]`,
chronological, each keeping the window's score), the clamp against an
oversized `expand_context`, and the non-negative window above. All three
fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

* test(event-backend): assert the expansion contract, not a ranking

The tests added in the previous commit discriminate, but they pin an
outcome: an exact episode list (`[tl-3, tl-4, tl-5]`) and exact score
values. Both are properties of the fixture's ranking and of the
backward/forward split `expand_context // 3`, neither of which the fix
claims -- change the split or the scoring and the tests fail while the
behaviour under test is still correct.

Restate them as the contract. Each episode now gets its own similarity from
an explicit search rank, with the match's four timeline neighbours ranked
last, so:

- no correct top-k can return those neighbours, and any nonzero window
  around the match reaches at least one of them whatever the split. The
  assertion is "expansion returned a neighbour the search itself would not",
  plus chronological order and the episode limit.
- the clamp is asserted on the call made to the segment store
  (0 <= backward + forward <= limit - 1, over several limit/expand_context
  pairs) rather than on which episodes come back.
- `expand_context == 0` is asserted as "matches only, best score first",
  without naming them.

Exact lists and score values stay in the unit tests, which own the fill
algorithm and the score-retention rule and are meant to track them. All
three expansion tests still fail against the code they cover.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

---------

Co-authored-by: Claude Fable 5 <[email protected]>

* test(event-backend): pin expand_context as a window of segments

Review on #1541: every wiring fixture uses PassthroughSegmenter, the one
configuration in which segments and episodes cannot come apart, so the
window-to-episode fold was exercised through search_scored only where it
is trivially one-to-one; and expand_context is spent in segments
(EventMemory splits it into the segment store's backward/forward window)
while the API describes it in episodes.

The unit stays segments. EventMemory and the segment store know segments,
not episodes; windowing by true episodes would need new store operations,
and the passthrough segmenter already gives the declarative backend's
behavior, one segment per episode. A splitting segmenter trades that for
chunked retrieval: the same window then reaches fewer episodes, the ones
its segments belong to. The clamp comment says so, since the clamp bounds
episodes as well as segments (a segment belongs to one episode).

The new wiring test runs the timeline fixture under both segmenters at
expand_context=3, limit 4, with TextSegmenter(max_chunk_length=9) splitting
every episode into three segments. Episodes keep the score of the first
window that contributed them, so the episodes at the match's score are
the match's own window: three neighbours under passthrough, one under
the splitting segmenter for any backward/forward split, and the fold
dedups one episode's segments into it. It fails against the pre-fix
long_term_memory.py (no neighbours at all).

_make_ltm_with_metric becomes _make_ltm(embedder, episodes, *, segmenter)
so the test can build self-contained stores per segmenter; the three
metric tests pass their FakeEmbedder explicitly. Mechanical.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

---------

Co-authored-by: Claude Fable 5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant