Skip to content

Overhaul segment store: shared tables with incarnation-scoped tenant keys (fixes #1544, #1546, #1549) (speedkick) - #1548

Merged
edwinyyyu merged 116 commits into
MemMachine:speedkickfrom
edwinyyyu:fix/segment-store-partition-delete-speedkick
Sep 2, 2026
Merged

edwinyyyu merged 116 commits into
MemMachine:speedkickfrom
edwinyyyu:fix/segment-store-partition-delete-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Copy of #1545 for speedkick — cherry-picked from fix/segment-store-partition-delete at c38d5433, all 114 commits, then merged with speedkick so this PR is not left conflicting. Every added and removed line of the ported change is identical to #1545's; the extra commits on top are the integration with speedkick, listed at the end. #1545 carries the closing references for the issues named below.

#1545 is still being worked on; if it gains further commits they are not in this copy yet.

Purpose of the change

Addresses #1544, #1546, and #1549 by construction: the segment store's per-tenant partitioned layout is replaced with shared tables and an incarnation-carrying tenant registry, on every dialect. Design record: design/segment_store_shared_tables.md.

Also fixes #1557, #1558, and #1559: #1462's UTC-normalization fixes are folded in for every SQL store (segment, episode, cluster), so no store ships with the SQLite timestamp corruption.

Why an overhaul instead of the earlier targeted fixes

This PR began as three targeted repairs to the partitioned layout (detach-before-drop for #1544, a management lock for the create/delete deadlock, child-table DML for #1546), each verified on its own. Further investigation showed the layout itself was the root cause and could not be fully repaired:

  • PostgreSQL's FK integrity triggers still addressed the partitioned parents internally, so writes and deletes kept a per-backend lock spike proportional to partition count even after every client statement was fixed (10 -> 166 -> 10 relation locks across the trigger's generic-plan attempt).
  • Writer-vs-lifecycle-DDL deadlocks through the shared parents survived every lock discipline: 41-83 deadlocks per 20-second churn on upstream and on every intermediate fix, in two cycle shapes captured from the server log, one of them a lock-upgrade cycle inherent to dropping a child that carries a cloned FK.
  • Scaling requirements ruled out every per-tenant-table variant: tenant creation must stay cheap at 1e5-1e6 tenants (a per-tenant table costs ~5 catalog relations plus disk files), with 1e4-1e7 rows per tenant.

Description

  • Shared tables everywhere. The ORM models are the physical schema on every dialect. PostgreSQL partitioning, per-tenant DDL, detach machinery, the store-wide management lock, and the per-partition entity cache are removed. With no partitions there is no generic-plan lock fan-out, and with no lifecycle DDL there is nothing for the deadlock class to live in: the churn smoke that measured 41-83 deadlocks/20s on every partitioned build measures zero, with ~60x more write throughput and ~200x more lifecycle throughput.
  • Incarnation-scoped tenant identity. segment_store_pt is the tenant registry: logical key (primary key) plus an incarnation, a random unique-constrained UUID. Data rows carry the incarnation alone, so a data query cannot be built without resolving the registry, and referencing the wrong tenant is structurally impossible. A deleted-and-recreated tenant never sees its predecessor's rows, even mid-purge. A colliding mint is rejected rather than left to probability: the unique constraint rejects a collision with a live incarnation, and the mint transaction re-checks the purge queue after its insert so an incarnation whose garbage is still awaiting purge is re-minted (race-free with the existing tables: the check runs after the registry insert, and no new queue entry for the minted value can appear before commit, since the only registry row carrying it is uncommitted). Any integrity rejection with no committed row under the key is likewise retried with a fresh incarnation, up to _MAX_MINT_ATTEMPTS; a persistent cause surfaces as SegmentStoreAttemptsExhaustedError with the database error chained. Rejections are warning-logged at the detection site, since a genuine collision is astronomically unlikely and the log is the signal for broken randomness or a persistent database error even when the re-mint heals it.
  • O(1) deletion plus a purge queue. create_partition is a row insert (microseconds, no DDL). delete_partition takes FOR UPDATE on the registry row (waiting out writers' FOR SHARE pins), enqueues the incarnation on segment_store_gc (logical key kept for forensics), and deletes the row: O(1) at any tenant size, the pool-model contract of "unreachable immediately, erased asynchronously".
  • Purge as an ABC capability. purge_deleted_partitions() -> bool is part of the SegmentStore interface. The caller's whole protocol is "call until False"; sizing is not a call argument, because callers cannot know engine-appropriate bounds. Deployments set them once at construction: purge_max_segments (default 10000) and purge_max_partitions (default 100, its own bound because retiring an empty entry costs ~0.95 ms, ~200x a segment row, and empty partitions are cheap to mass-create-and-delete). Each call is one transaction (measured ~147k-312k rows/s), committing its progress or nothing. Links follow by ON DELETE CASCADE, 50-68% faster than manual link deletion at one and four links per segment; cascade deletion saturates ~3M link rows/s (380k segments/s at 1 link/segment, 46k at 64), so heavily linked partitions purge in sub-second calls. Link fan-out is set by the deriver, which the store cannot reject after derivation. Before retiring an entry the call also reclaims any link rows that escaped referential integrity, in batches drawn from the same budget and logged as a warning; a link row deletes ~3x cheaper than a segment row (1.0 vs 3.3 us/row, batched), so the shared budget bounds the call without assuming a ratio. Entries are claimed one at a time, oldest first (enqueue stamped by the database clock, indexed), with FOR UPDATE SKIP LOCKED, so a call neither materializes nor locks the rest of the backlog and concurrent purgers from any process share the queue; only the claiming call touches a dead incarnation's rows, so reclamation is deadlock-free by construction. On SQLite, which drops locking clauses, purgers serialize on the write lock at the DELETE: two may claim the same entry, and the second finds no rows and re-retires it. The store never schedules purging; implementations whose deletes reclaim physically return False. The delete path does not purge: an inline drain was tried, first of the global queue and then scoped to the deleted key, and removed, since prompt physical erasure is not a promise the store can keep on every dialect (on SQLite any writer past the busy timeout fails) and the sweeper reclaims within its interval; a deployment must run the sweeper, and the ABC says so.
  • Fencing (Segment store partition handles are not fenced to a partition incarnation: stale handles silently operate on a recreated partition #1549). Writes pin the registry row under FOR SHARE with an incarnation predicate. Reads add the same predicate to their data statement as an EXISTS, so one statement checks liveness and reads; a read that finds no rows issues the registry check on its own, to tell an empty partition from a stale handle, and a windowed read ends with one registry read because its statements take separate snapshots. On SQLite the driver defers BEGIN until the first write, so a SELECT-only fence checks nothing; BEGIN IMMEDIATE would mean taking over transaction management of the caller's shared engine, so the fence is a self-checking registry-row UPDATE (same write lock, scoped to the transaction, match count as the staleness check), and deletion opens its transaction the same way. A handle that outlives its partition or its incarnation raises SegmentStorePartitionHandleStaleError on every dialect; SQLite previously had no stale-handle detection at all. Measured cost: none on reads that find rows, one registry round trip on reads that find none.
  • Foreign keys. The segment table's FK to the registry is removed on purpose (registry and data rows decouple so deletion is O(1)); the link table's FK to the segment table and its cascade remain, enforced by a single-row indexed check with no partition fan-out. On SQLite, enforcement is per-connection state registered at ENGINE creation (enable_sqlite_foreign_keys, applied by the DatabaseManager to its engines): a listener added later misses connections the shared engine already pooled, a pre-existing bug surfaced in review round 4. The store states the requirement on its engine parameter and does not verify it, which is the reference practice for SQLite foreign keys; an engine without the pragma leaves link rows that outlive their segments until the partition is dropped, when the purge reclaims them with a warning. The SQLite vector stores carry the same latent pattern; filed as SQLite per-connection state (foreign keys, sqlite-vec extension) is registered per-store on caller-supplied engines #1568.
  • Folded in: Fix timezone-aware datetime roundtrip on SQLite #1462's SQL-store fixes (fixes [Bug]: Segment store corrupts non-UTC timezone-aware timestamps on SQLite (write path + filter bounds) #1557, [Bug]: Episode store corrupts non-UTC timezone-aware created_at on SQLite (write path + time-range bounds) #1558, [Bug]: Cluster store corrupts non-UTC timezone-aware timestamps on SQLite (last_ts, pending created_at) #1559). App-supplied timezone-aware datetimes are normalized to UTC before persisting in every SQL store: segment timestamps, episode created_at and its start_time/end_time bounds, cluster last_ts and pending created_at. Without this, SQLite (whose DateTime(timezone=True) discards tzinfo) read non-UTC values back shifted by their offset. Datetime filter values are normalized to UTC-aware instants at Comparison/In construction, the filter language's contract stated on FilterExpr (a value denotes an instant; naive means UTC), so every consumer receives instants and compilers only choose a representation; the SQL column leaf binds values as-is, and remaining per-backend normalizations are idempotent defenses, removable separately. Two consumers change by design: the Neo4j compiler's datetime.timestamp() read a naive value in the server's local zone and now receives instants, with its ISO-string branch parsing under the same rule (pinned under a non-UTC process zone); the in-memory short-term evaluator tags a naive stored metadata datetime as UTC at comparison time, since stored user data is not rewritten. The episode store's bounds spell the convention inline as two explicit steps, ensure_tz_aware(...).astimezone(UTC), kept unwrapped so the naive-means-UTC decision stays visible at each site. PostgreSQL is unaffected. On SQLite, rows the episode and cluster stores previously wrote from non-UTC-offset datetimes hold the local wall clock as if it were UTC and read back shifted before and after this change alike; the offset was never stored, so only new writes are correct, and a range query spanning the upgrade mixes the two conventions.

Measurements

Layout comparison and scaled runs are in the design doc. Headlines (pgvector:pg16, one harness for all arms): tenant creation 0.006 ms vs 6-12 ms of per-tenant DDL; ingest and read parity or better at 40x2k and 1M-pair scales including a 500k-row tenant; O(1) delete at any size; churn 0 deadlocks vs 41-83.

Store-API ABAB against upstream main (3 interleaved rounds, one PG instance, medians, at the revision before the read fence was folded into the data statement): ingest +42% (11.1k vs 7.8k pairs/s), delete_segments +20%, event/derivative lookups 10-17% faster, windowed context expansion 8% faster (5.9 vs 6.4 ms), lifecycle create+open+delete 4.4x faster (2.9 vs 12.8 ms/cycle); seed context reads +0.19 ms (1.30 vs 1.11 ms) from the fence's separate registry round trip, since removed. Read-path ABAB of that fold (5 interleaved rounds, medians): seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms, derivative lookups 1.05 vs 1.33 ms, windowed context expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median -0.91 ms); reads that find nothing unchanged. Server-side (EXPLAIN ANALYZE, best of 30) the EXISTS conjunct plans as a one-time InitPlan costing ~3 us per statement, so the client-side gain is the round trip. Purge drains ~216k rows/s; the batched-delete pattern was benchmarked on a 100k-row incarnation with the cascade in place: uuid IN (SELECT ... LIMIT n) 252k rows/s vs a PostgreSQL-only ctid batch at 238k and an unbatched single-DELETE ceiling at 295k, so the portable pattern is also the fastest batched one. Claim and mint-check statements measure 215/156 us.

Testing

  • The behavioral suite runs unchanged against the new layout on both dialects: segment store file 82 in the -m "not integration" lane, 86 in the -m integration lane; full server suite 1926 passed, 3 skipped. Concurrency coverage runs on both dialects wherever the property exists on both: lifecycle churn, racing deletions (single-enqueue assertion), overlapping segment deletes, the mint-vs-deletion collision race, and O(1) deletion via recorded SQL.
  • Purge: test_purge_reclaims_oldest_garbage_first (FIFO, explicit stamps), test_purge_queue_stamps_enqueue_time_from_database_clock, test_purge_bounds_entries_processed_per_call, test_purge_batches_links_that_escaped_integrity (orphans staged through a second SQLite engine without the pragma; fails against the unbounded form), test_purge_skips_entries_claimed_by_concurrent_purger (fails with SKIP LOCKED ablated), test_concurrent_purges_reclaim_everything (both dialects), test_purge_claims_queue_entries_incrementally (fails against the claim-all form), test_purge_reclaims_only_dead_incarnations, test_purge_bound_comes_from_params, test_default_purge_bound_comes_from_params.
  • Incarnations: test_incarnation_with_garbage_left_is_never_reused (queue re-check) and test_incarnation_colliding_with_live_partition_is_never_reused (unique constraint) force a collision on both creation paths and both dialects, each failing with its guard ablated; test_mint_detects_collision_with_concurrent_deletion pins the insert-then-check ordering with a staged concurrent deletion (swapping the statements fails only this test); test_persistent_integrity_error_surfaces_with_cause pins the bounded re-mint with the driver error chained, both dialects; test_persistent_mint_failure_raises_instead_of_looping pins the attempt cap on both creation paths.
  • Fencing: test_stale_handle_raises_after_delete / ..._after_recreate (both dialects), test_reads_check_liveness_inside_the_data_statement, test_windowed_read_raises_when_partition_dies_between_statements (both dialects; fails without the closing registry read), test_recreated_partition_is_isolated_from_old_rows, test_delete_partition_touches_only_registry_and_queue (recorded SQL), test_write_pin_blocks_partition_delete, the SQLite fence test (a write racing delete-plus-purge can no longer orphan rows; the two SQLite race tests carry started-events so a loaded box cannot pass them vacuously).
  • Lock-necessity tests (integration lane, under a second): each locking property is pinned by staging the interleaving it serializes (blocked-ness observed via pg_stat_activity, no grace sleeps). Verified by per-lock ablation on both generations: removing a lock from the new store fails exactly its targeted tests, and the pre-overhaul store fails the churn and concurrent-delete tests with all its locks intact (reproducible DeadlockDetectedError in the two cycle shapes on Episodic search reads lock every segment-store partition once prepared statements go generic (Postgres lock-table exhaustion, 500s under load) #1546). The churn test also caught the _open_or_create_partition race, now a retry loop bounded by _MAX_MINT_ATTEMPTS.
  • Timestamps: Fix timezone-aware datetime roundtrip on SQLite #1462's regression tests for the episode and cluster stores are ported; test_timestamp_roundtrips_with_timezone and test_timestamp_filter_compares_instants_not_wall_clocks ([Bug]: Segment store corrupts non-UTC timezone-aware timestamps on SQLite (write path + filter bounds) #1557) are parametrized over non-UTC zones on both dialects; test_older_than_compares_instants_not_wall_clocks covers the semantic history bound; the filter-parser tests pin construction-time normalization; test_iso_string_coerces_as_utc_instant pins the Neo4j string branch under a non-UTC process zone (skipped where time.tzset is unavailable); test_datetime_metadata_filters_tag_naive_stored_values pins the in-memory evaluator. Each fails without its fix.
  • Erasure and scheduling: LongTermMemory deletes the collection and the partition and returns; the partition is unreachable at once and its rows are reclaimed by the sweeper within its interval. Its handles are nulled so later use raises (test_event_backend_unusable_after_drop_session_partition), and the wiring test asserts the delete path never purges. The resource manager runs one background purge task per store (bounded calls a short pause apart while the store reports a backlog, one idle call per interval otherwise; failures logged and retried; cancelled on close); the loop is a module-level coroutine so a pending task does not pin the manager (test_purge_task_does_not_pin_the_manager), and get_segment_store after close() raises ResourceManagerClosedError. Both purge bounds are validated at construction; the server uses the defaults, and config plumbing is future work.
  • Other review-round fixes, each with a test that fails pre-fix: unloadable codec config commits no registry row (both creation paths); trailing-newline partition keys rejected (re.fullmatch), and the session-id-to-partition-key mapping validates with the store's own validator instead of a drifted copy; the StaticPool guard raises ValueError (test_static_pool_engine_is_rejected); pre-3.35 SQLite runtimes are rejected at construction (test_old_sqlite_runtime_is_rejected); empty-input add_segments short-circuits like delete_segments. Stale-handle handling above the store (API status mapping, cross-replica cache eviction) is filed as Stale segment-store handles need cache eviction and an API status mapping #1571.
  • Retained: test_delete_partition_keeps_other_partitions_cascading and test_delete_partition_keeps_foreign_key_enforced now hold trivially, since no DDL can damage the constraint.
  • Verified on this branch, which already contains speedkick@3b0d9ec6: segment store file 86 in the -m "not integration" lane and 86 in the -m integration lane; full server suite 1925 passed / 16 skipped / 1507 deselected, against 1868 / 16 / 1463 for speedkick alone. ruff format --check over packages/server and ruff check packages/server clean. (The 82 / 86 / 1926-3 figures above are Overhaul segment store: shared tables with incarnation-scoped tenant keys (fixes #1544, #1546, #1549) #1545's own run on main; the counts differ because the two environments have different optional extras installed.) Integrating with speedkick needed three resolutions, in the last two commits: _sql_engine_kwargs (speedkick, fix(db): bound asyncpg statement and connect time so a dead socket recovers #1552) kept as the engine-keyword source with this branch's enable_sqlite_foreign_keys applied to the constructed engine; import logging and import math both kept in test_event_backend_wiring.py; and speedkick's test_get_segment_store_supplies_a_factory, which builds a ResourceManagerImpl via __new__, now also sets the _closed and _segment_store_purge_tasks attributes this branch has get_segment_store read.

Compatibility

No migration from the partitioned layout is provided: the event backend is opt-in and pre-GA, and existing databases recreate their schema (see the design doc). The partition-key contract ([a-z0-9_], max 32 bytes) is unchanged.

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • Breaking change (schema layout of the opt-in event backend)

🤖 Generated with Claude Code

https://claude.ai/code/session_01PVtg6Zea292Pb9L7GXnTJp

https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5

edwinyyyu and others added 5 commits August 28, 2026 16:28
…artition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD
Fixes MemMachine#1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>
The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD
Regression test for MemMachine#1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>
edwinyyyu and others added 13 commits August 28, 2026 17:07
This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>
The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>
…le waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD
…helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>
…tract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>
Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>
Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>
Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>
functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>
@edwinyyyu
edwinyyyu marked this pull request as draft August 31, 2026 18:41
edwinyyyu and others added 10 commits August 31, 2026 20:49
- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on MemMachine#1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>
inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>
Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>
Implements design/segment_store_shared_tables.md. Fixes MemMachine#1544, MemMachine#1546,
and MemMachine#1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>
Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>
The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>
The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>
new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>
Segment-store slice of MemMachine#1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in MemMachine#1462.

Co-Authored-By: Claude Fable 5 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 17, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 17, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…

@edwinyyyu edwinyyyu left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One defect found while auditing the main port, #1661 (line comment).


# Helpers
@override
@override

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Duplicate @override (this line and the next). Same on the port, #1661.

edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
The decorator was stacked twice in the speedkick merge (MemMachine#1548); one
application is the whole effect.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
The decorator was stacked twice in the speedkick merge (MemMachine#1548); one
application is the whole effect.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
The decorator was stacked twice in the speedkick merge (MemMachine#1548); one
application is the whole effect.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
The decorator was stacked twice in the speedkick merge (MemMachine#1548); one
application is the whole effect.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
The decorator was stacked twice in the speedkick merge (MemMachine#1548); one
application is the whole effect.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
The decorator was stacked twice in the speedkick merge (MemMachine#1548); one
application is the whole effect.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 24, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 24, 2026
The decorator was stacked twice in the speedkick merge (MemMachine#1548); one
application is the whole effect.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 25, 2026
…keys (fixes MemMachine#1544, MemMachine#1546, MemMachine#1549) (speedkick) (MemMachine#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the store
existed never receive the pragma, so cascade deletes silently leave
orphaned link rows for LIV…
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 25, 2026
The decorator was stacked twice in the speedkick merge (MemMachine#1548); one
application is the whole effect.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit that referenced this pull request Sep 25, 2026
…keys (port of #1548) (#1661)

* Overhaul segment store: shared tables with incarnation-scoped tenant keys (fixes #1544, #1546, #1549) (speedkick) (#1548)

* Fix: Detach segment store partitions before dropping them, and lock partition metadata on delete

Deleting a partition on PostgreSQL dropped its child tables with CASCADE.
The foreign key from segment_store_dv_ln to segment_store_sg is declared on
the partitioned parents, so the CASCADE dropped the parent-level constraint
rather than only the part belonging to the deleted partition. After the
first partition deletion the store stopped enforcing the link for every
remaining partition, and ON DELETE CASCADE stopped removing derivative
links with it, so delete_segments left orphaned rows that
get_derivative_uuids_by_segment_uuids still returned. Detaching each child
before dropping it keeps the constraint and the cascade intact.

delete_partition also took only a row lock on the partition row. ROW SHARE
does not conflict with the SHARE ROW EXCLUSIVE table lock the create paths
take, so a concurrent create and delete could reach segment_store_sg and
segment_store_dv_ln in opposite order and deadlock; that reproduced in 4 of
7 sampled interleavings against PostgreSQL 16, and in 0 of 7 once delete
takes the same table lock first. The row lock stays, because it is what
makes delete wait for in-flight writers holding FOR SHARE on that row.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: address partition child tables directly for segment DML

Fixes #1546. Parent-table queries carry the partition key as a bind
parameter; once asyncpg's prepared statement flips to a cached generic
plan (after five executions) PostgreSQL locks every child partition on
every execution before runtime pruning. With hundreds of partitions and
concurrent sessions this exhausts the lock table (searches fail 500
'out of shared memory') and saturates the database CPU with lock churn
and generic-plan startup.

The partition handle now maps the ORM entities onto its own child
tables (orm.aliased with adapt_on_names) and targets them for
insert/delete, so every plan references exactly one partition. SQLite
keeps the parent tables (it has no children).

Measured on a store with 316 partitions: max relation locks held by a
backend during a read loop drops 1276 -> 4; the 239 HTTP 500s in a
4-worker load test disappear; throughput at 128 concurrent requests
rises ~20-30% with PostgreSQL no longer pinned at its CPU cap.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Test: Pin the partition-delete lock ordering

The foreign-key half of this branch has a regression test; the lock
ordering did not. A deadlock test would be timing-dependent, so assert the
invariant the deadlock analysis rests on instead: delete_partition issues
LOCK TABLE segment_store_pt IN SHARE ROW EXCLUSIVE MODE, and issues it
before any DETACH or DROP of a child table. Without the lock the test fails
and reports the statement sequence it saw.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* Style: Apply ruff format to the lock-ordering test

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* test: segment DML must address only the partition's child tables

Regression test for #1546. Captures the SQL the partition handle emits
across add / seed read / windowed read / filtered read / uuid maps /
delete and asserts no statement references the partitioned parents --
the deterministic observable of the generic-plan lock explosion (lock
counts would need timing-dependent pg_locks sampling). Fails against
the parent-table implementation, passes with per-partition DML.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop annotation-only AliasedClass and Table imports

This SQLAlchemy version exports no public AliasedClass name (only the
aliased() factory), so the attribute annotations forced an import from
sqlalchemy.orm.util. The annotations were documentation only; the
branch comment already records that the attributes hold either the ORM
class or its child-table alias.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: derive partition table names from one helper and the models

The parent names come from the models' __tablename__ and the child
naming pattern lives in _pg_child_table_name, used by child-table
creation, teardown, and the per-partition DML targets, so the three
sites cannot drift apart. Physical names are unchanged; the regression
tests keep literal names to pin the on-disk naming contract.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Fix: Drop detached children, and stop holding the store-wide lock while waiting for writers

Three follow-ups from review of this branch.

The child-table probe asked whether the table exists, but the statement it
guards is DETACH PARTITION, which requires that the table be attached. A
child left detached by manual maintenance or an interrupted DETACH
PARTITION CONCURRENTLY (that form is not transactional) passed the probe and
made DETACH raise "is not a partition of", rolling back the transaction, so
the partition became neither deletable nor recreatable -- a state the
CASCADE drop this branch replaced used to clean up. Probe pg_inherits for
attachment instead, in one round trip for both children, and drop an
unattached child directly.

delete_partition took the partitions-table lock before the row lock that
waits for in-flight writers, so a slow writer on one partition stalled
open_or_create_partition for every partition, which runs on the request
path. Take the row lock first; the table lock only has to be held across the
child DDL for the deadlock argument to hold. Re-measured: 4/7 sampled
interleavings deadlock with no table lock, 0/7 with either ordering.

Tests: the foreign key is now asserted to be enforced after a partition
delete, not only that the cascade fires -- the PR's measurements list those
as separate things the CASCADE drop broke. A second test leaves a child
detached and requires deletion to succeed and the key to be reusable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01MpFwBnZ6SxdMe3pSuMTCHD

* fix: memoize partition entities per key; validate keys in the naming helper

Review follow-ups on the child-table change:

- The child Table objects and aliases are now built once per partition
  key (functools.cache) instead of per handle. SQLAlchemy's compiled-
  statement cache keys on the Table objects a statement references, so
  per-handle tables made every handle's statements recompile and
  polluted the cache for everything else (verified: cache keys differed
  across handles for the same partition; now identical).
- The tables carry columns only. The to_metadata copies dragged along
  foreign keys with unresolvable targets and duplicate index names --
  latent hazards for anything walking that MetaData.
- _pg_child_table_name validates the partition key itself, so every SQL
  string built from a child table name (including the DDL literals) is
  safe by construction rather than by call-site convention; the store's
  validator moved to module level beside it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make partition key validation part of the segment store contract

Partition keys are embedded in native storage identifiers by any
implementation, so the alphabet/length rule is interface-level, not a
SQLAlchemy detail: validate_partition_key now lives in the package's
data_types (exported from the package), the SegmentStore.create_partition
docstring states the contract, and the SQLAlchemy store imports it.
Deliberately NOT unified with the vector store's identical identifier
rule: the repo-wide naming contract is not wired through yet, so the
convergence is treated as incidental for now.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: move partition key validation to segment_store/utils.py

Mirrors the vector store's layout (validate_identifier in
vector_store/utils.py); the interface docstring states the key rule
plainly instead of referencing a code path.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state partition key naming constraints in the VectorStore format

Same ABC-level 'Naming constraints:' block the vector store uses, no
method-level restatement, and the length limit is enforced and
documented in bytes, matching validate_identifier.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: state the key rule as the regex, not a prose fragment

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: restore the ABC's original naming-constraints docstring

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: make validate_partition_key boolean, call sites raise

Mirrors the vector store's validate_identifier: the predicate returns
bool so callers can compose it, and each entry point raises its own
error in the vector store's message style.

Co-Authored-By: Claude Fable 5 <[email protected]>

* Revert "refactor: make validate_partition_key boolean, call sites raise"

This reverts commit 88be86603c36e4b759da0f876968e6dd6db8d913.

* fix: bound the partition-entity cache with lru_cache(4096)

functools.cache grew ~29 KiB per distinct partition key (measured) for
the life of the process, including deleted partitions. The LRU cap
bounds it at ~115 MiB per worker; eviction is harmless since a rebuilt
entry is identical and only costs recompiling that partition's
statements once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* style: drop point-in-time memory figures from the cache comment

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: address second review round

- Type the partition entities: _pg_partition_entities returns a
  NamedTuple with real field types instead of a positional tuple of
  object, which was adding 55 ty diagnostics (the CI static check would
  have failed) and forcing blind unpacking; callers and the memoization
  test use named fields, and the private _generate_cache_key assertion
  (redundant given object identity) is dropped.
- DROP TABLE gains IF EXISTS back, so an out-of-band drop landing
  between the state probe and the drop cannot leave a partition
  half-deleted.
- open_or_create_partition opens existing partitions without the
  store-wide management lock (double-checked: unlocked read, then lock
  and re-check only when creating), so request-path opens no longer
  serialize behind a concurrent deletion's DDL window; pinned by
  test_open_existing_partition_takes_no_management_lock.
- The engine's compiled-statement cache is raised from the default 500
  (per-partition statements would thrash it once enough partitions are
  live concurrently).
- Comments and the DML test docstring scope the lock claim honestly:
  PostgreSQL's FK integrity triggers still address the parents
  internally, costing a one-shot per-backend lock spike on
  writes/deletes when a trigger plan first goes generic (verified:
  ~10 locks steady, one spike at execution six, then back) -- tracked
  on #1546.
- The detach test cleans up its detached child unconditionally so a
  failure cannot poison the session-scoped container for later tests;
  the statement recorder is a shared fixture instead of copy-paste;
  the byte-length check encodes once.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop the typing casts from the partition entities

inspect(Model).columns yields real Column objects (the stubs type
__table__ as FromClause, which forced the cast), and SQLAlchemy's
typing convention represents an aliased entity as the mapped class
type, so the NamedTuple fields are type[SegmentRow] /
type[DerivativeLinkRow] and aliased() assigns without coercion.

Co-Authored-By: Claude Fable 5 <[email protected]>

* design: shared segment store tables with incarnation-scoped keys

Design record for replacing the per-tenant partitioned layout with
shared tables on every dialect: the tenant registry carries an
incarnation, data rows are keyed by <logical_key>@<incarnation>,
deletion is an O(1) registry write plus a purge queue, and fencing
fails stale handles loudly. Records the measured comparison against
PARTITION OF and standalone-table layouts and the scaling requirements
(cheap tenant creation at 1e5-1e6 tenants, 1e4-1e7 rows per tenant)
that decided it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: shared segment store tables, incarnation fencing, O(1) delete

Implements design/segment_store_shared_tables.md. Fixes #1544, #1546,
and #1549 by construction:

- The ORM models are the physical schema on every dialect; PostgreSQL
  partitioning, per-tenant DDL, the detach machinery, the store-wide
  management lock, and the per-partition entity cache are all removed.
  No partitions means no generic-plan lock fan-out (client or
  RI-trigger) and no DDL for lifecycle deadlocks to live in -- the
  churn smoke that measured 41-83 deadlocks per 20s on every
  partitioned build measures zero, with 60x more write throughput.
- segment_store_pt becomes the tenant registry: partition_key +
  incarnation. Data rows are keyed by <logical_key>@<incarnation>, so
  a deleted-and-recreated tenant never sees its predecessor's rows.
- create_partition is a row insert (no DDL); delete_partition is O(1):
  FOR UPDATE on the registry row (drains writer pins), enqueue the
  physical key on segment_store_gc, delete the row.
  purge_deleted_partitions reclaims rows in chunked background batches.
- Writes pin the registry row FOR SHARE with an incarnation predicate;
  reads check it too: a stale handle raises
  SegmentStorePartitionStaleError on every dialect, SQLite included.
  Measured cost: one extra registry round trip per read operation.
- The segment table's FK to the registry is removed (registry and data
  rows are deliberately decoupled for O(1) deletion); the link-table FK
  and cascade remain.

Same-moment ABAB vs the partitioned build: ingest and windowed reads at
parity, lifecycle cycles 4-6x faster, tenant creation ~1000x cheaper
(row insert vs DDL).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: identify tenant data rows by incarnation UUID alone

Data rows drop the composite <logical_key>@<incarnation> string for a
bare incarnation UUID column: a data query cannot be constructed
without resolving the registry, so referencing the wrong tenant is
structurally impossible; index entries narrow from a 41-byte varchar
to a native 16-byte uuid; random UUIDs are globally unique across
nodes without coordination, so tenant moves between databases carry
rows verbatim; and collisions among incarnations with live traces are
rejected by constraints (unique on the registry, primary key on the
purge queue) instead of left to probability. The purge queue keeps the
logical key for forensics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: finish the physical-key -> incarnation wording in the design doc

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: fence by incarnation alone

The incarnation is unique-constrained, so it resolves the registry row
by itself; the logical-key predicate was a leftover from the composite
string design and contradicted the rule that the incarnation is the
handle's sole authority.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename the stale error to name the handle, not the partition

The handle is what is stale -- the partition is deleted -- and
SegmentStorePartitionHandleStaleError follows the existing noun+state
convention (ConfigMismatch, AlreadyExists).

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: inline uuid4 for incarnation generation

new_incarnation() was a one-line wrapper adding indirection for no
behavior; the multi-node rationale lives in the design doc.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize segment timestamps to UTC before persisting

Segment-store slice of #1462: SQLite's DateTime(timezone=True) discards
tzinfo and stores wall-clock fields verbatim, so a non-UTC timezone-aware
timestamp written without UTC normalization read back shifted by its
offset (13:30:45-08:00 came back as 05:30:45-08:00). The read path
already assumed UTC and reapplies the separately stored offset; only the
write was missing the conversion. PostgreSQL timestamptz stores a true
instant, so this is a no-op there.

Regression test parametrized over UTC/-08:00/+05:30 runs on both
backends; verified the non-UTC params fail without the fix and pass
with it (sqlite 54, pg 58).

The companion filter-bound normalization lives in shared
sql_filter_util.py (used by episode and cluster stores too) and stays
in #1462.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: normalize datetime filter bounds to UTC in SQL filter compilation

Second half of the #1462 segment-store slice: timestamp columns now hold
the UTC instant, so comparison bounds must be named in the same frame.
On SQLite an aware datetime bind is rendered as wall-clock digits with
tzinfo dropped and compared lexically, so `timestamp <=
2024-01-01T08:00+08:00` excluded a row stored at 00:00Z -- the same
instant. _normalize_column_value converts datetime values (Comparison
and In leaves) to UTC before binding; PostgreSQL compares timestamptz
by instant either way, so the two backends now agree.

The helper lives in the shared sql_filter_util because that is where
column leaves are compiled; other stores' write paths (episode, cluster)
are intentionally not touched here.

Regression test parametrized over the same instant named in +00:00,
+08:00, and -08:00, on both backends; verified the non-UTC bounds fail
without the fix and pass with it (sqlite 57, pg 61; full server suite
1880 passed, 3 skipped).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: promote purge_deleted_partitions to the SegmentStore ABC

Physical reclamation of deleted partitions is now an ABC capability:
callers schedule it however they want; the store never schedules it
itself. delete_partition's contract notes that reclamation may be
deferred. Implementations whose deletes reclaim physically implement
it as a no-op returning False.

Signature review against the prior purge iterations (#1199/#1205):
- The old three-step orphan-derivative API (get_orphaned / mark /
  purge) existed only because derivative purging interleaved with
  vector-collection deletes between steps; incarnation purge is fully
  internal to the store, so a single method suffices.
- The old scheduling knob (purge_interval loop in ExtraMemory) lived
  in the consumer -- preserved: no scheduling in the store.
- The bound is max_segments (domain unit; derivative links ride along
  uncounted) rather than max_batches, which presumed chunked-transaction
  implementations. batch_size stays as a keyword on the SQLAlchemy
  implementation only, as a transaction-size tuning knob.
- Returns bool ("reclaimable work may remain") instead of rows deleted:
  a row count cannot distinguish "drained" from "stopped at the bound"
  when a dead incarnation has zero data rows, and the scheduling caller
  needs exactly the more-work signal.

New test pins the bound and the completion signal on both dialects.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: drop batch_size from purge_deleted_partitions

With max_segments as the caller's bound, a per-call batch_size is
redundant: bounded calls already cap every delete transaction at the
remaining budget, so the knob only governed the unbounded case --
where transaction sizing is engine policy, not caller policy. The
chunk is now an internal constant (_PURGE_CHUNK_SIZE); if a deployment
ever needs to tune it, it belongs in SQLAlchemySegmentStoreParams,
not per call. Tests exercise multi-chunk draining by patching the
constant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge is one atomic slice per call; fix open_or_create race

Purge contract resolved to atomic-slice-per-call: each
purge_deleted_partitions call is a single transaction that reclaims up
to max_segments rows and either commits that progress or nothing.
Draining a backlog is the caller's loop (call until False), so
reclamation never holds a long transaction, committed slices survive
interruption, and there is no internal chunking competing with the
caller's bound. max_segments=None means the store-chosen slice size
(_PURGE_SLICE_SEGMENTS), keeping engine-appropriate transaction sizing
out of callers' hands. Rationale over the alternatives: cross-call
atomicity is anti-useful for gc (a huge atomic purge is exactly the
long-transaction hazard, and an error would forfeit all progress),
while batch_size+max_batches exposes the store's transaction quantum
and bounds a call only as a product of two knobs.

Also fixes a TOCTOU in _open_or_create_partition caught by the new
lifecycle churn test: losing the insert race and then finding no row
(a concurrent delete removed the winner) raised RuntimeError; the
read-then-insert sequence now retries, since every retry implies
another actor changed the state. New deterministic fencing tests:
test_write_landing_during_delete_is_never_orphaned (the write pin means
rows can never land under an incarnation the purge queue no longer
tracks) and test_concurrent_remote_delete_yields_single_queue_entry
(the delete pin means racing deletions enqueue exactly once).

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: lock-necessity suite verified by per-lock ablation in both generations

New test_segment_store_locking.py (PostgreSQL integration lane) pins
each locking property through the public API, using only surface shared
with the pre-overhaul partitioned store so the module runs against both
generations. All interleavings are staged event-driven: blocked-ness is
decided by observing pg_stat_activity lock waits, not elapsed time;
there are no grace sleeps, and paused writers are released in finally
blocks so a failing assertion cannot wedge fixture teardown. The lane
adds under a second of CI time.

Ablation matrix (each lock removed one at a time via source-patched
variant trees, PYTHONPATH-shadowed; old = pre-overhaul partitioned
store at 14b8f0a2~1):

- write pin ablated (either generation): write-pin test fails, plus
  the no-orphaned-writes fencing test on the new store.
- delete pin ablated (new store): churn, concurrent-delete, and
  single-queue-entry tests fail (double-enqueue IntegrityError).
- delete row pin ablated (old store): write-pin test fails (the delete
  no longer waits out the in-flight writer).
- ordered delete_segments row locks ablated (either generation): no
  test fails -- identical DELETE shapes lock rows in identical orders
  on PostgreSQL (sorted scalar-array probes, TID-ordered bitmap scans),
  so the AB/BA cycle needs plan divergence the store never produces.
  The overlap test is kept as a regression canary and documented as
  such; whether to keep the pre-lock itself is a separate decision.
- old store with ALL locks intact: churn and concurrent-delete tests
  fail with DeadlockDetectedError in the two cycle shapes documented
  on #1546 (delete-vs-delete lock upgrade over the table mutex;
  create-vs-delete DDL cycles through the shared parents). Those
  deadlocks are inherent to the partitioned layout -- the property the
  shared-table overhaul removes, and these tests now pin.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries with SKIP LOCKED; correct lock-order rationale

Concurrent purgers were a real deadlock surface: two processes draining
the same dead incarnation delete overlapping row sets through unordered
scans. Claiming queue entries with FOR UPDATE SKIP LOCKED removes the
contention instead of ordering it -- racing purgers partition the queue,
and only the claiming call touches a dead incarnation's rows (writers
cannot; the fence pins live incarnations only), so reclamation is
deadlock-free by construction. This is the claiming half of the purger
scale-out design in the design doc; the ABC now states the contract
(concurrent calls, including cross-process, must neither error nor
deadlock).

Tests: test_purge_skips_entries_claimed_by_concurrent_purger stages a
purger from another process holding its claim uncommitted -- a
concurrent purge must skip the entry and complete without blocking;
verified to fail (blocks on the held queue row) with the claim ablated
and pass with it. test_concurrent_purges_reclaim_everything pins the
correctness property on both dialects: racing drain loops terminate
cleanly with full reclamation.

Also rewords the ordered-row-lock rationale in the locking suite: the
consistent acquisition order that makes the ablation unobservable is
current PostgreSQL executor behavior, not a guarantee any engine
documents -- the pre-lock imposes the order deliberately, and the canary
catches divergence if an engine or plan change ever produces it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: accept any task result type in _wait_until_blocked_or_done

The helper only observes done-ness; Task[None] rejected the purge
task (Task[bool]).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: reject minting an incarnation whose garbage is still awaiting purge

Data rows are keyed by incarnation alone, so a fresh mint colliding with
a dead-but-unpurged incarnation would adopt its garbage and then be
erased by the purger. The registry's unique constraint only guarded
collisions with live incarnations; the purge-queue case was guarded by
uuid randomness alone.

The mint (shared by create_partition and open_or_create_partition) now
re-checks the purge queue inside the insert transaction and re-mints on
collision. The check is race-free with the existing tables -- no ledger
table needed: it runs after the registry insert, so a concurrent
deletion moving a colliding row to the queue (the insert waited on its
uncommitted registry delete) is already visible, and no new queue entry
for the minted value can appear before commit because the only registry
row carrying it is uncommitted. The locking read sees latest-committed
state on dialects whose plain reads serve transaction-start snapshots;
SQLite serializes whole transactions. An incarnation value can therefore
never be reused while any trace of it remains within a database; across
databases, uniqueness still rests on random-uuid collision resistance.

test_incarnation_with_garbage_left_is_never_reused forces the collision
by stubbing the mint (both creation paths, both dialects); verified to
fail with the re-check ablated and pass with it.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: one collision error for live and garbage incarnation mints

Both collision causes look the same to the mint's callers and share the
same remedy -- mint a fresh incarnation and retry -- so they now share
one error, with the cause classification (key taken vs incarnation
collision) resolved inside _insert_partition_row: an IntegrityError with
a committed row under the key means the key is taken
(SegmentStorePartitionAlreadyExistsError: open or delete it instead);
without one, the incarnation collided with a live row. Errors are typed
by the decision the caller makes, not by the failing constraint, and
both call sites shrink to one remedy branch per error.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: fold locking tests into the segment store test file; drop "slice"

All SQLAlchemy segment store tests live in one file. The separate
locking module existed so the same tests could import against the
pre-overhaul partitioned store for the lock-ablation matrix; that
verification is done and recorded, so the split's constraint is spent.
The per-lock coverage map moves to a section comment.

Also replaces the "one slice per call" purge wording, which was
circular (a slice being defined as whatever one call does), with the
actual contract: each call reclaims at most max_segments segments --
in this store, one transaction that commits that progress or nothing --
and None means the store's default bound (_DEFAULT_PURGE_MAX_SEGMENTS,
renamed from _PURGE_SLICE_SEGMENTS).

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: correct design-doc drift; drop dead _is_postgresql flag

Accuracy review of the design doc against the code:
- The create bullet said "one row insert"; the mint transaction also
  re-checks the purge queue.
- The locking model omitted the purger's SKIP LOCKED queue claims and
  the mint's collision-case queue pin; it now lists every row lock and
  why reclamation cannot contend with anything.
- The consequences section claimed the only remaining dialect split is
  the LATERAL-vs-loop read strategy; the PostgreSQL-only ordered row
  locks in delete_segments and SQLite's foreign-key pragma are splits
  too.

_is_postgresql was assigned and never read -- dead since the overhaul
removed the DDL branches.

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: make the default purge bound a store construction parameter

SQLAlchemySegmentStoreParams.default_purge_max_segments (default 10000)
replaces the module constant: each purge call is one transaction, so
the right default bound is dialect- and deployment-dependent, and the
construction parameter lets an application set it once instead of
every purge caller reading configuration to pass max_segments.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: cover live-incarnation collision and purge-bound params wiring

Coverage audit of the recent additions found two unpinned paths:

- The live half of the mint's collision handling (registry unique
  violation classified by the key re-read, then re-mint) had no test --
  only the garbage half did. test_incarnation_colliding_with_live_
  partition_is_never_reused forces the collision on both creation paths
  and both dialects; verified to fail with the classification ablated
  (create_partition misreports AlreadyExists) and pass with it.
- The purge tests patched the store's default-bound attribute directly,
  leaving the SQLAlchemySegmentStoreParams.default_purge_max_segments
  wiring itself untested. test_default_purge_bound_comes_from_params
  constructs a store with a small configured bound and observes it
  govern an unbounded purge call.

Also converts the two override-method docstrings (delete_partition,
purge_deleted_partitions) to body comments: the contract lives on the
ABC; overrides keep only implementation mechanics.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename _logical_partition_key to _partition_key

The "logical" qualifier contrasted with the physical partition key of
the composite-key era; data rows now carry no key at all, so there is
nothing physical to distinguish from.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename PurgeRow to PurgeQueueRow

The model classes are named for what a row represents (PartitionRow,
SegmentRow, DerivativeLinkRow); a segment_store_gc row is not a purge
but an entry of the purge queue.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: restore @staticmethod on _resolve_segment_field

It became an instance method when field resolution went through the
handle's per-partition aliased entities; the shared-table overhaul
resolves against the module-level SegmentRow again, leaving self
unused. Call sites return to the original class-qualified form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the mint's insert-then-check statement ordering

The collision guard relies on checking the purge queue AFTER the
registry insert: under READ COMMITTED, the insert's unique-index wait
on a concurrent deletion's uncommitted registry delete is what forces
that deletion's queue entry to be committed -- and therefore visible to
the later check. Checked before the insert, the queue is read too
early and the mint commits a live partition whose incarnation is on
the purge queue, handing its rows to the purger.

Only a concurrent interleave distinguishes the orderings, so the
sequential collision tests cannot pin it: verified by swapping the two
statements -- the sequential tests all still pass (the opposite order
is correct for non-concurrent use), while the new
test_mint_detects_collision_with_concurrent_deletion fails (it is
incorrect for concurrent use). The test stages the interleave
deterministically: a raw-session deletion held uncommitted, the
colliding mint observed blocking on it via pg_stat_activity, then the
deletion committed.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: plain "maximum number of segment rows purged per call" wording

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: match params docstring to the pydantic field description

Convention in the class: the Attributes entry carries the field
description plus the default; the field expresses the default via its
default attribute only.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge claims queue entries one at a time

The claim SELECT had no limit: it materialized and row-locked every
unclaimed queue entry even when max_segments exhausted on the first
incarnation -- a mass-deletion backlog was fetched wholesale per call,
and the first purger claimed the entire queue, so concurrent purgers
skipped everything and exited instead of sharing the backlog.

Claims are now LIMIT 1 FOR UPDATE SKIP LOCKED, issued as the call
processes entries: a bounded call locks exactly what it works on.
Within the transaction each claimed entry is retired before the next
claim, so the call's own claims (which SKIP LOCKED does not skip)
cannot recur and the loop terminates.

test_purge_claims_queue_entries_incrementally pins the property via
recorded SQL: every queue claim carries LIMIT, and a call whose bound
exhausts on its first incarnation issues exactly one claim; verified
to fail against the previous claim-all form.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: batch, not chunk, for the purge deletion unit

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: purge runs on an engine connection for typed rowcount

AsyncSession.execute has no DML overload -- it is typed Result[Any] for
every non-typed statement, so reading rowcount needed an isinstance
narrowing to CursorResult (whose unreachable else-branch would have
fabricated a zero count). AsyncConnection.execute is typed CursorResult
in every overload, and the purge transaction is pure Core DML with no
session features, so it now runs on self._engine.begin(): the library's
own annotations carry the type and the narrowing disappears.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* fix: address code-review findings (round 3)

Six confirmed-or-verified defects from a second review session, each
fixed with a test verified to fail on the pre-fix code:

- SQLite write fence was a no-op: the driver defers BEGIN to the first
  data-modifying statement, so the fence SELECT ran outside the write
  transaction and a write racing a delete-plus-purge committed rows no
  queue entry tracked. The write fence (and deletion's row check) now
  issue a no-op registry UPDATE first, opening the write transaction so
  the check is transactional and racing deletions serialize.
- The shared column-leaf UTC normalization silently changed OTHER
  stores' datetime filters: their write paths still store wall clock,
  so on SQLite their filters stopped matching rows they had just
  written. Normalization is now an explicit compile_sql_filter opt-in
  (column_datetimes_are_utc) that only the segment store sets; other
  stores regain their previous behavior, and #1462 flips the opt-in
  for the stores whose write paths it fixes.
- Mint collision retries were unbounded: any persistent IntegrityError
  with the key absent became an infinite hot loop. Both creation paths
  cap consecutive collision retries (_MAX_MINT_ATTEMPTS) and re-raise
  the underlying error -- consecutive failures at that depth mean a
  permanent cause, not a race.
- purge_deleted_partitions accepted non-positive bounds and returned
  True unconditionally, spinning the documented drain loop; it now
  raises ValueError. Empty incarnations charge one segment of budget,
  so a backlog of empty tenants is bounded per call instead of drained
  in one unbounded transaction.
- open_or_create committed the registry row before materializing the
  payload codec, leaving an unopenable partition behind on codec
  failure; the codec is loaded before the insert again.
- validate_partition_key used re.match with $, accepting keys with a
  trailing newline; now re.fullmatch.

Also from the review: the ABC documents the stale-handle contract on
SegmentStorePartition and corrects purge's False semantics (work owned
by a concurrent purger is not counted); the blocked-or-done test helper
scopes pg_stat_activity to the current database.

Rejected findings, with grounds recorded in the PR discussion: the
read-fence round trip is the deliberate loud-fencing contract (#1549);
fence/live-check unification, the forensic enqueued_at column, and
FIFO claiming are declined as taste; the partition-key rule's overlap
with service_locator stays per the incidental-convergence ruling.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: LongTermMemory erasure drains the purge queue inline

delete_partition's physical reclamation is deferred by design, but the
review found its one production caller now leaked: session deletion
previously removed data physically (DROP on PostgreSQL, cascade on
SQLite) and nothing anywhere called purge_deleted_partitions. The
erasure path drains the queue inline before returning, restoring
physical removal semantics; background scheduling remains available to
other callers.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: UTC-normalize every SQL store; self-checking SQLite fence; FIFO purge

Four follow-ups to the review round, per direction:

- The datetime-normalization opt-in is gone: instead of scoping the
  shared compiler's UTC bound normalization to the segment store, every
  SQL store's write path is fixed honestly in this PR. #1462's episode
  and cluster fixes (created_at + start/end bounds; last_ts + pending
  created_at) are ported with their regression tests, and the compiler
  normalizes column datetime bounds unconditionally -- correct for all
  consumers, since the semantic-storage columns it also serves are
  server-generated UTC (func.now()). Fixes #1558 and #1559 here.

- The SQLite fence is one self-checking statement instead of a no-op
  UPDATE plus a SELECT: the proper primitive, BEGIN IMMEDIATE, is only
  expressible engine-wide in SQLAlchemy (it would put every read
  transaction behind the write lock), so the registry-row UPDATE
  acquires the same write lock scoped to the transaction, and its match
  count is the staleness check. Deletion opens its transaction the same
  way, with zero matches as the idempotent no-op case.

- The purge queue is FIFO: claims order by enqueued_at (indexed), so
  the oldest garbage is reclaimed first and the name is honest. Queue
  entries carry their own per-call bound
  (SQLAlchemySegmentStoreParams.purge_max_partitions, default 1000)
  instead of charging a fake segment of budget: their cost is round
  trips rather than row deletions, and empty partitions are cheap to
  mass-create-and-delete, normally or adversarially. Empty entries no
  longer consume max_segments.

- PostgreSQL-only concurrency coverage gains SQLite counterparts
  wherever the property exists on both dialects: lifecycle churn,
  racing deletions (plus a single-enqueue assertion), overlapping
  segment deletes now run on both; new SQLite tests pin the
  mint-vs-deletion collision race and O(1) deletion via recorded SQL.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: precise rationale for the SQLite fence primitive

BEGIN IMMEDIATE is expressible per-transaction in principle, but only
atop engine-wide rewiring (isolation_level=None plus a begin-event
hook) that the store cannot apply to a caller-owned, possibly shared
engine; say that instead of "only expressible engine-wide".

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: two-character index name tokens, matching the prior convention

pk_ev / pk_ts_ev_bk_ix / pk_su used two characters per indexed column;
in (incarnation) and ea (enqueued_at) follow, replacing inc and enq.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor!: purge takes no arguments; cap derivatives at ingestion

purge_deleted_partitions() -> bool. Callers cannot know
engine-appropriate transaction sizing -- the same argument that made
the default a construction parameter removes the per-call override:
the caller's whole protocol is "call until False", and every bound
(purge_max_segments, renamed from default_purge_max_segments;
purge_max_partitions) is implementation policy set once at
construction. Non-positive bounds are now impossible by pydantic
validation, superseding the runtime ValueError.

The derivative side is bounded where it is created, not where it is
reclaimed: purge keeps relying on the link-table ON DELETE CASCADE --
benchmarked against manual link deletion on the real schema and 50-68%
faster (1 link/segment: ~312k vs ~209k segs/s; 4 links: ~266k vs
~158k; the manual pattern's extra round trips and array shipping cost
more than the per-row indexed trigger probes) -- and ingestion rejects
more than max_derivatives_per_segment links per segment (default 100),
so one purge call's work is at most purge_max_segments segment rows
plus that many times the cap in link rows.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the store-level derivatives-per-segment cap

The cap rejected at add_segments time, when the caller has already
segmented and derived and can do nothing to obey it -- the bound on
link fan-out is ingestion-pipeline policy (deriver/segmenter design),
not a store contract. Performance also gives the cap no case: measured
across densities, cascade deletion saturates around 3M link rows/s
(380k segments/s at one link per segment, 302k at 4, 175k at 16, 46k
at 64 -- per-row cost FALLS with density, 1.3us/row at 1 link to
0.34us at 64), so a purge_max_segments=10000 call finishes in ~0.45s
even at 64 links per segment. The design doc records where the bound
lives and the measured sensitivity.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: state the purge contract's promise to callers

The bounds are implementation policy; what the caller is promised is
that a purge call does not noticeably degrade concurrent request
serving. The design doc also records why a store-level link cap would
be unactionable (only the deployment's segmenter/deriver choice can
change the ingested shape, so a dedicated error type would have no
useful handler).

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: rebalance the purge entry bound to measured cost

Retiring an empty queue entry measures ~0.95 ms through the store
(four round trips), roughly 200x a segment row at the measured purge
rate -- not the ~10x the old default implied. purge_max_partitions
drops from 1000 (a ~0.95 s transaction when saturated, 20x the row
bound's ~46 ms) to 50, putting a full-entry call and a full-row call
at comparable transaction duration. Backlog drain throughput is
unchanged (~1k entries/s regardless of slicing); only per-call
transaction length shrinks.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: public SegmentStorePermanentError; power-of-ten entry bound

Mint-retry exhaustion raised "last_collision.__cause__ or
last_collision" -- expedient plumbing that leaked either the wrapped
SQLAlchemy IntegrityError or the private collision type to callers.
Per the error-design principle (type by the caller's decision), the
decision here is "retrying will not fix this; diagnose", so both
creation paths now raise the ABC-declared SegmentStorePermanentError
with the underlying error chained as the cause. The ABC documents it
on create_partition and open_or_create_partition.

purge_max_partitions defaults to 100 instead of 50: sibling fields of
one config keep to the same numeric family (powers of ten, alongside
purge_max_segments=10000); a saturated entry call (~95 ms measured)
and a saturated row call (~46 ms) stay within the same order of
transaction duration.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: rename SegmentStorePermanentError to SegmentStoreRetriesExhaustedError

"Permanent" asserted a diagnosis the store cannot make -- sustained
adversarial churn could in principle clear on a later attempt. The
name now states only what happened (internal retries exhausted), with
the guidance phrased as likelihood: an immediate retry is unlikely to
succeed; diagnose the chained cause.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: drop illustrative examples from contract docstrings

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: bound open_or_create's lost-race arm with the same retry cap

The collision arm was capped but the AlreadyExists arm looped
unboundedly -- the reviewer's livelock finding. Both non-terminating
outcomes now count toward one retry budget, and exhausting it raises
SegmentStoreRetriesExhaustedError with the last error chained. With
this, every retry construct in the store is bounded: purge makes
guaranteed progress per call, deletion is a single idempotent
transaction, fences raise stale, and reads are single-pass -- the
creation paths were the only sites with retries to exhaust.

Co-Authored-By: Claude Fable 5 <[email protected]>

* refactor: attempts, not retries

SegmentStoreAttemptsExhaustedError, with the counter and docstrings
using the same word: "retry" is ambiguous between a re-attempt and the
whole attempt sequence, and _MAX_MINT_ATTEMPTS already counted
attempts.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: attempts vocabulary in the mint-exhaustion message

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: increase max mint attempts from 8 to 10

Signed-off-by: Edwin Yu <[email protected]>

* nit: manual formatting

Signed-off-by: Edwin Yu <[email protected]>

* review: fold the read fence into the data statement; harden creation, drain, purge

Third review round (15 findings; 12 acted on, 3 declined with grounds
in the PR body).

- Reads no longer issue a separate registry round trip: the liveness
  predicate rides in each data statement as an EXISTS conjunct (one
  statement, one snapshot -- a stale handle reads nothing), and the
  registry check is issued on its own only when a read returns no rows,
  to tell an empty partition from a stale handle. The write fence and
  the read check share one query builder (`_registry_row_query`) and
  one checker (`_ensure_partition_live(pin=...)`).
- `create_partition` materializes the payload codec before inserting,
  like `open_or_create_partition`; its mint loop uses the same
  attempt-counter idiom and message as the other path, which also
  removes the possibly-unbound `last_collision`.
- `drop_session_partition` nulls its handles before the inline drain,
  so a drain failure cannot leave them pointing at deleted resources;
  the drain's comment states exactly what it guarantees (the queue is
  global, the drain uncapped, and an entry a concurrent drain claimed
  is finished by that drain).
- The purge queue's enqueue stamp is the database clock (`now()`), so
  every server's entries order on one clock; the unreachable
  `remaining <= 0` guard is gone; the purge comment and design doc
  state SQLite's actual claiming behavior (plain read, serialized on
  the database write lock at the DELETE; duplicated round trips only).
- `startup()` refuses the old partitioned layout (registry without the
  incarnation column) with a directive to recreate the schema, instead
  of letting create_all leave the old tables in place for an opaque
  missing-column error later.
- Cluster store reads use the shared `ensure_tz_aware`; the private
  clone is deleted. Contract wording: "every data operation" raises
  the stale-handle error (the config property never did).

Tests: unloadable codec guard parametrized over both creation paths,
FIFO pinned with explicit stamps set against insertion order, the
database-clock stamp and the folded liveness check pinned via recorded
SQL, the startup probe on both dialects, and the LTM nulling order
under a failing drain. The codec and nulling tests were each verified
to fail with their fix ablated.

Read-path ABAB against the previous HEAD (interleaved rounds, medians):
seed context reads 1.17 vs 1.42 ms, event lookups 1.04 vs 1.61 ms,
derivative lookups 1.05 vs 1.33 ms (5 rounds), windowed context
expansion 8.74 vs 9.67 ms (8 rounds x 600 reps, paired median
-0.91 ms); reads that find nothing unchanged (two statements either
way). Server-side EXPLAIN ANALYZE: the EXISTS conjunct plans as a
one-time InitPlan (~3 us per statement); a windowed read's 3
statements execute in 0.069 ms vs the previous 4 statements' 0.067 ms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: purge bound scales with link fan-out; queue stamp is transaction time

Follow-up notes from the review session: the purge_max_segments
description says its derivative links cascade uncounted, so a call's
transaction also scales with the deployment's links per segment (the
promise in the ABC is kept by sizing this bound with that fan-out in
mind, which is the deployment's knob, not the caller's); the enqueue
stamp comment records that PostgreSQL's now() is transaction-start
time and that one deletion per transaction makes it one stamp per
entry.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: a deleted partition's handle is permanently invalid

"Obtain a fresh handle to continue" read as if deletion-and-recreation
were a routine flow; the contract is simply that deletion permanently
invalidates the handle, including against a later same-key creation.

Co-Authored-By: Claude Fable 5 <[email protected]>

* revert: drop the old-layout startup probe

Handling pre-existing partitioned-layout deployments is out of scope
for the opt-in, pre-GA event backend; existing databases recreate
their schema, and startup stays a plain create_all.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: slim the purge claiming comment

The code comment keeps only the invariants the loop relies on; the
full rationale stays in design/segment_store_shared_tables.md, which
the comment now points at.

Co-Authored-By: Claude Fable 5 <[email protected]>

* nit: params docstring matches field descriptions, defaults at the end

The purge bounds' field descriptions carry the full text and the
docstring repeats them verbatim, with (default: N) moved to the end
of each description per the params convention.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: second-round fixes across locator, stores, and race tests

Second review round from the local review session (9 findings; 7
acted on here, the background-purger suggestion lands separately, and
the unbounded link-retire guard is kept with its tradeoff stated in a
comment -- bounding it would add a budget for a case that indicates a
broken schema).

- partition_key_for_session validated with a drifted private copy of
  the store's key contract: its re.match passed a trailing-newline
  session id through unhashed, and the store's re.fullmatch then
  hard-rejected it, failing session creation where hashing would have
  succeeded. The copy is deleted; the locator (and its tests) now call
  the store's own validate_partition_key, and the hash slice length
  comes from the now-public PARTITION_KEY_MAX_BYTES, so the two can
  never disagree again. Regression test verified to fail pre-fix.
- The SegmentStorePartition contract states that a call with empty
  input does no work and returns without checking the handle -- the
  empty-set guards return before any fence, which the docstring's
  "from then on" overstated.
- delete_partition on SQLite resolves the incarnation in the pin
  UPDATE itself via RETURNING; the locking select is PostgreSQL's path
  only, removing SQLite's extra round trip and its unreachable
  row-is-None branch.
- _open_or_create_partition loads the payload codec only on the create
  path (still before any registry write); opening an existing
  partition no longer materializes a codec it discards.
- Episode-store reads use ensure_tz_aware instead of an inline clone
  in the same file that imports it for writes.
- The purge's link-retire guard comment states it is normally a
  zero-row delete and unbounded only if referential integrity was
  actually broken.
- The two SQLite race tests gained started-events proving the racing
  task ran before the sample, so a loaded box cannot pass them
  vacuously by never scheduling it; the remaining grace periods are
  annotated (SQLite exposes no lock-wait state to observe).

Co-Authored-By: Claude Fable 5 <[email protected]>

* feat: background purge tick in the resource manager

Nothing but the inline drain in drop_session_partition ever called
purge_deleted_partitions, so a drain interrupted by a crash or a
dropped connection left its queue entry (and the partition's rows)
waiting for the next session deletion anywhere in the deployment.

The resource manager -- the component that owns each segment store --
now runs one background task per store: one bounded purge call per
fixed tick, exceptions logged and retried next tick, cancelled in
close() before the stores shut down. One call per tick keeps the
background work bounded by construction (a backlog drains over
successive ticks), and no purger coordination is needed at any
instance count because the store's claiming already makes racing
purgers safe. The store itself still never schedules reclamation;
this loop is the caller-side scheduler the ABC contract calls for.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: purge loop reads the backlog signal

purge_deleted_partitions() returning True is the API's statement that
more work remains; discarding it drained a backlog at one bounded
call per tick (~167 rows/s at the defaults). The loop now runs
bounded calls back-to-back while the store reports more and sleeps
one tick only when it reports done or a call fails -- full-rate
recovery, still bounded per call, still one idle call per tick.
Pinned by a test that drains a three-call backlog under a
deliberately huge tick interval.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: empty-input calls MAY skip the handle check

The contract permits the shortcut rather than mandating it; an
implementation that checks anyway still conforms.

Co-Authored-By: Claude Fable 5 <[email protected]>

* test: pin the config-mismatch guard directly

The guard, its error type, and the ABC declaration predate this
branch, but no test staged a mismatch -- only the lifecycle-churn
test tolerated it as a domain outcome. Plaintext is the only concrete
codec config, so the test stands in a subclass for a future variant
(pydantic instances survive validation unrevalidated and compare
unequal by class). Verified to fail with the guard ablated.

Co-Authored-By: Claude Fable 5 <[email protected]>

* purge: bound the integrity-escape link delete; warn on it and on collisions

The retire-path guard delete was the one unbounded statement in the
purge, unbounded precisely when it was not a no-op. It is now batched
under the same per-call budget as the segment rows: a full batch
leaves the queue entry for the next call (the existing call-until-
False contract absorbs it, callers unchanged), the normal case still
costs one zero-row statement, and reclaiming rows there logs a
warning naming the incarnation, since it means referential integrity
failed somewhere. Pinned by a test that stages orphan link rows
through a second SQLite engine without the foreign-key pragma and
drains them in warned batches; verified to fail against the unbounded
form.

The module's logger also gains the only other events worth an
operator's attention: a minted incarnation colliding (with garbage or
in the registry) is warning-logged at the detection site -- a genuine
collision is astronomically unlikely, so the log marks either broken
randomness or a misclassified persistent database error, visible even
when retries eventually succeed. Everything else either raises to the
caller or is normal operation, and stays unlogged.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: shared purge budget is measured, not a guessed ratio

Measured with the purge's batched-delete shape on 100k rows each
(3 interleaved rounds): a segment row deletes at ~3.3 us and a
derivative-link row at ~1.0 us, so link rows are about 3x cheaper --
they are narrower, carry fewer indexes, and fire no cascade. That is
why integrity-escaped links draw count-for-count on the segment
budget instead of getting their own limit: one budget calibrated on
the most expensive row type upper-bounds the call, whereas a separate
link limit would be safe only under an assumed cost ratio.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the row-cost direction is the shared budget's precondition

The shared purge budget stays conservative only while a link row
deletes cheaper than a segment row; widening the link table or adding
indexes to it revisits the choice.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: pace the purger, state the drain's real guarantee, close cleanly

Round-3 findings 1, 2, 5 and 6 -- the first three introduced by the
background purger itself.

- The purge loop now pauses briefly after every productive call
  instead of running delete transactions back-to-back: the pause
  yields the database (and SQLite's single write lock) to request
  serving, while a backlog still drains at one bounded call per pause
  and an idle store costs one call per tick. This also removes the
  in-process busy-timeout window between the background task and the
  inline drain on SQLite.
- The inline drain's comment claimed "the server schedules no other
  purger", which the purger commit falsified, and "reclaimed before
  returning", which SKIP LOCKED claiming never strictly guaranteed
  under any concurrent purger. Comment, design doc, and PR body now
  state the actual promise: rows are reclaimed promptly -- normally
  before the drain returns, and otherwise within the bounded call of
  whichever purger claimed the entry, moments later.
- close() clears the purge-task list and the store registry, so a
  second close is a no-op and a post-close get_segment_store can no
  longer hand back a shut-down store that silently never purges.
- The design doc no longer implies deployments can already tune the
  purge bounds through server configuration: the server constructs
  its stores with the defaults, and config plumbing is future work.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: one datetime-normalization rule per filter path

Round-3 findings 3 and 11.

- The properties_json In branch bound raw values while its Comparison
  sibling normalized through _cast_properties_json_value; a datetime
  In list would bind datetime objects against the stored ISO-string
  form (an InterfaceError on Python 3.14's sqlite3, a never-matching
  comparison on PostgreSQL). Both leaf shapes now cast and normalize
  through the one function, which also aligns the float and bool
  casts the old branch fell through to as_string/as_integer. Same
  defensive-reachability status as the column-leaf In normalization
  kept deliberately: unreachable by In's declared value types,
  reachable at runtime.
- The episode store's start_time/end_time bounds re-implemented the
  UTC normalization inline; sql_filter_util's normalize_column_value
  is now public and both bounds use it, so the storage convention has
  one definition across compiled filters and dedicated bounds.

Co-Authored-By: Claude Fable 5 <[email protected]>

* review: StaticPool guard raises; empty add_segments short-circuits

Round-3 findings 9 and 10.

- The params validator's StaticPool guard was a bare assert, stripped
  under python -O -- and the design depends on multiple connections
  (the registry fence, deletion waiting out writers, SKIP LOCKED
  claiming all degrade on one shared connection). It now raises
  ValueError like the ephemeral-SQLite check beside it; pinned by a
  test, and the check stays a ValueError because pydantic converts
  only ValueError/AssertionError into a ValidationError.
- add_segments returns early on empty input, matching delete_segments
  and the ABC's empty-input permission; previously it opened a
  transaction and, on SQLite, took the write lock to insert nothing.
  The stale-handle test pins the shortcut.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: datetime values normalize to UTC at node construction

The filter language now owns datetime semantics: a value denotes an
instant, and a naive value means UTC. Comparison.__post_init__
normalizes datetime values to UTC-aware instants (In gets the same
defensively -- its declared types exclude datetimes, but runtime lists
are unchecked), so every consumer -- parsed trees and programmatically
built ones, SQL compilers and vector stores alike -- receives
normalized instants by construction, and compilers only choose a
representation.

This is where the rule the recent fixes kept restating per leaf
actually belongs: the same aware-to-UTC-or-naive-means-UTC conversion
appeared in the SQL column leaf, the properties_json leaf, the episode
bounds, and twice in the Milvus store, and two of the drifted copies
were bugs fixed this round. With the invariant at the node, the SQL
column leaf's re-normalization became redundant and is reverted (it
binds tree values as-is); the properties_json leaf keeps its routing
because datetime-to-ISO-string is representation, not normalization;
the episode start/end bounds keep the shared helper because they are
raw API values outside any tree; other backends' now-idempotent
defenses are left for separate cleanup.

Pinned by tests that a programmatically built Comparison and a parsed
date() literal with a non-UTC offset both carry the UTC instant.

Co-Authored-By: Claude Fable 5 <[email protected]>

* filter: drop normalize_column_value; contract stated on the protocol

With datetime normalization at node construction, the compiler-side
helper had no filter role left, and its one remaining consumer -- the
episode store's start/end bounds, which arrive outside any filter
tree -- now spells the convention inline as the two explicit steps,
ensure_tz_aware(...).astimezone(UTC). A composed to_utc() helper was
considered and rejected: the name does not pin the naive-means-UTC
tagging decision (an alternative design under the same name could
reject naive datetimes entirely), so the explicit steps are clearer
at each site.

The FilterExpr protocol docstring now states the construction-time
contract where the next value-carrying node's author will read it:
such a node normalizes datetime values to UTC-aware instants, and
compilers bind instants without re-normalizing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: the two-step datetime spelling is deliberate

Record the datetime convention in the design doc so a future cleanup
does not consolidate the repeated ensure_tz_aware(...).astimezone(UTC)
sequences back into the composed helper b636d62a deliberately removed:
a name that pins only the conversion, not the naive-means-UTC tagging,
hides a real design decision, so the repetition is load-bearing.

Co-Authored-By: Claude Fable 5 <[email protected]>

* docs: second ground for rejecting the composed datetime helper

A shared helper earns its place only when the name honestly pins the
unit AND the composition structurally prevents half-applied
normalization. The second condition fails here regardless of naming:
read paths legitimately need the tagging step alone (segment reads
reapply the stored original offset; cluster and episode reads only tag
naive database values), so ensure_tz_aware stays independently
available and the helper could not have removed the partial-use error
class.

Co-Authored-By: Claude Fable 5 <[email protected]>

* fix: SQLite foreign keys enforced from engine creation

Round-4 findings 1, 4 and 5. The store registered its foreign_keys
pragma as a per-store connect listener, which has two structural
faults: connections the caller's shared engine pooled before the…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant