This is the full developer documentation for TypeGraph # What is TypeGraph? > A TypeScript-first embedded knowledge graph library TypeGraph is a **TypeScript-first, embedded knowledge graph library** that brings property graph semantics and ontological reasoning to applications using standard relational databases. Rather than introducing a separate graph database, TypeGraph lives inside your application as a library, storing graph data in your existing SQLite or PostgreSQL database. ## Architecture ![TypeGraph Architecture: Your application imports TypeGraph as a library dependency. TypeGraph uses Drizzle ORM to store graph data (nodes, edges, schema, ontology) in your existing SQLite or PostgreSQL database. No separate graph database required.](../../assets/typegraph-architecture.svg) ## Core Capabilities ### 1. Type-Driven Schema Definition Zod schemas are the single source of truth. From one schema definition, TypeGraph derives: - Runtime validation rules - TypeScript types (inferred, not duplicated) - Database storage requirements - Query builder type constraints ```typescript const Person = defineNode("Person", { schema: z.object({ fullName: z.string().min(1), email: z.string().email().optional(), dateOfBirth: z.date().optional(), }), }); ``` ### 2. Semantic Layer with Ontological Reasoning Type-level relationships enable sophisticated inference: | Relationship | Meaning | Use Case | | -------------- | ----------------------------------------- | ---------------------- | | `subClassOf` | Instance inheritance (Podcast IS-A Media) | Query expansion | | `broader` | Hierarchical concept (ML broader than DL) | Topic navigation | | `equivalentTo` | Same concept, different name | Cross-system mapping | | `disjointWith` | Cannot be both (Person ≠ Organization) | Constraint validation | | `implies` | Edge entailment (marriedTo implies knows) | Relationship inference | | `inverseOf` | Edge pairs (manages/managedBy) | Bidirectional queries | ### 3. Self-Describing Schema (Homoiconic) The schema and ontology are stored in the database as data, enabling: - Runtime schema introspection - Versioned schema history - Self-describing exports and backups - Migration tooling ### 4. Type-Safe Query Compilation Queries compile to an AST before targeting SQL: - Consistent semantics across SQLite and PostgreSQL - Type-checked at compile time - Query results have inferred types ### 5. Temporal and Bitemporal History Every node and edge has a valid-time window (`validFrom` / `validTo`), so you can ask what was true at a domain instant with `.temporal("asOf", T)` or `store.asOf(T)`. Stores created with `{ history: true }` also capture recorded time for TypeGraph-managed writes, the system-time axis that remembers when the graph wrote each fact down. `store.asOfRecorded(T)` reconstructs what the graph captured at a recorded instant, and `store.asOf(validT).asOfRecorded(recordedT)` pins both axes independently. Use it for audit trails, agent decision replay, effective-dated policies, and breach forensics. See [Temporal queries](/queries/temporal) and the [Bitemporal Time Travel](/examples/bitemporal-time-travel) example. ## Design Philosophy ### Embedded, Not External TypeGraph is a library dependency, not a networked service. TypeGraph initializes with your application, uses your database connection, and requires no separate deployment. ### Schema-First, Type-Driven Define your schemas once with Zod, and TypeGraph handles validation, type inference, and storage. No duplicate type definitions or manual synchronization. ### Explicit Over Implicit TypeGraph favors explicit declarations: - Relationships are declared, not inferred from foreign keys - Semantic relationships are explicit in the ontology - Cascade behavior is configured, not assumed ### Portable Abstractions The query builder generates portable ASTs that can target different SQL dialects. The same query code works with SQLite and PostgreSQL. ## What TypeGraph Is Not TypeGraph deliberately excludes: - **Broad graph analytics suites**: Focused PageRank, connectivity, and deterministic label-propagation primitives are built in; modularity optimization and most centrality measures are not - **Distributed storage**: Single-database deployment only These exclusions keep TypeGraph focused and maintainable. Note: TypeGraph **does support** semantic search via native database vector engines: pgvector for PostgreSQL, sqlite-vec for the local (better-sqlite3) SQLite backend, and libSQL's built-in vectors for the libSQL / Turso backend. See [Semantic Search](/semantic-search) for details. Note: TypeGraph **does support** fulltext search — native BM25 on SQLite (FTS5) and `tsvector` + GIN on PostgreSQL, with a query-builder `n.$fulltext.matches()` predicate that composes with any other predicate. Combine with semantic search for hybrid RAG retrieval. See [Fulltext Search](/fulltext-search) for details. Note: TypeGraph does support **variable-length paths** via `.recursive()` with configurable depth limits, optional path/depth projection, and explicit cycle policy. Cycle prevention is the default. See [Recursive Traversals](/queries/recursive) for details. Note: TypeGraph ships **Tier 1 graph algorithms** (shortest path, reachability, neighborhoods, and degree) on `store.algorithms.*`. Traversal calls use a set-based BFS frontier, while degree uses a single count query. See [Graph Algorithms](/graph-algorithms) for details. Note: TypeGraph supports **runtime schema induction** via graph extensions. An LLM or ingestion agent can propose a typed schema as a JSON-serializable document, an operator approves it, and `store.evolve()` atomically commits a new schema version — no redeploy, full Zod validation, restart parity. See [Graph Extensions](/graph-extensions) for the agent-driven workflow. Note: TypeGraph ships **graph merge** — fork a store into isolated working copies, let many writers (parallel agents, importers, reviewers) edit independently, then reconcile them into one canonical graph with deterministic entity resolution (exact / blocking / fulltext / vector / hybrid), edge repointing, conflict reporting, and provenance. `mergeIncremental()` folds new sources into a *live* graph without creating duplicates — the primitive for multi-agent knowledge-graph construction and continuous ingestion. See [Graph Merge](/graph-merge) for the full guide. Note: TypeGraph supports **bitemporal graph reads**. Valid time answers "when was this fact true in the domain?" Recorded time answers "when did the graph record it?" Together they reconstruct prior captured state after corrections, replay agent decisions against the graph they actually saw, and traverse access graphs at a breach instant for TypeGraph-managed writes. See [Temporal queries](/queries/temporal) and the [Agent Decision Replay](/examples/agent-decision-replay) example. ## Why TypeGraph? ### Compared to Graph Databases (Neo4j, Amazon Neptune) Graph databases are powerful but come with operational overhead: | Aspect | Graph Database | TypeGraph | |--------|---------------|-----------| | **Deployment** | Separate service to manage, scale, and monitor | Library in your app, uses existing database | | **Network** | Additional latency for every query | In-process, no network hop | | **Transactions** | Separate transaction scope from your SQL data | Same ACID transaction as your other data | | **Learning curve** | New query language (Cypher, Gremlin) | TypeScript you already know | | **Graph algorithms** | Broad suites (PageRank, shortest path, community detection) | Focused algorithms (shortest path, reachability, neighborhoods, degree, WCC, label propagation, PageRank/PPR) | | **Scale** | Optimized for billions of nodes | Best for thousands to millions | **Choose TypeGraph** when your graph is part of your application domain (knowledge bases, org charts, content relationships) rather than a standalone analytical system. ### Compared to ORMs (Prisma, Drizzle, TypeORM) ORMs model relations through foreign keys, which works well for simple associations but lacks graph semantics: | Aspect | Traditional ORM | TypeGraph | |--------|----------------|-----------| | **Relationships** | Foreign keys, eager/lazy loading | First-class edges with properties | | **Traversals** | Manual joins or N+1 queries | Fluent traversal API, compiled to efficient SQL | | **Inheritance** | Table-per-class or single-table | Semantic `subClassOf` with query expansion | | **Constraints** | Foreign key constraints | Disjointness, cardinality, implications | | **Schema** | Migrations alter tables | Schema versioning, JSON properties | **Choose TypeGraph** when you need to traverse relationships, model type hierarchies, or enforce semantic constraints beyond what foreign keys provide. ### Compared to Triple Stores (RDF, SPARQL) Triple stores and RDF provide rich ontological modeling but have practical challenges: | Aspect | Triple Store | TypeGraph | |--------|-------------|-----------| | **Type safety** | Runtime validation, stringly-typed | Full TypeScript inference | | **Query language** | SPARQL (powerful but verbose) | TypeScript fluent API | | **Schema** | OWL/RDFS (complex specification) | Zod schemas (familiar, composable) | | **Integration** | Separate system, data sync required | Embedded in your app | | **Inference** | Full reasoning engines available | Precomputed closures, practical subset | **Choose TypeGraph** when you want ontological concepts (subclass, disjoint, implies) without the complexity of full semantic web stack. ### The TypeGraph Sweet Spot TypeGraph is designed for applications where: 1. **The graph is your domain model** — not a separate analytical system 2. **You already use SQL** — and don't want another database to manage 3. **Type safety matters** — you want compile-time checking, not runtime surprises 4. **Semantic relationships help** — inheritance, implications, constraints add value 5. **Scale is moderate** — thousands to millions of nodes, not billions ## When to Use TypeGraph TypeGraph is ideal for: - **Knowledge bases** with typed entities and relationships - **Organizational structures** with hierarchies and roles - **Content graphs** with topics, articles, and references - **Domain models** requiring semantic constraints - **RAG applications** combining graph traversal with vector search - **Multi-source ingestion & entity resolution** — reconcile parallel agent or importer outputs into one canonical graph with [graph merge](/graph-merge) - **Auditable AI systems and forensics** — reconstruct the graph an agent or investigator saw at a recorded instant with [bitemporal reads](/queries/temporal#recorded-time-bitemporal) TypeGraph is not ideal for: - Large-scale graph analytics requiring distributed processing - Social networks with billions of edges - Real-time streaming graph data - Applications requiring a broad graph-data-science suite such as community detection or betweenness centrality (use Neo4j or a graph library; focused algorithms—including PageRank and weighted shortest path—ship on `store.algorithms.*`) # Quick Start > Set up TypeGraph and build your first knowledge graph Get TypeGraph running in your project with this minimal example. ## 1. Install ```bash npm install @nicia-ai/typegraph zod drizzle-orm better-sqlite3 npm install -D @types/better-sqlite3 ``` > **Edge environments or libsql:** Skip `better-sqlite3` and use > `@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql` with `@libsql/client`, or > `@nicia-ai/typegraph/adapters/drizzle/sqlite` with your edge-compatible driver (D1, bun:sqlite). > See [Backend Setup](/backend-setup#libsql--turso) and [Edge and Serverless](/integration#edge-and-serverless). ## 2. Create Your First Graph ```typescript import { z } from "zod"; import { defineNode, defineEdge, defineGraph } from "@nicia-ai/typegraph"; import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local"; // Define your schema const Person = defineNode("Person", { schema: z.object({ name: z.string(), role: z.string().optional() }), }); const Project = defineNode("Project", { schema: z.object({ name: z.string(), status: z.enum(["active", "done"]) }), }); const worksOn = defineEdge("worksOn"); const graph = defineGraph({ id: "my_app", nodes: { Person: { type: Person }, Project: { type: Project } }, edges: { worksOn: { type: worksOn, from: [Person], to: [Project] } }, }); // Provision an in-memory database and create the store const store = await createLocalSqliteStore(graph); // Use it! const alice = await store.nodes.Person.create({ name: "Alice", role: "Engineer" }); const project = await store.nodes.Project.create({ name: "Website", status: "active" }); await store.edges.worksOn.create(alice, project, {}); // Query with full type safety const results = await store .query() .from("Person", "p") .traverse("worksOn", "e") .to("Project", "proj") .select((ctx) => ({ person: ctx.p.name, project: ctx.proj.name })) .execute(); console.log(results); // [{ person: "Alice", project: "Website" }] ``` That's it! You have a working knowledge graph. Read on for the complete setup guide. This managed entrypoint returns the complete typed `Store` while keeping its public declaration surface independent of Drizzle. Use `@nicia-ai/typegraph/postgres/pglite` for the same setup with in-process PostgreSQL. If your application owns the database connection or needs direct driver access, use the adapter entrypoints described below instead. --- ## Complete Setup Guide This section covers production setup with SQLite and PostgreSQL in detail. ### Installation ```bash npm install @nicia-ai/typegraph zod drizzle-orm better-sqlite3 npm install -D @types/better-sqlite3 ``` > `better-sqlite3` is optional. For libsql/Turso, use `@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql`. > For D1 or bun:sqlite, use `@nicia-ai/typegraph/adapters/drizzle/sqlite` with the matching Drizzle driver. ### SQLite Setup TypeGraph provides two ways to set up SQLite: #### Managed Store (Recommended) Use the managed Store when TypeGraph should own the connection and provision its schema: ```typescript import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local"; const store = await createLocalSqliteStore(graph, { path: "./my-app.db" }); // The Store owns the connection. await store.close(); ``` The return value is a `Store`: typed node and edge collections, queries, algorithms, graph-owned transactions, schema evolution, and schema-derived property types are all available. Adapter-native handles and caller-owned transaction adoption are absent by design; opt into `AdapterStore` through a Drizzle adapter entrypoint when application tables must share a transaction. #### Quick Setup (Recommended for Development) Use the backend wrapper when you also need the underlying Drizzle database or want to choose how the Store is created. > **Note:** `createLocalSqliteBackend` requires `better-sqlite3` and only works in Node.js. > For edge environments, see [Manual Setup](#manual-setup-full-control) with > `/adapters/drizzle/sqlite`. ```typescript import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; // In-memory database (data lost on restart) const { backend } = createLocalSqliteBackend(); // File-based database (persistent) const { backend, db } = createLocalSqliteBackend({ path: "./my-app.db" }); ``` The function returns both the `backend` (for use with `createStore`) and `db` (the underlying Drizzle instance for direct SQL access if needed). #### Manual Setup (Full Control) For production deployments or when you need full control over the database configuration: ```typescript import Database from "better-sqlite3"; import { drizzle } from "drizzle-orm/better-sqlite3"; import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; // Create database connection const sqlite = new Database("my-app.db"); // Run TypeGraph migrations (creates required tables) sqlite.exec(generateSqliteMigrationSQL()); // Create Drizzle instance const db = drizzle(sqlite); // Create the backend const backend = createSqliteBackend(db); ``` #### libsql / Turso Setup For Turso, embedded replicas, or sharing a libsql connection with other libraries: ```typescript import { createClient } from "@libsql/client"; import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql"; const client = createClient({ url: "file:app.db" }); const { backend } = await createLibsqlBackend(client); ``` `createLibsqlBackend` handles DDL automatically. The caller owns the client and is responsible for closing it. See [Backend Setup](/backend-setup#libsql--turso) for remote Turso URLs and caveats. #### Edge-Compatible Setup (D1, bun:sqlite) For Cloudflare Workers or Bun, use the driver-agnostic backend: ```typescript import { drizzle } from "drizzle-orm/d1"; // or bun-sqlite import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; // D1 example const db = drizzle(env.DB); const backend = createSqliteBackend(db); ``` Use [drizzle-kit managed migrations](/integration#drizzle-kit-managed-migrations-recommended) to set up the schema. #### Drizzle-Kit Managed Migrations If you already use `drizzle-kit` for migrations, see [Drizzle-Kit Managed Migrations](/integration#drizzle-kit-managed-migrations-recommended) for how to import TypeGraph's schema into your `schema.ts` file. ## Defining Your Schema ### Step 1: Define Node Types Nodes represent entities in your graph. Each node type has a name and a Zod schema: ```typescript import { z } from "zod"; import { defineNode } from "@nicia-ai/typegraph"; const Person = defineNode("Person", { schema: z.object({ name: z.string().min(1), email: z.string().email().optional(), bio: z.string().optional(), }), }); const Project = defineNode("Project", { schema: z.object({ name: z.string(), description: z.string().optional(), status: z.enum(["planning", "active", "completed"]), }), }); const Task = defineNode("Task", { schema: z.object({ title: z.string(), priority: z.enum(["low", "medium", "high"]), completed: z.boolean().default(false), }), }); ``` ### Step 2: Define Edge Types Edges represent relationships between nodes: ```typescript import { defineEdge } from "@nicia-ai/typegraph"; const worksOn = defineEdge("worksOn", { schema: z.object({ role: z.string().optional(), since: z.string().optional(), }), }); const hasTask = defineEdge("hasTask", { schema: z.object({}), }); const assignedTo = defineEdge("assignedTo", { schema: z.object({ assignedAt: z.string().optional(), }), }); // Unconstrained edge — connects any node to any node const related = defineEdge("related"); ``` ### Step 3: Create the Graph Definition Combine nodes, edges, and ontology into a graph: ```typescript import { defineGraph, disjointWith } from "@nicia-ai/typegraph"; const graph = defineGraph({ id: "project_management", nodes: { Person: { type: Person }, Project: { type: Project }, Task: { type: Task }, }, edges: { worksOn: { type: worksOn, from: [Person], to: [Project] }, hasTask: { type: hasTask, from: [Project], to: [Task] }, assignedTo: { type: assignedTo, from: [Task], to: [Person] }, related, // any→any }, ontology: [ // A Person cannot be a Project or Task disjointWith(Person, Project), disjointWith(Person, Task), disjointWith(Project, Task), ], }); ``` ### Step 4: Create the Store The store connects your graph definition to the database: ```typescript import { createStore } from "@nicia-ai/typegraph"; const store = createStore(graph, backend); ``` #### Store Creation: Which Function to Use | Function | Schema Handling | Use Case | | -------------------------------- | ---------------------------------------------- | -------------------------------------------------- | | `createLocalSqliteBackend` | Automatic | Quick start, development, tests (Node.js) | | `createLibsqlBackend` | Automatic | libsql/Turso (Node.js, Workers, browser) | | `createLocalPgliteBackend` | Automatic | In-process Postgres, embedded apps, pgvector tests | | `createStore` + manual migration | None | When you manage migrations externally | | `createStoreWithSchema` | Auto-creates tables, validates & auto-migrates | **Recommended for production** | :::caution[Fulltext requires `createStoreWithSchema`] If your graph has any `searchable()` fields, you must boot through `createStoreWithSchema` once at startup. It durably materializes the fulltext storage; bare `createStore()` is an attach-only path and throws `StoreNotInitializedError` on the first fulltext operation. Graphs without `searchable()` fields are unaffected. ::: For production, use `createStoreWithSchema` to validate and auto-apply safe schema changes: ```typescript import { createStoreWithSchema } from "@nicia-ai/typegraph"; const [store, result] = await createStoreWithSchema(graph, backend); if (result.status === "initialized") { console.log("Schema initialized at version", result.version); } else if (result.status === "migrated") { console.log(`Migrated from v${result.fromVersion} to v${result.toVersion}`); } // Other statuses: "unchanged", "pending", "breaking" // See Schema Migrations for full details ``` #### Graph ID Every graph has a unique `id` that scopes its data: ```typescript const graph = defineGraph({ id: "my_app", // Scopes all nodes/edges to this graph // ... }); ``` **Key behaviors:** - All nodes and edges are stored with this `graph_id` in the database - Multiple graphs can share the same database tables (isolated by `graph_id`) - Changing the ID creates a new, empty graph (existing data is orphaned) See [Multiple Graphs](/multiple-graphs) for multi-graph deployments. ## Working with Data ### Creating Nodes ```typescript const alice = await store.nodes.Person.create({ name: "Alice Smith", email: "alice@example.com", }); const project = await store.nodes.Project.create({ name: "Website Redesign", status: "active", }); const task = await store.nodes.Task.create({ title: "Design mockups", priority: "high", }); ``` ### Creating Edges Pass node objects directly to create edges: ```typescript await store.edges.worksOn.create(alice, project, { role: "Lead Designer" }); await store.edges.hasTask.create(project, task, {}); await store.edges.assignedTo.create(task, alice, { assignedAt: new Date().toISOString() }); ``` ### Retrieving Nodes ```typescript const person = await store.nodes.Person.getById(alice.id); console.log(person?.name); // "Alice Smith" ``` ### Updating Nodes ```typescript const updated = await store.nodes.Task.update(task.id, { completed: true }); ``` ### Deleting Nodes ```typescript await store.nodes.Task.delete(task.id); ``` ## Querying Data TypeGraph provides a fluent query builder: ```typescript // Find all active projects const activeProjects = await store .query() .from("Project", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ctx.p) .execute(); // Find people working on a project const teamMembers = await store .query() .from("Project", "p") .traverse("worksOn", "e", { direction: "in" }) .to("Person", "person") .select((ctx) => ({ project: ctx.p.name, person: ctx.person.name, })) .execute(); // Multi-hop traversal: find tasks for a person const myTasks = await store .query() .from("Person", "person") .whereNode("person", (p) => p.name.eq("Alice Smith")) .traverse("worksOn", "e1") .to("Project", "project") .traverse("hasTask", "e2") .to("Task", "task") .select((ctx) => ({ project: ctx.project.name, task: ctx.task.title, priority: ctx.task.priority, })) .execute(); ``` `whereNode()` and `whereEdge()` constrain graph matches while traversal is built. Use `.where((ctx) => ...)` when a condition should filter completed rows, including optional or recursive results. Each successful traversal combination is one row, so fanout can repeat a source entity; project an identity and call relation `.distinct()` when the intended result is one row per entity. Traversal continues from the latest target by default. Reusable branching fragments should state their source explicitly with `{ from: "alias" }` so their behavior does not depend on which traversal preceded them. Direction, ontology expansion, and temporal coordinates retain their ordinary query defaults. ## Transactions Group operations in transactions for atomicity: ```typescript await store.transaction(async (tx) => { const project = await tx.nodes.Project.create({ name: "New Feature", status: "planning", }); const task1 = await tx.nodes.Task.create({ title: "Research", priority: "high", }); const task2 = await tx.nodes.Task.create({ title: "Implementation", priority: "medium", }); await tx.edges.hasTask.create(project, task1, {}); await tx.edges.hasTask.create(project, task2, {}); }); ``` ## Error Handling TypeGraph provides specific error types: ```typescript import { ValidationError, NodeNotFoundError, DisjointError, RestrictedDeleteError } from "@nicia-ai/typegraph"; try { await store.nodes.Person.create({ name: "" }); // Invalid: empty name } catch (error) { if (error instanceof ValidationError) { console.log("Validation failed:", error.message); } } try { await store.nodes.Project.delete(project.id); } catch (error) { if (error instanceof RestrictedDeleteError) { console.log("Cannot delete: edges exist"); } } ``` ## PostgreSQL Setup TypeGraph also supports PostgreSQL for production deployments with better concurrency and JSON support. For in-process Postgres during local development or tests, see [PGlite in Backend Setup](/backend-setup#pglite-postgres-in-wasm). ### Installation ```bash npm install @nicia-ai/typegraph zod drizzle-orm pg npm install -D @types/pg ``` ### Database Setup ```typescript import { Pool } from "pg"; import { drizzle } from "drizzle-orm/node-postgres"; import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; // Create connection pool const pool = new Pool({ connectionString: process.env.DATABASE_URL, max: 20, // Connection pool size }); // Run TypeGraph migrations await pool.query(generatePostgresMigrationSQL()); // Create Drizzle instance and backend const db = drizzle(pool); const backend = createPostgresBackend(db); ``` If you use `drizzle-kit` for migrations, see [Drizzle-Kit Managed Migrations](/integration#drizzle-kit-managed-migrations-recommended). ### PostgreSQL Advantages - **JSONB**: Native JSON type with efficient indexing - **Connection pooling**: Better concurrency handling - **Partial indexes**: More efficient uniqueness constraints - **Full transactions**: ACID guarantees across operations ### Using with Connection Pools For production, always use connection pooling: ```typescript import { Pool } from "pg"; const pool = new Pool({ connectionString: process.env.DATABASE_URL, max: 20, idleTimeoutMillis: 30000, connectionTimeoutMillis: 2000, }); // Graceful shutdown process.on("SIGTERM", async () => { await pool.end(); }); ``` ## Next Steps - [Project Structure](/project-structure) - Organize your graph definitions as your project grows - [Schemas & Types](/core-concepts) - Deep dive into nodes, edges, and schemas - [Ontology](/ontology) - Learn about semantic relationships - [Query Builder](/queries/overview) - Query patterns and traversals - [Schemas & Stores](/schemas-stores) - Complete API documentation # Schemas & Types > Defining nodes, edges, and leveraging TypeScript inference TypeGraph's power comes from its type system. Define your schema once with Zod, and get: - **Runtime validation** on every create and update - **TypeScript types** inferred automatically (no duplication) - **Query builder constraints** that prevent invalid queries at compile time ## Contents - [Nodes](#nodes) — Entities with properties and metadata - [Defining Node Types](#defining-node-types) - [Schema Features](#schema-features) - [Node Operations](#node-operations) - [Edges](#edges) — Relationships between nodes - [Defining Edge Types](#defining-edge-types) (domain/range constraints) - [Edge Constraints](#edge-constraints) (cardinality) - [Edge Operations](#edge-operations) - [Graph Definition](#graph-definition) — Combining nodes, edges, and ontology - [Delete Behaviors](#delete-behaviors) — Restrict, cascade, disconnect - [Uniqueness Constraints](#uniqueness-constraints) — Enforcing unique values - [Type Inference](#type-inference) — Extracting TypeScript types from schemas ## Nodes Nodes represent entities in your graph. Each node has: - **Type**: The type of node (e.g., "Person", "Company") - **ID**: A unique identifier within the graph - **Props**: Properties defined by a Zod schema - **Metadata**: Version, timestamps, and soft-delete state ### Defining Node Types ```typescript import { z } from "zod"; import { defineNode } from "@nicia-ai/typegraph"; const Person = defineNode("Person", { schema: z.object({ fullName: z.string().min(1), email: z.string().email().optional(), dateOfBirth: z.string().optional(), tags: z.array(z.string()).default([]), }), description: "A person in the system", // Optional }); ``` ### Schema Features TypeGraph supports all Zod validation features: ```typescript const Product = defineNode("Product", { schema: z.object({ // Required string name: z.string().min(1).max(200), // Optional with default status: z.enum(["draft", "active", "archived"]).default("draft"), // Number with constraints price: z.number().positive(), // Array with items validation categories: z.array(z.string()).min(1), // Regex pattern sku: z.string().regex(/^[A-Z]{2,4}-\d{4,8}$/), // Nullable field description: z.string().nullable(), // Transform on validation slug: z.string().transform((s) => s.toLowerCase().replace(/\s+/g, "-")), }), }); ``` ### Node Operations ```typescript // Create with auto-generated ID const node = await store.nodes.Person.create({ fullName: "Alice Smith" }); // Create with specific ID const node = await store.nodes.Person.create({ fullName: "Alice Smith" }, { id: "person-alice" }); // Retrieve const person = await store.nodes.Person.getById("person-alice"); // Update (partial) const updated = await store.nodes.Person.update("person-alice", { email: "alice@example.com", }); // Delete (soft delete by default) await store.nodes.Person.delete("person-alice"); // Hard delete (permanent removal) - use carefully! await store.nodes.Person.hardDelete("person-alice"); ``` ### Node Object Shape A node returned from the store has this structure: ```typescript const alice = await store.nodes.Person.create({ name: "Alice", email: "a@example.com" }); // alice = { // id: "01HX...", // Generated ULID (or your custom ID) // kind: "Person", // The node type name // name: "Alice", // Schema property (flattened to top level) // email: "a@example.com", // Schema property // meta: { // version: 1, // createdAt: "2024-01-15T10:30:00.000Z", // updatedAt: "2024-01-15T10:30:00.000Z", // deletedAt: undefined, // validFrom: "2024-01-15T10:30:00.000Z", // defaults to createdAt when omitted, // // unless a stated past validTo makes // // the row "born already ended" (undefined) // validTo: undefined, // } // } ``` Schema properties are flattened to the top level for ergonomic access (`alice.name` instead of `alice.props.name`). System metadata lives under `meta`. ### Soft Delete vs Hard Delete By default, `delete()` performs a **soft delete**—it sets the `deletedAt` timestamp but preserves the record: ```typescript await store.nodes.Person.delete(alice.id); // Sets deletedAt, keeps the record ``` For permanent removal, use `hardDelete()`: ```typescript await store.nodes.Person.hardDelete(alice.id); // Removes from database ``` **When to use each:** | Method | Use Case | |--------|----------| | `delete()` | Standard deletions, audit trails, undo capability | | `hardDelete()` | GDPR erasure, storage cleanup, removing test data | **Warning:** `hardDelete()` is irreversible. It also removes associated uniqueness entries and embeddings. Consider using soft delete for most use cases. ## Edges Edges represent relationships between nodes. Each edge has: - **Type**: The type of relationship (e.g., "worksAt", "knows") - **ID**: A unique identifier - **From**: Source node (type + ID) - **To**: Target node (type + ID) - **Props**: Properties defined by a Zod schema ### Defining Edge Types ```typescript import { defineEdge } from "@nicia-ai/typegraph"; // Edge with properties const worksAt = defineEdge("worksAt", { schema: z.object({ role: z.string(), startDate: z.string().optional(), isPrimary: z.boolean().default(true), }), }); // Edge without properties const knows = defineEdge("knows"); // Equivalent to: defineEdge("knows", { schema: z.object({}) }) ``` #### Unconstrained Edges Edges defined without `from` and `to` are **unconstrained** — they can connect any node type to any node type. When used directly in `defineGraph`, they are automatically allowed for all node types in the graph: ```typescript const sameAs = defineEdge("sameAs"); const related = defineEdge("related", { schema: z.object({ reason: z.string() }), }); const graph = defineGraph({ id: "my_graph", nodes: { Person: { type: Person }, Company: { type: Company }, }, edges: { sameAs, // any→any (Person↔Person, Person↔Company, Company↔Company) related, // any→any, with properties worksAt: { type: worksAt, from: [Person], to: [Company] }, // constrained }, }); // All of these work: await store.edges.sameAs.create(alice, bob, {}); // Person→Person await store.edges.sameAs.create(alice, acme, {}); // Person→Company await store.edges.sameAs.create(acme, alice, {}); // Company→Person ``` This is useful for semantic relationships like `sameAs`, `seeAlso`, `related`, or `tagged` that apply broadly across node types. #### Domain and Range Constraints Edges can include built-in domain (source types) and range (target types) constraints directly in their definition. This makes edge definitions self-contained and reusable: ```typescript // Edge with built-in domain/range constraints const worksAt = defineEdge("worksAt", { schema: z.object({ role: z.string(), startDate: z.string().optional(), }), from: [Person], // Domain: only Person can be the source to: [Company], // Range: only Company can be the target }); // Edge connecting multiple types const mentions = defineEdge("mentions", { from: [Article, Comment], to: [Person, Company, Topic], }); ``` Any edge type can be used directly in `defineGraph` without an `EdgeRegistration` wrapper. Constrained edges use their built-in `from`/`to`; unconstrained edges allow all node types: ```typescript const graph = defineGraph({ nodes: { Person: { type: Person }, Company: { type: Company } }, edges: { worksAt, // Constrained - uses built-in from/to sameAs, // Unconstrained - connects any node to any node }, }); ``` You can still use `EdgeRegistration` to narrow (but not widen) the constraints: ```typescript const worksAt = defineEdge("worksAt", { from: [Person], to: [Company, Subsidiary], // Allows both Company and Subsidiary }); const graph = defineGraph({ edges: { // Narrow to only Subsidiary targets in this graph worksAt: { type: worksAt, from: [Person], to: [Subsidiary] }, }, }); ``` Attempting to widen beyond the edge's built-in constraints throws a `ConfigurationError`: ```typescript const worksAt = defineEdge("worksAt", { from: [Person], to: [Company], }); // This throws ConfigurationError - OtherEntity is not in the edge's range defineGraph({ edges: { worksAt: { type: worksAt, from: [Person], to: [OtherEntity] }, }, }); ``` #### Source-Dependent Targets An array-valued `to` allows every combination of the source and target types. When the valid target depends on the source, use a map instead: ```typescript const Employee = defineNode("Employee", { schema: z.object({ name: z.string() }) }); const Student = defineNode("Student", { schema: z.object({ name: z.string() }) }); const Department = defineNode("Department", { schema: z.object({ name: z.string() }) }); const Course = defineNode("Course", { schema: z.object({ name: z.string() }) }); const assignedTo = defineEdge("assignedTo", { from: [Employee, Student], to: { Employee: [Department], Student: [Course], }, }); const graph = defineGraph({ id: "assignments", nodes: { Employee: { type: Employee }, Student: { type: Student }, Department: { type: Department }, Course: { type: Course }, }, edges: { assignedTo }, }); ``` This permits `Employee → Department` and `Student → Course`. It rejects `Employee → Course` and `Student → Department`. Using `to: [Department, Course]` would permit all four combinations. Map keys are the literal node kind names (`Employee.kind`), not aliases used to register nodes in a graph. Every kind in `from` must have a map entry, no other keys are allowed, and each target array must be nonempty. You can also use computed keys such as `[Employee.kind]: [Department]`. The map syntax works in an explicit graph registration too: ```typescript edges: { assignedTo: { type: assignedTo, from: [Employee], to: { Employee: [Department] }, }, } ``` A registration may narrow the built-in allowed pairs, but it cannot introduce new pairs. Replacing a map with arrays is valid only when every resulting combination is already allowed by the edge definition. At runtime, both endpoints must match the **same** declared pair, including `subClassOf` assignability. A source matching several source entries can use the targets allowed by any of those entries. An undeclared pair fails with [`EndpointPairError`](/errors#endpointpairerror); an invalid source kind still fails with `EndpointError`. Malformed declarations fail with `ConfigurationError`. Typed collection writes preserve the source/target relationship; dynamic writes and imports enforce it at runtime. Bulk writes reject invalid pairs atomically. Import pair validation remains active even when reference validation is disabled; imports retain their own documented error-handling and partial-success behavior. See [collection types](/types#typededgecollectionr) for inference limits, [graph extensions](/graph-extensions#edges) for runtime declarations, and [schema management](/schema-management#endpoint-pair-changes) for schema changes. ### Edge Constraints #### Cardinality Control how many edges can exist: ```typescript const graph = defineGraph({ edges: { // Default: no limit knows: { type: knows, from: [Person], to: [Person], cardinality: "many" }, // At most one edge of this type from any source node currentEmployer: { type: currentEmployer, from: [Person], to: [Company], cardinality: "one", }, // At most one edge between any (source, target) pair rated: { type: rated, from: [Person], to: [Product], cardinality: "unique" }, // At most one active edge (valid_to IS NULL) from any source currentRole: { type: currentRole, from: [Person], to: [Company], cardinality: "oneActive", }, }, }); ``` | Cardinality | Description | |-------------|-------------| | `"many"` | No limit (default) | | `"one"` | At most one edge of this type from any source node | | `"unique"` | At most one edge between any (source, target) pair | | `"oneActive"` | At most one edge with `valid_to IS NULL` from any source | #### Enforcement Timing Cardinality constraints are checked at edge **creation time**, before the insert: ```typescript // With cardinality: "one" on currentEmployer: await store.edges.currentEmployer.create(alice, acme, {}); // OK await store.edges.currentEmployer.create(alice, other, {}); // Throws CardinalityError ``` The check queries existing edges and throws `CardinalityError` if violated. For `oneActive`, only edges with `validTo` unset count toward the limit. ### Edge Operations ```typescript // Create edge - pass nodes directly const edge = await store.edges.worksAt.create(alice, acme, { role: "Engineer" }); // Retrieve edge const e = await store.edges.worksAt.getById(edge.id); // Delete edge await store.edges.worksAt.delete(edge.id); ``` ## Graph Definition The graph definition combines all components: ```typescript import { defineGraph } from "@nicia-ai/typegraph"; const graph = defineGraph({ // Unique identifier for this graph id: "my_application", // Node registrations nodes: { Person: { type: Person, onDelete: "restrict", // Default behavior }, Company: { type: Company, onDelete: "cascade", }, Employment: { type: Employment, onDelete: "disconnect", }, }, // Edge registrations edges: { worksAt: { type: worksAt, from: [Person], to: [Company], cardinality: "many", }, employedAt: { type: employedAt, from: [Company], to: [Employment], cardinality: "many", }, }, // Semantic relationships ontology: [subClassOf(Company, Organization), disjointWith(Person, Company)], }); ``` ## Delete Behaviors Control what happens when nodes are deleted: ### Restrict (Default) Blocks deletion if any edges are connected: ```typescript nodes: { Author: { type: Author }, // onDelete defaults to "restrict" } // This throws RestrictedDeleteError if Author has edges await store.nodes.Author.delete(authorId); ``` ### Cascade Automatically deletes all connected edges: ```typescript nodes: { Book: { type: Book, onDelete: "cascade" }, } // Deletes the book and all edges connected to it await store.nodes.Book.delete(bookId); ``` ### Disconnect Soft-deletes edges (preserves history): ```typescript nodes: { Review: { type: Review, onDelete: "disconnect" }, } // Marks connected edges as deleted (deleted_at is set) await store.nodes.Review.delete(reviewId); ``` ## Uniqueness Constraints Ensure unique values within node types: ```typescript const graph = defineGraph({ nodes: { Person: { type: Person, unique: [ { name: "person_email", fields: ["email"], where: (props) => props.email.isNotNull(), scope: "kind", collation: "caseInsensitive", }, ], }, Company: { type: Company, unique: [ { name: "company_ticker", fields: ["ticker"], scope: "kind", collation: "binary", }, ], }, }, }); ``` ### Scope Options - `"kind"`: Unique within this exact type only - `"kindWithSubClasses"`: Unique across this type and all subclasses ### Collation Options - `"binary"`: Case-sensitive comparison - `"caseInsensitive"`: Case-insensitive comparison ## Type Inference TypeGraph infers TypeScript types from Zod schemas—you never duplicate type definitions. ### Extracting Types from Definitions ```typescript import { z } from "zod"; import { defineNode, type Node, type NodeProps, type NodeId } from "@nicia-ai/typegraph"; const Person = defineNode("Person", { schema: z.object({ name: z.string(), email: z.string().email().optional(), age: z.number().optional(), }), }); // For functions that work with full nodes (id, kind, metadata, props): type PersonNode = Node; // { id: NodeId; kind: "Person"; name: string; email?: string; version: number; createdAt: Date; ... } // For functions that only need the property data: type PersonProps = NodeProps; // { name: string; email?: string; age?: number } // For type-safe node IDs (prevents mixing IDs from different node types): type PersonId = NodeId; // string & { readonly [__nodeId]: typeof Person } ``` Use `Node` when your function needs the full node with metadata. Use `NodeProps` when you only care about the schema properties (e.g., for form validation or API payloads). ### Typed Store Operations ```typescript // Create returns a fully typed Node const alice: Node = await store.nodes.Person.create({ name: "Alice", email: "alice@example.com", }); // TypeScript knows the structure alice.id; // NodeId - branded string alice.name; // string alice.email; // string | undefined alice.age; // number | undefined alice.version; // number alice.createdAt; // Date // Type errors caught at compile time await store.nodes.Person.create({ name: 123, // Error: Type 'number' is not assignable to type 'string' invalid: "field", // Error: Object literal may only specify known properties }); ``` ### Typed Query Results ```typescript // Result type is inferred from your select projection const results = await store .query() .from("Person", "p") .select((ctx) => ({ name: ctx.p.name, // TypeScript knows: string email: ctx.p.email, // TypeScript knows: string | undefined id: ctx.p.id, // TypeScript knows: NodeId })) .execute(); // results: Array<{ name: string; email: string | undefined; id: NodeId }> // Invalid property access is caught .select((ctx) => ({ invalid: ctx.p.nonexistent, // TypeScript error! })) ``` ### Typed Edge Operations Edge endpoints are constrained to valid node types: ```typescript // Edge definition: worksAt goes from Person → Company const graph = defineGraph({ // ... edges: { worksAt: { type: worksAt, from: [Person], to: [Company] }, }, }); // TypeScript enforces valid endpoints await store.edges.worksAt.create(alice, acmeCorp, { role: "Engineer" }); // OK await store.edges.worksAt.create(acmeCorp, alice, { role: "Engineer" }); // Error: Argument of type 'Node' is not assignable to parameter of type 'Node' ``` # Backend Setup > Configure SQLite and PostgreSQL backends for TypeGraph TypeGraph stores graph data in your existing relational database using Drizzle ORM adapters. This guide covers setting up SQLite, PostgreSQL, and PGlite backends. :::note[Custom indexes] TypeGraph migrations create the core tables and built-in indexes. For application-specific indexes on JSON properties (and Drizzle/drizzle-kit integration), see [Indexes](/performance/indexes). ::: ## SQLite SQLite is ideal for development, testing, single-server deployments, and embedded applications. ### Quick Setup For development and testing, use the convenience function that owns the connection and provisions TypeGraph's base tables: ```typescript import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; import { createStore } from "@nicia-ai/typegraph"; // In-memory database (resets on restart) const { backend } = createLocalSqliteBackend(); const store = createStore(graph, backend); // File-based database (persisted) const { backend, db } = createLocalSqliteBackend({ path: "./app.db" }); const store = createStore(graph, backend); ``` The local backend owns its connection, so it applies performance pragmas at open: `journal_mode=WAL`, `synchronous=NORMAL`, and a 5s `busy_timeout`. On file databases this makes single-operation writes roughly 5× faster than the driver defaults (rollback journal, `synchronous=FULL`). Override individual values or opt out entirely: ```typescript // Override one value, keep the other defaults createLocalSqliteBackend({ path: "./app.db", pragmas: { busyTimeoutMs: 10_000 } }); // Keep better-sqlite3's driver defaults untouched createLocalSqliteBackend({ path: "./app.db", pragmas: false }); ``` :::caution[Fulltext and embeddings require `createStoreWithSchema`] `createLocalSqliteBackend` creates the base tables but does not durably materialize strategy-owned storage. If your graph has `searchable()` or `embedding()` fields, boot with `const [store] = await createStoreWithSchema(graph, backend);` instead of bare `createStore()` — otherwise the first fulltext or embedding operation throws `StoreNotInitializedError`. ::: ### Manual Setup For full control over the database connection: ```typescript import Database from "better-sqlite3"; import { drizzle } from "drizzle-orm/better-sqlite3"; import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; import { createStoreWithSchema } from "@nicia-ai/typegraph"; // Create and configure the database const sqlite = new Database("app.db"); sqlite.pragma("journal_mode = WAL"); // Recommended for performance sqlite.pragma("foreign_keys = ON"); // Create Drizzle instance and backend const db = drizzle(sqlite); const backend = createSqliteBackend(db); // createStoreWithSchema auto-creates tables on first run const [store] = await createStoreWithSchema(graph, backend); // Clean up when done process.on("exit", () => sqlite.close()); ``` For a fresh database whose DDL is managed externally, use `generateSqliteMigrationSQL()` with `createStore()` instead: ```typescript sqlite.exec(generateSqliteMigrationSQL()); const store = createStore(graph, backend); ``` The generated script is complete installation DDL and stamps the current deployment-wide base-schema marker last; it is not an incremental upgrade planner. Existing databases attached only through the zero-DDL runtime factories must apply release-specific additive migrations through their migration tool. See [Upgrading deployment-wide base storage](#upgrading-deployment-wide-base-storage) for the exact SQLite and PostgreSQL statements. A privileged `createStoreWithSchema()` open adopts missing release storage once, then stamps a deployment-wide base-schema marker. Warm opens read that marker and issue no base-adoption DDL. ### SQLite with Vector Search For semantic search, use the sqlite-vec extension. `createLocalSqliteBackend()` wires the `sqliteVecStrategy` automatically when the extension loads. For a bring-your-own connection, load the extension and pass the strategy explicitly: ```typescript import Database from "better-sqlite3"; import { drizzle } from "drizzle-orm/better-sqlite3"; import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; import { sqliteVecStrategy } from "@nicia-ai/typegraph"; const sqlite = new Database("app.db"); // Load sqlite-vec extension sqlite.loadExtension("vec0"); // Run migrations (core tables) sqlite.exec(generateSqliteMigrationSQL()); const db = drizzle(sqlite); const backend = createSqliteBackend(db, { vector: sqliteVecStrategy }); ``` sqlite-vec stores embeddings in `vec0` virtual tables and supports the `cosine` and `l2` metrics. Per-field vector tables are provisioned by `createStoreWithSchema` at boot (not by the generated migration SQL), and the runtime asserts a durable marker rather than issuing DDL on first write — see [Database roles & least privilege](#database-roles--least-privilege). See [Semantic Search](/semantic-search) for query examples. ### libsql / Turso For edge deployments, shared-driver setups, or Turso cloud databases, use the first-class libsql backend: ```bash npm install @libsql/client ``` ```typescript import { createClient } from "@libsql/client"; import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql"; import { createStore } from "@nicia-ai/typegraph"; // Local file const client = createClient({ url: "file:app.db" }); // Or remote Turso database // const client = createClient({ url: "libsql://my-db.turso.io", authToken: "..." }); const { backend, db } = await createLibsqlBackend(client); const store = createStore(graph, backend); ``` `createLibsqlBackend` handles DDL execution and configures the correct async execution profile automatically. It returns both the `backend` and the underlying Drizzle `db` instance for direct SQL access. The caller retains ownership of the client and is responsible for closing it when done — this allows sharing a single client across TypeGraph and other libraries. Its installation is complete: the factory publishes the deployment-wide base-schema marker, and when it encounters a pre-0.52 edge table it applies the focused match-identity storage adoption before retrying the idempotent installation script. The local SQLite factory has the same behavior. The libsql backend has native vector and hybrid search, wired automatically via `libsqlVectorStrategy` — no extension to load. It uses libSQL's built-in engine (`F32_BLOB(N)` storage, `vector_distance_cos` / `vector_distance_l2`, and DiskANN approximate nearest neighbor via `libsql_vector_idx` + `vector_top_k`) and supports the `cosine` and `l2` metrics. See [Semantic Search](/semantic-search) for query examples. :::caution[In-memory databases and transactions] libsql's `file::memory:` creates a separate database per connection. Since transactions open a new connection, the original database is destroyed after a transaction completes ([tursodatabase/libsql-client-ts#229](https://github.com/tursodatabase/libsql-client-ts/issues/229)). Use a file-based database (`file:path.db`) or remote URL when transactions are needed. ::: ### API Reference #### `createLocalSqliteBackend(options?)` Creates a SQLite backend with automatic database and schema setup. ```typescript function createLocalSqliteBackend(options?: { path?: string; // Database path, defaults to ":memory:" tables?: SqliteTables; /** * Override the fulltext strategy. Defaults to `fts5Strategy` (SQLite's * built-in FTS5 virtual table). Pass `false` to disable fulltext support * entirely — the backend then advertises no `capabilities.fulltext` and * omits the fulltext CRUD/search methods, and the managed installation * never creates the fulltext table. Forwarded to both the installation * DDL and `createSqliteBackend`. */ fulltext?: FulltextStrategy | false; }): { backend: GraphBackend; db: BetterSQLite3Database }; ``` #### `createSqliteBackend(db, options?)` Creates a SQLite backend from an existing Drizzle database instance. Pass `vector` to enable vector search (for example `sqliteVecStrategy` after loading the sqlite-vec extension). ```typescript function createSqliteBackend( db: BetterSQLite3Database, options?: { tables?: SqliteTables; /** * Override the fulltext strategy. Defaults to `fts5Strategy` (SQLite's * built-in FTS5 virtual table). Pass `false` to disable fulltext * support entirely — the backend then advertises no * `capabilities.fulltext` and omits the fulltext CRUD/search methods, * mirroring `vector` left unset. Required for a SQLite build without * FTS5 compiled in. */ fulltext?: FulltextStrategy | false; vector?: VectorStrategy; capabilities?: BundledBackendCapabilityOverrides; }, ): GraphBackend; ``` Pass `{ fulltext: false }` on a SQLite build without FTS5 compiled in, or whenever the graph has no `searchable()` fields and you would rather skip the virtual table than carry it unused: ```typescript const backend = createSqliteBackend(db, { fulltext: false }); ``` #### `generateSqliteMigrationSQL()` Returns complete fresh-installation SQL for creating TypeGraph tables and stamping the current deployment-wide base-schema marker in SQLite. ```typescript function generateSqliteMigrationSQL( tables?: SqliteTables, fulltextStrategy?: FulltextStrategy | false, ): string; ``` `generateSqliteDDL()` is the lower-level table/index statement array used by backend bootstrap. It deliberately omits the deployment-wide marker row and is therefore not a complete installation script. Use `generateSqliteMigrationSQL()` when the resulting database will be opened through `createVerifiedStore()` or the DML-only graph-template APIs. #### `createLibsqlBackend(client, options?)` Creates a SQLite backend from a `@libsql/client` instance. Runs DDL automatically. The caller retains ownership of the client and is responsible for closing it. ```typescript async function createLibsqlBackend(client: Client, options?: { tables?: SqliteTables }): Promise<{ backend: GraphBackend; db: LibSQLDatabase }>; ``` ## PostgreSQL PostgreSQL is recommended for production deployments with concurrent access, large datasets, or when you need advanced features like pgvector. `createPostgresBackend` is driver-agnostic. Pick the Drizzle adapter that matches your runtime, and TypeGraph works the same way against each. ### Choosing a PostgreSQL driver | Runtime | Recommended driver | Drizzle adapter | | --------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | -------------------------------------------------------- | | Long-lived Node server (Fly, Render, Cloud Run, containers) | `pg` (node-postgres) or `postgres` (postgres-js) | `drizzle-orm/node-postgres` or `drizzle-orm/postgres-js` | | Node serverless (Vercel Functions, AWS Lambda, Netlify Functions) | `postgres` (postgres-js) — faster cold start, lower per-query overhead | `drizzle-orm/postgres-js` | | Bun server | `postgres` (postgres-js) or Bun's built-in SQL | `drizzle-orm/postgres-js` or `drizzle-orm/bun-sql` | | Edge runtime (Cloudflare Workers, Vercel Edge, Netlify Edge) — needs transactions | `@neondatabase/serverless` Pool over WebSockets | `drizzle-orm/neon-serverless` | | Edge runtime — single-statement reads/writes only | `@neondatabase/serverless` `neon(url)` over HTTP | `drizzle-orm/neon-http` | | Cloudflare Hyperdrive | `pg` or `postgres` (through the Hyperdrive pooler) | `drizzle-orm/node-postgres` or `drizzle-orm/postgres-js` | | Embedded apps, local development, Postgres dialect tests | `@electric-sql/pglite` | `drizzle-orm/pglite` | :::note[Neon HTTP vs WebSocket] Both Neon drivers work with TypeGraph. They have different tradeoffs: - **`drizzle-orm/neon-http`** uses HTTP per statement. Lowest cold-start cost; survives Workers' per-request isolation. **Cannot hold a session across statements**, so multi-statement transactions are unavailable — TypeGraph auto-detects this driver and sets `capabilities.execution.interactiveTransactions = false`, so `store.transaction(...)` refuses rather than pretending to provide rollback. Eligible atomic-batch operations remain available when the transport is certified for them. A schema-managed Store's write fuses its schema fence into the write's own statement when the write fuses, and fails closed otherwise — see [The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for which writes fuse and the reasons a write that cannot refuses with. - **`drizzle-orm/neon-serverless`** uses a WebSocket Pool. Holds a session, supports full transactional semantics, but the WebSocket connection lifecycle needs care in serverless / per-request contexts (you typically want a fresh Pool per request). Pick HTTP for stateless reads and for the fused schema-managed writes. Pick WebSockets for schema migrations, and for any write outside that fused envelope. ::: ### node-postgres (pg) The default choice for long-lived Node servers. Widest ecosystem and most deployment documentation. ```typescript import { Pool } from "pg"; import { drizzle } from "drizzle-orm/node-postgres"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { createStoreWithSchema } from "@nicia-ai/typegraph"; const pool = new Pool({ connectionString: process.env.DATABASE_URL, max: 20, }); const db = drizzle(pool); const backend = createPostgresBackend(db); const [store] = await createStoreWithSchema(graph, backend); ``` For a fresh database managed externally, use `generatePostgresMigrationSQL()` with `createStore()`: ```typescript import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; await pool.query(generatePostgresMigrationSQL()); const store = createStore(graph, backend); ``` As with SQLite, this is complete installation DDL rather than an incremental upgrade plan. Apply the [base-schema upgrade](#upgrading-deployment-wide-base-storage) to an existing database, or let a privileged `createStoreWithSchema()` preparation adopt the storage before runtime workers use `createStore()`. ### postgres-js A leaner Postgres client with lower per-query overhead and smaller bundle size. Good default for Node serverless platforms and Bun. Fully tested against TypeGraph's adapter and integration suites. ```bash npm install postgres drizzle-orm ``` ```typescript import postgres from "postgres"; import { drizzle } from "drizzle-orm/postgres-js"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { createStoreWithSchema } from "@nicia-ai/typegraph"; const sql = postgres(process.env.DATABASE_URL, { max: 10, idle_timeout: 30, }); const db = drizzle(sql); const backend = createPostgresBackend(db); const [store] = await createStoreWithSchema(graph, backend); ``` Transactions go through `sql.begin(fn)`; TypeGraph handles this automatically via Drizzle's `db.transaction()`. Isolation levels are honored the same way as with node-postgres. ### Neon serverless (WebSockets) For edge runtimes like Cloudflare Workers, Vercel Edge, and Netlify Edge — anywhere native TCP sockets aren't available. Neon's `@neondatabase/serverless` driver speaks the Postgres wire protocol over WebSockets and exposes a pg-Pool-compatible API. ```bash npm install @neondatabase/serverless drizzle-orm ``` ```typescript import { Pool } from "@neondatabase/serverless"; import { drizzle } from "drizzle-orm/neon-serverless"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { createStoreWithSchema } from "@nicia-ai/typegraph"; const pool = new Pool({ connectionString: env.NEON_DATABASE_URL }); const db = drizzle(pool); const backend = createPostgresBackend(db); const [store] = await createStoreWithSchema(graph, backend); ``` When running under Node.js (for local testing), install `ws` and configure it once before connecting: ```typescript import { neonConfig } from "@neondatabase/serverless"; import ws from "ws"; neonConfig.webSocketConstructor = ws; ``` Edge runtimes expose `WebSocket` globally and need no extra setup. ### Neon HTTP For stateless edge workloads where you don't need transactional writes. The HTTP driver issues one request per query — lowest cold-start cost, no session lifecycle to manage. TypeGraph auto-detects this driver and sets `capabilities.execution.interactiveTransactions` to `false` and `capabilities.execution.unitOfWork` to `"batch"`. On a raw Store, `store.transaction(...)` refuses rather than silently falling through to sequential execution. A schema-managed or verified Store's first write does not universally fail closed here — it depends on whether the write fuses. A singleton node create, update, `upsertById`, or delete fuses on a kind with no declared unique constraint (a create takes a generated or a caller-supplied id) — except a node delete, which fuses even when the kind DOES carry a declared unique constraint, because the atomic delete program releases that claim in the same statement. A singleton edge create fuses when the kind's cardinality is `"many"`, and edge update and delete fuse the same way (`EdgeCollection` has no `upsertById`). So do `bulkInsert`/`bulkCreate`/`bulkDelete`/`bulkReplaceById`/`bulkUpsertById`, and a constrained write inside an atomic program's claim envelope. Each of these asserts the active schema version inside the statements neon-http submits together, and `transaction(queries)` commits or rejects that submission as a whole. A write that cannot fuse either fails closed with `BATCH_WRITE_UNSUPPORTED` naming a proven reason (an interactive callback, a probe-then-write constraint check, Operational Identity, history, or a schema commit), or — for a write that simply doesn't fit the fused shape, such as a singleton create, update, or `upsertById` on a uniquely-constrained kind, or a supplied-id tombstone resurrection — fails closed with the plain `SCHEMA_WRITE_FENCE_UNSUPPORTED` limitation and no named reason. See [The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for the shared guard and the full reason table. Schema commits stay refused regardless: `commitSchemaVersion` and `setActiveVersion` require holding one transaction across their compare-and-swap read and activating write to eliminate the orphan-row crash window they exist to fix, so they refuse with a typed `ConfigurationError` on non-transactional backends. Run schema migrations from a process with a transactional driver (`drizzle-orm/neon-serverless`, regular `pg`, etc.); the edge worker can keep using neon-http for reads and for the fused writes above. A raw `createStore()` remains available for writes outside that envelope when the application explicitly accepts they are not fenced against schema changes. ```bash npm install @neondatabase/serverless drizzle-orm ``` ```typescript import { neon } from "@neondatabase/serverless"; import { drizzle } from "drizzle-orm/neon-http"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { createStore } from "@nicia-ai/typegraph"; const sql = neon(env.NEON_DATABASE_URL); const db = drizzle({ client: sql }); const backend = createPostgresBackend(db); const store = createStore(graph, backend); // backend.capabilities.execution.interactiveTransactions === false (auto-detected) ``` Use `neon-http` for reads and for the fused schema-managed writes listed above. Run schema migrations, and any write outside that envelope, through `neon-serverless`, regular `pg`, or another transactional driver. ### PGlite (Postgres-in-WASM) [PGlite](https://pglite.dev/) is a full Postgres compiled to WebAssembly that runs in-process — in Node, Bun, Deno, or the browser — with no server and no native addon. It's ideal for local development, embedded apps, and running the real Postgres dialect (including pgvector) in tests without Docker. `@electric-sql/pglite` is an optional peer dependency. Vector support additionally needs `@electric-sql/pglite-pgvector` (PGlite ≥ 0.5 ships pgvector as a separate package): ```bash npm install @electric-sql/pglite @electric-sql/pglite-pgvector ``` The batteries-included helper constructs the engine, loads pgvector, runs the schema DDL, and returns a ready backend — the Postgres analog of `createLocalSqliteBackend`: ```typescript import { createLocalPgliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres/pglite"; import { createStore } from "@nicia-ai/typegraph"; // In-memory by default, with pgvector enabled. const { backend, db, client } = await createLocalPgliteBackend(); const store = createStore(graph, backend); // backend.close() disposes the PGlite engine. ``` ```typescript // Persistent on disk: const { backend } = await createLocalPgliteBackend({ dataDir: "./pgdata" }); // No embeddings? Skip the extension (no pgvector dependency needed): const { backend } = await createLocalPgliteBackend({ vector: false }); // Pass an explicit pgvector extension object: import { vector } from "@electric-sql/pglite-pgvector"; const { backend } = await createLocalPgliteBackend({ vector }); ``` If you construct PGlite yourself, pass its Drizzle database straight to `createPostgresBackend` — the execution fast path detects PGlite and routes it correctly: ```typescript import { PGlite } from "@electric-sql/pglite"; import { vector } from "@electric-sql/pglite-pgvector"; import { drizzle } from "drizzle-orm/pglite"; import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; const client = await PGlite.create({ extensions: { vector } }); await client.exec(generatePostgresMigrationSQL()); const backend = createPostgresBackend(drizzle(client)); ``` PGlite is single-connection and serial: there is no pooling, so concurrent `store.transaction()` calls queue rather than run in parallel. It complements, rather than replaces, a Docker-based Postgres for CI — PGlite exercises the SQL dialect and pgvector, but not driver-specific behavior (node-postgres statement naming, postgres-js, pgbouncer, real concurrency). ### PostgreSQL with Vector Search For semantic search, enable pgvector. `createPostgresBackend` defaults to `pgvectorStrategy`, so no extra wiring is required: ```typescript import { Pool } from "pg"; import { drizzle } from "drizzle-orm/node-postgres"; import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; const pool = new Pool({ connectionString: process.env.DATABASE_URL }); // Migration SQL enables the pgvector extension await pool.query(generatePostgresMigrationSQL()); // Runs: CREATE EXTENSION IF NOT EXISTS vector; const db = drizzle(pool); const backend = createPostgresBackend(db); ``` pgvector stores embeddings in per-field typed `vector(N)` tables (provisioned by `createStoreWithSchema` at boot — the generated migration SQL creates no embedding table) with HNSW or IVFFlat indexes, and supports the `cosine`, `l2`, and `inner_product` metrics. See [Semantic Search](/semantic-search) for query examples. ### Refreshing planner statistics after bulk loads `importGraph()` refreshes planner statistics automatically after an import that created or updated rows, and `store.materializeIndexes()` does the same on SQLite after creating indexes (pass `refreshStatistics: false` to opt out). On PostgreSQL, `materializeIndexes()` builds with `CREATE INDEX CONCURRENTLY` and skips the automatic refresh — call `store.refreshStatistics()` after materializing. `bulkCreate` and `bulkInsert` on nodes and edges also refresh automatically when a single autocommit call writes 1,000 rows or more. Tune or disable this with the `autoRefreshStatistics` store option: ```typescript // Refresh after any autocommit bulkCreate of 5,000+ rows const store = createStore(graph, backend, { autoRefreshStatistics: 5000 }); // Never refresh automatically after bulkCreate const store = createStore(graph, backend, { autoRefreshStatistics: false }); ``` Bulk writes inside a `store.transaction(...)` block never auto-refresh — statistics collected mid-transaction cannot see the uncommitted rows — so refresh manually after the transaction commits. The same applies to loops of small `bulkCreate` batches that never individually reach the threshold, and to backend-level batch inserts — the loop example below covers that pattern. PostgreSQL's query planner relies on table statistics to choose between multi-column indexes on `typegraph_edges` (forward vs reverse vs cardinality), and when those statistics are stale the planner can pick a reverse-index scan with a filter — turning a 0.5ms forward traversal into a 5ms one. SQLite's planner is similarly sensitive: without `sqlite_stat1` data, some FTS5 fulltext queries fall back to a plan that's roughly 30× slower. Autovacuum / background statistics collection will catch up eventually, but refreshing explicitly gives correct latencies immediately. ```typescript for (const batch of batches) { await store.nodes.Document.bulkCreate(batch); } await store.refreshStatistics(); ``` The implementation runs `ANALYZE` against the TypeGraph-managed tables in the configured backend — the call is safe regardless of custom table names or fulltext / embedding configuration. Cloudflare D1 and Durable Object SQLite reject the performance-only `PRAGMA analysis_limit` tuning statement through their authorizer. TypeGraph recognizes only that `SQLITE_AUTH` failure and continues with scoped `ANALYZE`; workerd permits `ANALYZE`, so planner statistics are still refreshed but without bounded sampling. Unexpected PRAGMA or ANALYZE failures stay visible through the existing caller warning or rejection. If you need to bypass the API for an unusual deployment (for example issuing `ANALYZE` over a separate admin connection), call `backend.execute()` with raw SQL as the escape hatch. ### pgbouncer / transaction-pool mode By default, the node-postgres / neon-serverless fast path issues server-side prepared statements (`client.query({name, text, values})`) so PostgreSQL caches the parsed plan per session. This is incompatible with pgbouncer in transaction-pool mode: pgbouncer routes successive statements over different backend connections, so a `name` registered on one connection isn't visible on the next. Pass `prepareStatements: false` to fall back to unnamed positional queries: ```typescript const backend = createPostgresBackend(db, { prepareStatements: false, // pgbouncer transaction-pool compatibility }); ``` The in-process cache that maps SQL text → statement name is LRU-bounded (default 256 entries, override via `preparedStatementCacheMax`). Eviction never recycles a name, because a live connection may still retain that name for its original SQL. Therefore this setting does not bound server-side prepared statement memory. For a high-cardinality stream of SQL text, use `prepareStatements: false` instead. ### Adopted schema transactions `store.withEvolvedTransaction(nativeTx, plan, callback, { waitBudgetMs })` requires an initialized adapter Store and a live caller-owned transaction. Plan the extension outside that transaction with `store.planEvolution()`. For change plans, the default exclusive schema-fence wait budget is 5,000 ms; a `SchemaFenceTimeoutError` requires rollback and retry of the complete application transaction. Omit `waitBudgetMs` for no-op plans, which use ordinary adoption without the exclusive fence and refuse that option. Interactive PostgreSQL adapters validate the active session, retain the existing schema advisory lock → schema row → recorded-write lock order, and use transaction-scoped advisory locks. This lock lifetime is suitable for transaction poolers such as Hyperdrive. The adapter restores temporary timeout settings before the callback. Noninteractive HTTP drivers cannot adopt schema transactions. SQLite schema adoption requires an active transaction on the backend's exact native connection with an observable `inTransaction` state, as provided by better-sqlite3. Drivers without that evidence refuse schema adoption; ordinary transaction support alone does not imply support for this operation. A deferred SQLite transaction acquires the writer slot before validating the schema plan. Adapters default to a DML-only schema provisioning policy. Plans requiring new vector slots or identity work refuse before taking a mutating fence, running DDL, or changing schema rows. A privileged adapter configured with `schemaProvisioning: "transactional"` can apply those plans: it revalidates storage on the pinned caller session and provisions identity relations, vector tables, and durable contribution markers inside that same transaction. The caller must roll back the entire native transaction if any step fails. ```typescript const backend = createPostgresBackend(db, { schemaProvisioning: "transactional", }); ``` Use a connection with permission to run the required DDL for this adapter; keep the default policy for a runtime role limited to DML. Bootstrap base storage before this request path; missing bootstrap tables refuse rather than being created lazily. Database permissions still determine whether transactional DDL succeeds. Generic eager index materialization, including concurrent PostgreSQL indexes, remains an explicit post-commit operation on the refreshed Store. Custom adapters must implement `adoptSchemaWriteTransaction` with the same session-bound fencing, finite-wait, and CAS guarantees to support change plans. See [Graph Extensions](/graph-extensions) for callback and receipt usage. ### Authoritative command sessions Store create paths use the backend's `commands` port for writes whose decision and mutation must share one command boundary. First-party paths pass an explicit command context: a root port owns any internal transaction it needs and cannot inherit caller coordination, while a transaction-scoped backend uses the active caller or Store transaction. A transaction command may additionally carry a coordination token only after it has acquired the graph's advisory lock; the token is bound to that graph and transaction session and cannot authorize work on another connection. On PostgreSQL, the lock statement also observes the effective transaction isolation and binds it to the same token. Match-key convergence therefore accepts only read committed or serializable based on database state, not the caller-requested option or the server's assumed default. `GraphBackend.commands` is a required member as of the authoritative command port release. Custom backends must expose `{ session, execute }` and implement the `node.create`, `edge.create`, and `edge.converge-create` commands, or return a typed `unsupported` result for dimensions they do not provide. The former optional managed-create and specialized edge-insert hooks are no longer a complete backend implementation; migrate those branches into the command port before upgrading. For a custom backend, the migration shape is: ```typescript const commands: GraphCommandPort = { session: "transaction", // use "root" for a single-statement backend execute(command, context) { // Apply every requested dimension, or explicitly refuse the command. switch (command.kind) { case "node.create": { return { outcome: "unsupported", entity: "node", dimensions: ["claims"] }; } case "edge.create": { return { outcome: "unsupported", entity: "edge", dimensions: ["endpointPredicate"], }; } case "edge.converge-create": { return { outcome: "unsupported", entity: "edge", dimensions: ["convergence"] }; } } }, }; const backend: GraphBackend = { ...members, commands }; ``` Every command port caller must provide the explicit context. TypeGraph-owned write paths use the command helper, which verifies that any coordination token belongs to the active graph and transaction session and carries a supported effective isolation before executing convergence. The portable PostgreSQL graph-lock path records that isolation automatically. A custom implementation of `lockSchemaVersionAndGraphWrite` must return the normalized `GraphCommandIsolation` observed by its combined lock statement. When decorating a first-party backend with `deriveBackend`, a same-session `commands` override retains the session identity. A wrapper that changes session or forwards to a different connection is a new command boundary and cannot reuse a token from the original port. These are four different execution guarantees; do not use “atomic” as a catch-all: - **Interactive transaction** (`store.transaction(...)`) pins one session and can make several Store operations commit or roll back together. The `runOptionallyInTransaction` callback receives `{ mode: "interactive-transaction" }` when this boundary was opened, or `{ mode: "sequential" }` on a backend without transaction support. - **Static internal adapter batch** is an adapter implementation detail (for example, a D1 batch or a bind-budgeted multi-row insert). It may make one precompiled set of statements atomic, but it is not a public Store transaction and does not make an arbitrary sequence of Store calls atomic. - **Certified atomic SQL program** is the backend-authoring transport seam for a closed, ordered sequence of statements. A backend earns this capability by passing the framework-agnostic conformance runner: result slots and bound parameters must be preserved, a failure in a later statement must leave no primary or sidecar writes, and an empty program must be a no-op. Certification is separate from semantic mutation eligibility; a transport alone does not authorize a mutation family. Bundled recognized PostgreSQL drivers provide this boundary either through Neon HTTP's transaction batch or a pinned interactive transaction; an unrecognized driver leaves it unavailable. - **Authoritative one-statement command** is the `commands.execute` port. A command returns a created/found/rejected/unsupported result after the database statement itself owns the decision and mutation. It is the transactionless path for eligible durable edge `matchIdentity` convergence; it is not a promise that every command or side effect can be fused. Operational Identity, single-edge claim/cardinality checks, and undeclared dynamic `matchOn` convergence remain interactive-transaction contracts. A custom or non-transactional backend must refuse those dimensions rather than silently falling through to a sequence of independent statements. Eligible direct edge batches on bundled roots are a narrower static-program contract: the insert and cardinality sidecars execute in one native atomic exchange. A declared durable edge `matchIdentity` is different for endpoint convergence: its canonical key has a database arbiter, so the eligible root create/found command can be authoritative in one statement. Backend implementations may also expose the optional `findEdgesByMatchIdentity` read capability for bounded merge planning. It must match the complete `(graphId, kind, name, key)` tuple and return tombstoned owners as well as active rows; omitting it keeps the portable full-clone path. Custom Drizzle operation strategies can opt in by supplying the corresponding owner-query builder. A strategy without that builder does not expose the capability, so callers can detect and retain the portable path. ### Connection Pooling For production, always use connection pooling: ```typescript import { Pool } from "pg"; const pool = new Pool({ connectionString: process.env.DATABASE_URL, max: 20, // Maximum pool size idleTimeoutMillis: 30000, // Close idle connections after 30s connectionTimeoutMillis: 2000, // Timeout for new connections }); // Handle pool errors pool.on("error", (err) => { console.error("Unexpected pool error", err); }); // Graceful shutdown process.on("SIGTERM", async () => { await pool.end(); process.exit(0); }); ``` ### API Reference #### `createPostgresBackend(db, options?)` Creates a PostgreSQL backend adapter. Accepts any Drizzle PostgreSQL database instance, regardless of the underlying driver. Tested with `drizzle-orm/node-postgres`, `drizzle-orm/postgres-js`, `drizzle-orm/neon-serverless`, `drizzle-orm/neon-http`, and `drizzle-orm/pglite`. The neon-http driver is auto-detected and `capabilities.execution.interactiveTransactions` is set to `false` (HTTP can't hold a session); use `drizzle-orm/neon-serverless` if you need transactional writes. ```typescript function createPostgresBackend( db: AnyPgDatabase, options?: { tables?: PostgresTables; /** * Override the fulltext strategy. Defaults to `tsvectorStrategy`. * Pass a custom `FulltextStrategy` to swap the fulltext stack, or * `false` to disable fulltext support entirely — the backend then * advertises no `capabilities.fulltext` and omits the fulltext * CRUD/search methods, mirroring `vector: false`. */ fulltext?: FulltextStrategy | false; /** * Override the vector search strategy. Defaults to * `pgvectorStrategy`. Pass a custom `VectorStrategy` to change the * storage / index engine, or `false` to disable vector support. */ vector?: VectorStrategy | false; /** * Override specific backend capabilities. Useful for HTTP-style * drivers or test scenarios. neon-http already has * `execution.interactiveTransactions: false` auto-applied — pass * this to override that or to disable other capabilities for custom * drivers. */ capabilities?: BundledBackendCapabilityOverrides; /** * Use server-side prepared statements on the node-postgres / * neon-serverless fast path. Default `true`. Set to `false` when * pooling through pgbouncer in transaction-pool mode (named * statements are invisible across pooled connections). */ prepareStatements?: boolean; /** * LRU cap on the number of distinct SQL strings tracked for * prepared-statement naming. Default 256. Worst-case server-side * footprint is roughly `cap × pool size` prepared statements. * Ignored when `prepareStatements` is `false`. */ preparedStatementCacheMax?: number; }, ): GraphBackend; ``` Pass `{ fulltext: false }` when the graph has no `searchable()` fields and you would rather skip the fulltext table (`typegraph_node_fulltext`) and its GIN index than carry them unused: ```typescript const backend = createPostgresBackend(db, { fulltext: false }); ``` #### `createPostgresTransactionBackend(tx, options?)` Creates a full backend on a Drizzle PostgreSQL transaction opened by the application. Use it when TypeGraph's tables share a transaction with other application tables, especially when TypeGraph uses prefixed table names. Pass the same `PostgresBackendOptions` as `createPostgresBackend`: ```typescript import { createPostgresTransactionBackend, createPostgresTables, } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; const graphTables = createPostgresTables({ nodes: "app_graph_nodes" }); await db.transaction(async (tx) => { const backend = createPostgresTransactionBackend(tx, { tables: graphTables, }); // Use the backend or a Store built from it within this callback. }); ``` The factory requires a transaction handle and serializes TypeGraph statements on its single pinned connection, including concurrent reads started by the same Store operation. Backends created for the same transaction handle share one queue. The application owns commit and rollback and must await all work using these backends before its transaction callback returns. `createPostgresBackend(tx)` also routes a PostgreSQL transaction handle to the transaction-scoped backend automatically. Use `createPostgresTransactionBackend` when you want the transaction-scoped intent to be explicit; a regular database handle passed to `createPostgresBackend(db)` still creates the pooled backend. #### `createLocalPgliteBackend(options?)` Creates an in-process PGlite backend with automatic engine construction, schema DDL, and optional pgvector loading. The returned backend owns the PGlite engine; call `backend.close()` when the process or test is done. ```typescript async function createLocalPgliteBackend(options?: { /** * PGlite data directory. Omit for an in-memory database, pass a filesystem * path for persistence, or use a runtime-specific scheme such as `idb://`. */ dataDir?: string; tables?: PostgresTables; /** * Omit to load @electric-sql/pglite-pgvector, pass `false` to disable vector * support, or pass a PGlite Extension object to control the extension import. */ vector?: false | Extension; /** * Override the fulltext strategy. Defaults to `tsvectorStrategy`. Pass * `false` to disable fulltext support entirely — the backend then * advertises no `capabilities.fulltext` and omits the fulltext CRUD/search * methods, and the installation DDL never creates the fulltext table. */ fulltext?: FulltextStrategy | false; }): Promise<{ backend: GraphBackend; db: PgliteDatabase; client: PGlite; }>; ``` #### `generatePostgresMigrationSQL()` Returns complete fresh-installation SQL for creating TypeGraph tables and stamping the current deployment-wide base-schema marker in PostgreSQL. It includes the pgvector extension. The vector-disabled local PGlite factory uses the same installation builder internally while omitting only that extension. ```typescript function generatePostgresMigrationSQL( tables?: PostgresTables, fulltextStrategy?: FulltextStrategy | false, ): string; ``` #### `generatePostgresDDL(tables?)` Returns individual DDL statements (CREATE TABLE, CREATE INDEX) as an array. Useful when you need per-statement control, for example to execute them in separate transactions or log them individually. This low-level array deliberately omits the deployment-wide marker row, so joining it does not produce a complete installation. Use `generatePostgresMigrationSQL()` for a database that will be opened through `createVerifiedStore()` or the DML-only graph-template APIs. ```typescript function generatePostgresDDL( tables?: PostgresTables, fulltextStrategy?: FulltextStrategy | false, ): string[]; ``` #### `generatePostgresDropSQL(tables?, fulltextStrategy?)` Returns one `DROP TABLE IF EXISTS` statement for the base and fulltext tables that `generatePostgresDDL()` would create. Use it to clean up an isolated, prefixed PostgreSQL table set after closing every backend connected to it. Pass the same tables and fulltext strategy used at installation. The statement does not use `CASCADE`: PostgreSQL refuses the drop if an application-owned object depends on one of these tables. It does not drop graph-scoped vector tables materialized later at runtime, so a working copy using those tables needs additional graph-scoped cleanup. ```typescript function generatePostgresDropSQL( tables?: PostgresTables, fulltextStrategy?: FulltextStrategy | false, ): string; ``` ### Upgrading deployment-wide base storage Skip this section when `createStoreWithSchema()` or `createAdapterStoreWithSchema()` owns schema preparation: the bundled SQLite and PostgreSQL adapters adopt each numbered base-schema release automatically on the first privileged open. No separate bootstrap command is needed. The deployment invariant is ordering: that privileged open must finish before any DML-only runtime worker starts. Base-schema version 1 includes the durable graph template relation and edge match-identity storage. It is required even for graphs without a `matchIdentity` declaration because every edge write names the two nullable columns. When database DDL is managed externally, apply the matching migration before a runtime worker opens the new graph schema. Apply the marker write last: it is the durable proof that every preceding step succeeded. The examples use the default TypeGraph table names. Replace every occurrence consistently when the adapter uses custom table names. For SQLite, run this migration exactly once. SQLite has no portable `ADD COLUMN IF NOT EXISTS`, so a migration tool must record whether it has already applied the two `ALTER TABLE` statements. Fresh and published schemas include the nullable-pair `CHECK` below. Privileged adoption accepts an externally managed table that already has both columns without that defensive constraint: SQLite does not expose structural CHECK metadata or support adding one without a full table rebuild, while TypeGraph writes always bind both values or neither. ```sql CREATE TABLE IF NOT EXISTS "typegraph_graph_templates" ( "template_id" TEXT PRIMARY KEY NOT NULL, "schema_hash" TEXT NOT NULL, "schema_doc" TEXT NOT NULL, "created_at" TEXT NOT NULL ); ALTER TABLE "typegraph_edges" ADD COLUMN "match_identity_name" TEXT; ALTER TABLE "typegraph_edges" ADD COLUMN "match_identity_key" TEXT CHECK (("match_identity_name" IS NULL) = ("match_identity_key" IS NULL)); CREATE UNIQUE INDEX IF NOT EXISTS "typegraph_edges_match_identity_uq" ON "typegraph_edges" ( "graph_id", "kind", "match_identity_name", "match_identity_key" ); CREATE TABLE IF NOT EXISTS "typegraph_base_schema_versions" ( "installation" INTEGER PRIMARY KEY NOT NULL, "version" INTEGER NOT NULL, "updated_at" TEXT NOT NULL, CONSTRAINT "typegraph_base_schema_versions_singleton_check" CHECK ("installation" = 1) ); INSERT INTO "typegraph_base_schema_versions" ("installation", "version", "updated_at") VALUES (1, 1, CURRENT_TIMESTAMP) ON CONFLICT ("installation") DO UPDATE SET "version" = excluded."version", "updated_at" = excluded."updated_at" WHERE "typegraph_base_schema_versions"."version" <= excluded."version"; ``` For PostgreSQL, the adoption statements are idempotent: ```sql CREATE TABLE IF NOT EXISTS "typegraph_graph_templates" ( "template_id" TEXT PRIMARY KEY NOT NULL, "schema_hash" TEXT NOT NULL, "schema_doc" JSONB NOT NULL, "created_at" TIMESTAMPTZ NOT NULL ); ALTER TABLE "typegraph_edges" ADD COLUMN IF NOT EXISTS "match_identity_name" TEXT; ALTER TABLE "typegraph_edges" ADD COLUMN IF NOT EXISTS "match_identity_key" TEXT; DO $$ BEGIN IF NOT EXISTS ( SELECT 1 FROM pg_constraint WHERE conrelid = to_regclass('"typegraph_edges"') AND conname = 'typegraph_edges_match_identity_pair_check' ) THEN ALTER TABLE "typegraph_edges" ADD CONSTRAINT "typegraph_edges_match_identity_pair_check" CHECK ( ("match_identity_name" IS NULL) = ("match_identity_key" IS NULL) ); END IF; END $$; CREATE UNIQUE INDEX IF NOT EXISTS "typegraph_edges_match_identity_uq" ON "typegraph_edges" ( "graph_id", "kind", "match_identity_name", "match_identity_key" ); CREATE TABLE IF NOT EXISTS "typegraph_base_schema_versions" ( "installation" INTEGER PRIMARY KEY NOT NULL, "version" INTEGER NOT NULL, "updated_at" TIMESTAMPTZ NOT NULL, CONSTRAINT "typegraph_base_schema_versions_singleton_check" CHECK ("installation" = 1) ); INSERT INTO "typegraph_base_schema_versions" ("installation", "version", "updated_at") VALUES (1, 1, NOW()) ON CONFLICT ("installation") DO UPDATE SET "version" = excluded."version", "updated_at" = excluded."updated_at" WHERE "typegraph_base_schema_versions"."version" <= excluded."version"; ``` The conditional update makes marker publication monotonic: replaying an older migration can never claim that storage prepared by a newer TypeGraph release is older. The fresh-installation generators use `DO NOTHING` instead because they are not upgrade planners; an existing stale marker remains stale until the numbered privileged adoption lifecycle runs. `createVerifiedStore`, `assertSchemaCurrent`, and the DML-only graph-template APIs read this marker and throw `BaseSchemaMigrationError` when it is missing, stale, or newer than the running library. They never attempt repair. A plain `createStore` remains a synchronous zero-I/O attach; if it reaches an edge write on legacy storage, the write fails with `ConfigurationError` and `details.code === "EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE"` rather than a raw missing-column error. Provisioning the columns does not authorize re-keying existing data. Adding, removing, renaming, or changing the fields of a declared `matchIdentity` remains a breaking graph-schema change while that edge kind has any physical rows, including tombstones. Export the affected edges, hard-delete them, publish the new schema, and import them again so every row receives a key under the new declaration. ### Base-schema version 5: byte-ordered `graph_id` indexes (PostgreSQL) Version 5 adds one index to each relation `listGraphIds` seeks (`nodes`, `edges` and `schema_versions`), ordering `graph_id` by bytes instead of by the database collation so that a page's cursor, prefix and limit bound the walk. SQLite already keeps text indexes in byte order, so its step only advances the marker. The index is a single-column `graph_id` index: PostgreSQL deduplicates the repeated values, so it stays small (about 7 MB beside a 97 MB `nodes` heap of one million rows) and adds 1 to 2% to writes on the relation it lands on (single creates and 1,000-row bulk writes alike). The privileged open builds the three indexes with a plain `CREATE INDEX`, which blocks writes to the table while it runs (about 0.1 second per million `nodes` rows on the measurement hardware). This happens inline at boot even when `systemIndexes: "skip"` is set: that option only defers system index materialization, not base-schema adoption. For a large deployment, build the indexes first with `CONCURRENTLY`; the adoption step is `IF NOT EXISTS` and then finds them in place. Run each statement outside a transaction, and never run the same concurrent build from two sessions at once. Use the adapter's table names throughout: ```sql CREATE INDEX CONCURRENTLY IF NOT EXISTS "typegraph_nodes_graph_id_bytes_idx" ON "typegraph_nodes" ("graph_id" COLLATE "C"); CREATE INDEX CONCURRENTLY IF NOT EXISTS "typegraph_edges_graph_id_bytes_idx" ON "typegraph_edges" ("graph_id" COLLATE "C"); CREATE INDEX CONCURRENTLY IF NOT EXISTS "typegraph_schema_versions_graph_id_bytes_idx" ON "typegraph_schema_versions" ("graph_id" COLLATE "C"); ``` `CREATE INDEX CONCURRENTLY` can leave an invalid index behind if it is interrupted; drop it and rerun. `listGraphIds` checks only that each index exists and is valid, not its definition, so an index you create by hand under one of these names with a different definition is trusted and makes the walk slow rather than wrong. Create them exactly as shown. Advancing the marker to 5 is a one-way step: a library release that predates version 5 refuses a database stamped 5, so roll forward rather than back once any process has adopted it. Externally managed DDL applies the same statements, then advances the marker to 5 with the monotonic `INSERT ... ON CONFLICT` shown above. Until the indexes exist `listGraphIds` still returns correct pages, by reading and de-duplicating the anchor relations instead of walking them. ## Drizzle-Free Entrypoints TypeGraph keeps its public core and backend contracts independent of Drizzle: - `@nicia-ai/typegraph/core` exports graph definition helpers and their schema-derived types for packages that only define or share schemas. - `@nicia-ai/typegraph/backend` exports the complete backend, dialect, SQL-fragment, fulltext, and vector strategy contracts for adapter authors. - `@nicia-ai/typegraph/sqlite/local` and `@nicia-ai/typegraph/postgres/pglite` create managed Stores without exposing adapter-native handles. Application code can continue importing the complete portable Store API from `@nicia-ai/typegraph`. Use the `/adapters/drizzle/...` entrypoints only when the application deliberately owns a Drizzle connection or needs native transaction interop. Custom insert builders must apply the same born-ended validity rule as the built-in adapters. Import its public owner instead of duplicating the bound comparison: ```typescript import { resolveStampedValidityLowerBound } from "@nicia-ai/typegraph/backend"; const validFrom = resolveStampedValidityLowerBound( params.validFrom, params.validTo, writeInstant, ); ``` Use the same `writeInstant` for the decision and the row's creation/update stamp. This keeps custom node and edge inserts, plus node resurrection paths that reset the validity window, aligned with Store and interchange semantics at the zero-width boundary. Edge resurrection retains its stored lower bound and does not use this stamping helper. ## Managed Store Entrypoints For local applications that do not need direct database access, TypeGraph can own the connection, provision its schema, and return the complete typed Store: - `@nicia-ai/typegraph/sqlite/local` — Node-only SQLite through the native better-sqlite3 addon - `@nicia-ai/typegraph/postgres/pglite` — in-process PostgreSQL through PGlite's WebAssembly runtime ```typescript import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local"; import { createLocalPgliteStore } from "@nicia-ai/typegraph/postgres/pglite"; const sqliteStore = await createLocalSqliteStore(graph, { path: "./graph.db" }); const postgresStore = await createLocalPgliteStore(graph, { vector: false }); ``` These entrypoints expose no adapter-native database handle. The returned `Store` keeps the complete graph API, including graph-owned `store.transaction(...)`, but intentionally omits `withTransaction` and `withRecordedTransaction`, which require a caller-owned adapter handle. The Store owns its connection, so call `store.close()` during shutdown. Its declaration surface is safe for strict TypeScript consumers that do not install unused database drivers. PGlite vector support is enabled by default and loads the optional `@electric-sql/pglite-pgvector` package. Install that package when using vector fields, or pass `{ vector: false }` as above for a smaller non-vector setup. Both managed entrypoints also accept `fulltext: false`, which skips the fulltext table at bootstrap and returns a backend with no `capabilities.fulltext`. Both factories accept `store` and `schemaManagement` groups, so the managed path supports the same hooks, history/revision tracking, custom SQL schema, query defaults, and migration policy as `createStoreWithSchema`: ```typescript import { createSqlSchema } from "@nicia-ai/typegraph"; const store = await createLocalSqliteStore(graph, { path: "./graph.db", pragmas: { busyTimeoutMs: 10_000 }, store: { history: true, schema: createSqlSchema({ nodes: "app_nodes", edges: "app_edges", fulltext: "app_fulltext", uniques: "app_uniques", }), }, schemaManagement: { systemIndexes: "skip" }, }); ``` When a custom SQL schema is supplied, the managed factory provisions those same physical table names; no separate Drizzle table configuration is needed. `drizzle-orm` is an optional peer dependency for these two managed entrypoints: they load it only when their factory is called and, when it is absent, reject with a typed `ConfigurationError` (`MISSING_PEER_DEPENDENCY`) naming the package and the install command (`npm install drizzle-orm`) rather than a bare module-resolution stack. The explicit `/adapters/drizzle/...` entrypoints below expose Drizzle-native backends, connections, or schema builders — or, for `/adapters/drizzle/engine`, the factory that assembles a backend from a caller-supplied engine profile — and load `drizzle-orm` when the module is evaluated. Importing one without the peer installed therefore surfaces the raw module-resolution error, which names the same package. ## Drizzle Adapter Entrypoints TypeGraph exposes Drizzle adapters through public entrypoints: - `@nicia-ai/typegraph/adapters/drizzle/indexes` — Drizzle schema-builder helpers for TypeGraph index declarations - `@nicia-ai/typegraph/adapters/drizzle/sqlite` — Generic SQLite adapter (any Drizzle SQLite driver) - `@nicia-ai/typegraph/adapters/drizzle/sqlite/local` — Batteries-included better-sqlite3 wrapper (Node.js only) - `@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql` — Batteries-included libsql wrapper (Node.js, Workers, browser) - `@nicia-ai/typegraph/adapters/drizzle/postgres` — PostgreSQL adapter (any Drizzle Postgres driver) - `@nicia-ai/typegraph/adapters/drizzle/postgres/working-copy` — PostgreSQL table-backed working-copy manager - `@nicia-ai/typegraph/adapters/drizzle/postgres/pglite` — Batteries-included PGlite (Postgres-in-WASM) wrapper - `@nicia-ai/typegraph/adapters/drizzle/engine` — `createSqlBackend`, `deriveEngineProfile`, the bundled builders, `SqlEngineProfile` Import from the entrypoint matching your database: ```typescript import { createSqliteBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql"; import { createPostgresBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { createLocalPgliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres/pglite"; ``` ### Engine profiles `createPostgresBackend` and `createSqliteBackend` are each `createSqlBackend` applied to a profile built by `buildPostgresEngineProfile` / `buildSqliteEngineProfile`, both exported alongside `createSqlBackend` and `deriveEngineProfile` from the engine entrypoint: ```typescript import { buildPostgresEngineProfile, createSqlBackend, } from "@nicia-ai/typegraph/adapters/drizzle/engine"; const backend = createSqlBackend(buildPostgresEngineProfile(db, options)); ``` Most callers adapting a bundled backend want `deriveEngineProfile`, which builds a variant of a bundled profile — a different lock spelling, a looser declared capability, a replaced resource-audit verdict — without hand-copying every other field. A profile written from scratch is not constructible today: the assembly constructor is unexported, and `createSqlBackend` refuses a hand-built assembly. See [Authoring an engine profile](/backend-authoring) for the derivable-field table, the refusals a custom profile can hit, a worked example, and [what is not derivable yet](/backend-authoring#what-is-not-derivable-yet). A profile owns everything that genuinely differs between engines: dialect tokens, the execution adapter, transaction framing, its `fenceSql` lock spelling (see [Write fence declaration](#write-fence-declaration-writefence)), provisioning DDL, strategies, and limits. `createSqlBackend` owns everything that is the same for every SQL engine: deriving the final capabilities, resolving the write-fence decision once, assembling the mirrored member groups, auditing the backend's resource shape, and applying the trust marks. A backend minted this way earns the marks its own declarations back: the schema-fenced-insert mark only when the resolved fence plan actually fences writers, the root-autocommit mark only when the profile declares single-statement durability, and the atomic-program registrations only when its capabilities support root atomic batching — whether the profile is for a PostgreSQL-wire engine with a different locking story or an embedded engine with a different transaction model. `createSqlBackend` refuses a profile whose resolved capabilities omit `writeFence` — every mark and registration it applies assumes a resolvable write-fence decision, and a profile that does not declare one cannot back that decision soundly. `buildPostgresEngineProfile` and `buildSqliteEngineProfile` are the reference profiles to read when modeling a new one. A profile's `provisioning.catalog` supplies the backend's optional `catalog` member: physical-schema introspection — table and index existence, each column's normalized type family and raw declared type (a `CatalogColumn` is `{ name, kind, declaredType }`; `declaredType` is required, and every custom `columnTypes` implementation must populate it), and this engine's index-build facts — for the handful of store paths that need to read the engine catalog directly instead of compiling a portable query: index materialization (`store.materializeIndexes()` refuses only once its empty-candidate short circuit and the status-table ensure step have already run; `store.materializeSystemIndexes()`, which has no candidate short circuit, refuses only once that same status-table ensure step has run), the recorded-time schema check, and the recorded-time migration's column read. A profile that leaves `catalog` unset builds a backend with no `catalog` member at all; those paths then refuse with a `ConfigurationError` naming `catalog` rather than guessing at engine-specific SQL. A dialect also declares `subgraphMembershipStrategy`, naming a decision the dialect adapter makes, not one a profile supplies directly — the dialect adapters are a fixed record keyed by `SqlDialect`, and each adapter's capabilities (`DialectCapabilities`) declares `subgraphMembershipStrategy`, so a profile inherits whichever of the two dialects its own `dialect` field names. It is the plan-shape choice behind `store.subgraph()`'s reachable-node filter: `"materialized-ids"` fetches the traversal closure once and filters both the node and edge queries against that fixed id list (the shape PostgreSQL uses, trading one extra round trip for a parameter-driven plan), while `"inline-cte"` embeds the recursive closure in each fetch instead (the shape SQLite uses, where an in-process traversal is cheap and a growing parameter list would pressure the bind budget). This is a control-flow and prepared-plan decision, not SQL text a shared token could express identically on both shapes, so it lives on `DialectCapabilities` rather than in the query compiler. `instantiateStatement` — a member of the profile's `graphTemplateRuntime` bag, and so one of the fields `deriveEngineProfile` can override — is a required builder cloning a durable schema template into a fresh graph. Given the template and target graph's ids and schema hashes (`InstantiateGraphTemplateSqlParams`: `templateId`, `templateSchemaHash`, `graphId`, `schemaHash`, and the three physical table names it reads), it must return the statement that inserts the target graph's `schema_versions` row from the template's stored document and copies the template's contribution-marker rows into the target graph — taking the target graph's write lock, the same key the schema-commit fence takes, co-atomically with the insert on an engine that fences with locks. An engine whose dialect can compose a data-modifying CTE beside the schema INSERT (PostgreSQL) folds the marker copy and the lock into that one statement; an engine that cannot (SQLite) instead supplies the optional `copyContributionMarkers` dep, which runs the marker copy as a second statement once the schema row is confirmed. The bundled `postgresInstantiateGraphTemplateStatement` and `sqliteInstantiateGraphTemplateStatement` builders (`graph-template-sql.ts`) are what `createPostgresBackend` and `createSqliteBackend` supply to their own profiles; neither is exported, so a from-scratch profile reaches the same shape only by copying a bundled profile and adapting its statement, while a derived profile can replace the whole `graphTemplateRuntime` bag through `deriveEngineProfile`. `FenceSql` (see [Write fence declaration](#write-fence-declaration-writefence)) declares `advisoryLockExpression` and `isolationFactExpression` as the two composable, no-`SELECT` forms an `advisory`-mechanism backend author supplies; TypeGraph derives the standalone-statement counterparts (`acquireKeyed`, `acquireKeyedWithIsolation`, `isolationFact`) from them. A statement that must compose a lock or an isolation read INSIDE a larger query it builds itself — a CTE, a data-modifying statement — embeds the bare expression directly, rather than running the derived standalone form as its own preceding statement. The schema write fence's fused schema + graph-write statement (`postgres-schema-write-fence.ts`) is the one site that needs this: it puts the expression in its own CTE's `SELECT ... AS "lock_token"`, resolving the fence target's OWN `FenceSql` — the bundled `postgresFenceSql` for a bundled backend, a derived profile's own override otherwise — so a custom spelling backs this fused statement exactly as it backs every ordinary lock site. The graph-template instantiation statement's `locked AS (SELECT ...)` CTE is a DIFFERENT case, not a `FenceSql` consumer at all: it composes the baked single-argument `advisoryLockSingleExpression` directly. The ONE lock form with no override point is `advisoryLockSingleExpression`, the ONE-argument form on a bare key: PostgreSQL stores it in a lock space distinct from every namespaced two-argument lock, and the schema-commit fence and graph-template instantiation both take it on the same key so the two mutually exclude. It is not a `FenceSql` member — both bundled builders bake it in directly, and a custom profile has no way to replace it. ## Cloudflare D1 TypeGraph supports Cloudflare D1 for edge deployments, with some limitations. Cloudflare D1 has no interactive transaction primitive, so it cannot commit TypeGraph schema versions: `commitSchemaVersion` / `setActiveVersion` need to hold one transaction across their compare-and-swap read and activating write, and D1 has no session to hold it on. Apply the base DDL with Wrangler / drizzle-kit. `capabilities.execution.unitOfWork` reports `"batch"`. A singleton node create, update, `upsertById`, or delete fuses on a kind with no declared unique constraint (a create takes a generated or a caller-supplied id) — except a node delete, which fuses even when the kind DOES carry a declared unique constraint, because the atomic delete program releases that claim in the same statement. A singleton edge create fuses when the kind's cardinality is `"many"`, and edge update and delete fuse the same way (`EdgeCollection` has no `upsertById`). So do `bulkInsert`/`bulkCreate`/`bulkDelete`/`bulkReplaceById`/`bulkUpsertById`, and a constrained write inside an atomic program's claim envelope. Each of these asserts the active schema version inside the statements `D1Database.batch()` runs together, so these succeed on a schema-managed Store. A write that cannot fuse either fails closed with `BATCH_WRITE_UNSUPPORTED` naming a proven reason (a probe-then-write constraint check, an interactive callback, Operational Identity, history, or a schema commit), or — for a write that simply doesn't fit the fused shape, such as a singleton create, update, or `upsertById` on a uniquely-constrained kind, or a supplied-id tombstone resurrection — fails closed with the plain `SCHEMA_WRITE_FENCE_UNSUPPORTED` limitation and no named reason; see [The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for the full reason table. Use a raw `createStore()` only when the application accepts unfenced writes for the remaining paths: ```typescript import { drizzle } from "drizzle-orm/d1"; import { createStore } from "@nicia-ai/typegraph"; import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; export default { async fetch(request: Request, env: Env) { const db = drizzle(env.DB); const backend = createSqliteBackend(db); const store = createStore(graph, backend); // Use store... }, }; ``` This raw Store does not validate or fence a committed TypeGraph schema version. For schema commits and multi-statement schema-managed writes on Cloudflare, use **Durable Objects** (below), whose SQLite storage exposes an interactive transaction runner. **Important:** D1 has no interactive transaction primitive (`D1Database.batch(...)` is transactional, but batch-only — not an interactive runner), so `store.transaction()` refuses on D1 before invoking its callback. See [Limitations](/limitations) for details. For a transactional Cloudflare SQLite store, use **Durable Objects** (below) instead. For the same reason, a write guarded by a **declared constraint** — edge cardinality other than `many`, a `disjointWith` axiom, a shared-scope unique, or dynamic `getOrCreateByEndpoints` convergence — is refused on D1 with `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` rather than committed unfenced. See [Declared constraints require an interactive transaction](#declared-constraints-require-an-interactive-transaction). ## Cloudflare Durable Objects (SQLite) A store backed by `drizzle(ctx.storage)` inside a Durable Object is **auto-detected** as `transactionMode: "do-sqlite"` and reports `capabilities.execution.interactiveTransactions: true` — no `executionProfile` hint needed. Unlike D1, Durable Objects expose an interactive storage transaction runner, so adapter stores can provide fully atomic `store.transaction()` and `store.withTransaction()` operations. The runtime authorizer forbids temporary tables, so the same profile reports `capabilities.graphAnalytics.supported: false`. Traversal algorithms such as `shortestPath`, `reachable`, and `weightedShortestPath` automatically use their inline fallback; temporary-table-only analytics such as `weaklyConnectedComponents` throw `UnsupportedBackendCapabilityError`. The authorizer also rejects SQLite's `analysis_limit` tuning PRAGMA. Statistics refresh catches that specific authorization error and still runs scoped `ANALYZE`; this affects refresh cost only, not query results. The same profile advertises Cloudflare's 100-bound-parameter query limit. TypeGraph uses that hard ceiling for its managed write batches and recorded-history flushes; capability overrides may lower it but cannot raise it. Literal `.in()` and `.notIn()` query lists are packed into one JSON-bound parameter, so the list itself does not exhaust the Durable Object budget. ```typescript import { drizzle } from "drizzle-orm/durable-sqlite"; import { createAdapterStoreWithSchema } from "@nicia-ai/typegraph"; import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; export class MyObject { constructor(private ctx: DurableObjectState) {} async handle() { const db = drizzle(this.ctx.storage); const backend = createSqliteBackend(db); // Boots schema/DDL outside any storage transaction (no DDL in the // business transaction); the schema-version commit uses the // do-sqlite runner. const [store] = await createAdapterStoreWithSchema(graph, backend); // Atomic across TypeGraph + the product's own relational tables: await store.transaction(async (tx) => { await tx.nodes.Document.update(documentId, props); if (tx.sqlAvailability !== "available") { throw new Error(`Native transaction unavailable: ${tx.sqlAvailability}`); } const sqlTx = tx.sql; await sqlTx.insert(documentVersions).values(versionRow); }); } } ``` TypeGraph delegates to the async storage runner `ctx.storage.transaction(async …)` (Drizzle's own `db.transaction()` on Durable Objects is `ctx.storage.transactionSync` and cannot span an `await`, so it is not used). See the [Cross-Store Transactions recipe](/recipes#cross-store-transactions-drizzle--typegraph) for the caller-owned (`withTransaction`) and graph-owned (`tx.sql`) shapes. ## Backend Capabilities Check what features a backend supports: ```typescript const backend = createSqliteBackend(db); const store = createStore(graph, backend); if (store.capabilities.execution.interactiveTransactions) { await store.transaction(async (tx) => { /* ... */ }); } else { // Handle non-transactional execution } if (store.capabilities.vector?.supported) { // Vector similarity queries available } ``` `store.capabilities` is the portable runtime source of truth; adapter authors can inspect the same object as `backend.capabilities`. The shape is: | Field | Meaning | | -------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | | `execution` | Execution boundaries: `interactiveTransactions`, exact-resource `atomicBatch` support, and derived `unitOfWork` | | `windowFunctions` | SQL window functions such as `ROW_NUMBER()` are available | | `orderedAggregates?` | Ordered scalar and record collection aggregation; absent means unsupported | | `constraintClaims?` | The backend carries the claim relations that fence declared constraints without a lock (see below) | | `durableEdgeMatchIdentity?` | Edge writes persist and atomically arbitrate a schema-declared endpoint/property identity | | `graphAnalytics?.{supported,mathFunctions}` | Static support for whole-graph temporary-table iteration, plus availability of deferred transcendental-math algorithms | | `vector?.metrics` / `vector?.indexTypes` / `vector?.maxDimensions` | Vector strategy capabilities (present once a vector strategy is configured) | | `fulltext?.{supported,languages,phraseQueries,prefixQueries,highlighting}` | Fulltext strategy capabilities | | `recursiveTraversal?.{supported,reason}` | Whether the engine can compute a bounded transitive closure of a relation in one round trip — a recursive CTE, or a graph-native expansion operator. **Absent means supported** | | `writeFence?.{mechanism,drain}` | How this engine excludes concurrent writers, and how far a caller can drain a table lock — see [Write fence declaration](#write-fence-declaration-writefence) | `expr.collect()` requires `orderedAggregates: true`. Bundled PostgreSQL supports it. Supported preparable synchronous SQLite clients and the dedicated async `createLibsqlBackend()` factory are probed when the backend is created; the query itself adds no discovery statement. Other unprobed SQLite connections default to unsupported. If you have verified that your engine supports aggregate-local ordering, declare `capabilities: { orderedAggregates: true }` in the bundled backend options. Older or unsupported engines must retain `false`; collection queries are refused before execution. SQLite introduced `NULLS FIRST` / `NULLS LAST` ordering in version 3.30 and aggregate-local ordering in [version 3.44](https://www.sqlite.org/releaselog/3_44_0.html). Collection compilation avoids depending on JSON object subtype preservation during sorting. Existing reads continue to work when ordered aggregates are unavailable. Custom dialect adapters implement `orderedScalarJsonArray` with one required object argument: `{ value, valueType, orderBy, filter }`. Migrate positional implementations by destructuring that object, applying `filter` as an aggregate `FILTER (WHERE ...)` before wrapping the aggregate in the empty-input `COALESCE`, and leaving it off when `filter` is `undefined`. The aggregate must preserve included NULL operands and return `[]` for empty input. The `filter` key itself is required in the adapter contract, even though its value may be `undefined`, which requires old positional implementations to migrate explicitly. The dialect adapter contract also requires `orderedRecordJsonArray` for `expr.collect({ field: scalarExpression }, options)`. Custom adapters must add this method when upgrading. It builds an ordered JSON array of flat records with the named projected scalar fields, honors the required ordering tuple and optional aggregate-local filter, preserves admitted SQL NULL fields within each record, and returns `[]` for empty input. The same `orderedAggregates` capability governs scalar and record collections; a custom adapter must supply both SQL emitters before declaring that capability. The former top-level `capabilities.transactions` override is not interpreted as an alias. Bundled factories refuse it with `LEGACY_CAPABILITY_OVERRIDE`, including for JavaScript and already-compiled callers, because transaction availability and atomic batching are now independent facts. Move the value to `capabilities.execution.interactiveTransactions`; root `atomicBatch` support is discovered from the bundled transport and cannot be claimed through factory overrides. Bundled PostgreSQL transaction factories may expose `atomicBatch: "session"` on the exact already-open transaction object. That declaration is paired with fresh transport and semantic registrations and is never inherited by an ordinary derived backend. `execution.unitOfWork` is derived, never declared by a factory or override: `"optimistic-retry"` when `interactiveTransactions` is `true` AND the resolved write fence is `{ mechanism: "row", conflict: "commit-time" }` (see [Write fence declaration](#write-fence-declaration-writefence) below); `"interactive"` when `interactiveTransactions` is `true` otherwise; else `"batch"` when `atomicBatch` is not `"none"` (an HTTP-only driver such as `drizzle-orm/neon-http`, which cannot hold an open session but does support a native atomic program); else `"none"`. Only the root capability derivation ever resolves the write-fence plan needed for the `"optimistic-retry"` arm — a derived or session-scoped capabilities object (a `store.transaction` session, a projected backend) has no way to re-resolve that plan for itself, but it carries the root's answer forward instead of losing it: it reads whether its own source object was already `"optimistic-retry"` and keeps the tier for as long as `interactiveTransactions` stays `true`, falling back to `"interactive"` only for a capabilities object whose source never carried the tier to begin with. Two further internal readers key off the `"batch"` value: the batch-tier write verdict (`resolveBatchWriteVerdict`) that produces `BATCH_WRITE_UNSUPPORTED` refusals, and the autocommit single-statement eligibility gate that decides whether a supplied-id singleton create can fuse its schema fence into one statement. Even absent `"optimistic-retry"`, `unitOfWork` exists so any caller can tell the execution shapes apart without re-deriving the same distinction from `interactiveTransactions` and `atomicBatch` separately. Under `"optimistic-retry"`, every TypeGraph-owned transaction that acquires a fence row replays a real commit-time conflict as a whole unit, up to `OPTIMISTIC_RETRY_ATTEMPTS` (3) attempts, and only exhausting that budget (or a non-retryable failure) surfaces `TransactionConflictError` to the caller — see [Retrying on conflict](/schemas-stores#retrying-on-conflict). That covers every store-owned write (collection create/update/delete, bulk paths, `importGraph`, identity maintenance, contribution rebuild, index materialization) as well as the two backend-owned transactions that acquire the schema-commit fence row directly, outside the store's own write path: graph-template instantiation and a schema commit (`commitSchemaVersion` and its three siblings, via `runSchemaWriteTransaction`). A nested write running inside an existing transaction (`store.transaction`, an adopted transaction) never retries on its own: it cannot restart a transaction it does not own, so its conflict propagates unchanged to the outermost store-owned write, or to `store.transaction` itself. This tier therefore changes behavior only for a transaction that opens its own top-level connection. An `"optimistic-retry"` backend requires `node:async_hooks`' `AsyncLocalStorage` to detect a retried unit nested inside another one; on a runtime where it is unavailable, the first retried unit `runRetriedUnit` opens is refused with `OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` rather than degrading to independent, unsafe per-unit retries, while interactive backends are unaffected and keep retrying (`store.transaction`'s own `retry` option) with no async-context support at all. `graphAnalytics.supported` describes the backend shape, not mutable PostgreSQL session state. A hot standby or a role without the database `TEMP` privilege can still reject the working-table transaction that the iterative graph algorithms open: a standby refuses the read-write transaction itself, and a role without `TEMP` refuses the `CREATE TEMP TABLE` inside it. Both refusals reach the caller as `UnsupportedBackendCapabilityError`, with the PostgreSQL error retained as its `cause`. ### Durable edge match identity capability `capabilities.durableEdgeMatchIdentity: true` is a correctness promise. A custom backend making it must provide all of these guarantees: - Every edge write carrying `InsertEdgeParams.matchIdentity` stores both the name and key with the row. They are either both absent or both present. - A database constraint atomically owns uniqueness over `(graph_id, kind, match_identity_name, match_identity_key)`. Soft deletion keeps the key; physical hard deletion releases it. - `commands.execute()` handles a durable `edge.converge-create` as one database decision and returns the authoritative `created` or `found` row. Returning `unsupported` fails closed with `DURABLE_EDGE_MATCH_IDENTITY_COMMAND_UNSUPPORTED`; TypeGraph does not fall back to a read-then-write race. - Storage exists before runtime writes. Implement `ensureEdgeMatchIdentityStorage` for privileged schema adoption, or provision the columns, pair constraint, and unique arbiter independently before setting the capability. `insertEdgesDurableBatchReturning` is an optional throughput member. When implemented, every input must carry a durable identity, conflicts are omitted from the returned rows, and returned rows identify exactly which inputs were created. Omitting it preserves correctness through per-row authoritative commands, but loses the set-oriented bulk/import fast path. `findEdgesByHeterogeneousEndpointSet` is likewise an optional set-read optimization. An input carrying `opposite` requests an exact directed endpoint pair, not every edge incident to the first endpoint. One call may contain only incident inputs or only exact-pair inputs; mixing the two modes is refused. A backend that omits the member retains the exact per-pair fallback. ### Validity-end clearing capability Custom backends must advertise `capabilities.clearValidTo: true` only when both `updateNode` and `updateEdge` apply `clearValidTo: true` by storing SQL `NULL` in `valid_to`. The built-in SQLite and PostgreSQL adapters do. An explicit clear on a backend without that promise is refused with `ConfigurationError` code `CLEAR_VALID_TO_UNSUPPORTED` before coalescing or writes, so the result does not depend on whether the target row is already open. Omission still means preserve; custom backends that do not support clearing remain compatible with all writes that omit the option. ### Recorded-table migration DDL (`recordedTableDdl`) `GraphBackend.recordedTableDdl` is an optional, synchronous callback used only by the offline timestamp-only recorded-time preview migration. The migration calls it twice, once with temporary table names and once with the final names, and expects DDL for `recordedClock`, `recordedNodes`, and `recordedEdges`. The backend owns this callback because table creation, indexes, and named constraints are dialect-specific and must not pull Drizzle into portable entrypoints. A custom backend can omit the callback unless it created data in the old preview schema. If `migrateLegacyRecordedTime` discovers that schema and the callback is absent, it throws `UnsupportedBackendCapabilityError` with `details.capability: "recordedTableDdl"`. When the engine names primary-key constraints, the temporary and final callback results must either both name the constraint or both omit it; a one-sided result throws `ConfigurationError` code `RECORDED_DDL_CONSTRAINT_NAME_MISMATCH`. The callback only describes DDL. It must not execute statements or inspect the catalog, because the migration invokes it inside its transaction. See [Migrating Preview Recorded Time](/schema-management#migrating-preview-recorded-time) for the operator workflow. ### Recursive traversal capability Both bundled backends declare `capabilities.recursiveTraversal: { supported: true }`. **Absent means supported** — mirroring `returning`, not `constraintClaims`: every existing custom backend already runs the six recursive-CTE emission sites unconditionally, so absence meaning unsupported would refuse traversals that work today. ```typescript const capabilities: Partial = { recursiveTraversal: { supported: false, reason: "engine has no WITH RECURSIVE / equivalent" }, }; ``` A backend that genuinely lacks the primitive declares `{ supported: false, reason }`. A factory refuses a contradictory declaration — `supported: false` with no `reason`, or `supported: true` with a dangling `reason` — with `ConfigurationError` details code `CAPABILITY_DECLARATION_CONTRADICTION`. Five operations refuse when unsupported: variable-length (`traverse`) queries, `store.subgraph()`, historical identity class reads, identity-expanded historical queries, and the identity window-ledger read — each throwing `ConfigurationError` code `RECURSIVE_TRAVERSAL_UNSUPPORTED` with `details.operation` naming the site and `details.reason` echoing the declaration. `weightedShortestPath` is the one exception: on a backend with temporary statements but no recursion, it **falls back** to a per-hop predecessor walk instead of refusing, issuing `pathLength + 1` extraction statements for the path a recursive CTE would have returned in one round trip. The unweighted `shortestPath` (along with `reachable`, `canReach`, and `neighbors`) emits no recursive CTE at all — it routes through the iterative working-table or inline path instead — so it neither refuses nor falls back regardless of this declaration. ### Write fence declaration (writeFence) TypeGraph serializes a family of writes — Operational Identity's mutations, and the TypeGraph-owned recorded-clock allocation behind `history` / `revisionTracking` — behind a per-graph fence rather than trusting the engine's default isolation. `capabilities.writeFence` declares what this backend can provide, as two independent facts, and `resolveWriteFencePlan` is the one place that declaration turns into a plan every lock site consumes instead of re-deriving: ```typescript const capabilities: Partial = { writeFence: { mechanism: "advisory", drain: "table-lock" }, }; ``` `mechanism` is how the backend excludes concurrent writers. `writeFence` is a discriminated union on `mechanism`, and `drain` is a field of the `"advisory"` and `"row"` shapes only — `"engine-serialized"` and `"caller-serialized"` declarations carry no `drain` key at all: | `mechanism` | Meaning | | --- | --- | | `"advisory"` | A keyed `pg_advisory_xact_lock`-style lock a caller takes explicitly. Needs `fenceSql` (below) and a `drain`. | | `"row"` | A keyed exclusion spelled by TypeGraph itself against the never-dropped fences relation, for an engine with no advisory-lock primitive. Needs a `drain` and a `conflict` (below); a `fenceSql.isolationFactExpression` is optional (absent means recorded capture and match-key convergence fail closed on an unknown isolation fact, exactly as they do for a target that supplies neither). | | `"engine-serialized"` | The engine serializes writers by construction — SQLite's single writer slot. No lock statement, no `fenceSql`, no `drain`. | | `"caller-serialized"` | A deployment-level promise, not an engine fact — see below. No lock statement, no `drain`; a `fenceSql` the backend still carries is used only for its isolation-fact read (recorded capture's isolation guard). | `drain` (on `mechanism: "advisory"` or `"row"` only) is a separate fact: whether a caller that already excluded other writers can additionally take a relation-wide lock on the table a drain site protects: | `drain` | Meaning | | --- | --- | | `"table-lock"` | Yes — a `LOCK TABLE`-style statement is available and the drain site takes it. | | `"quiescent"` | The resource is already exclusive for some other reason (an advisory or row lock layered under a deployment's own `caller-serialized` promise, for instance), so the drain site takes NO statement — one it does not need rather than one it cannot spell. | | `"none"` | Neither — a drain site refuses, naming the drain. | `conflict` (on `mechanism: "row"` only) is the engine fact for what happens when two writers acquire the SAME fence row: | `conflict` | Meaning | | --- | --- | | `"wait"` | A lock-based engine — the second acquirer's statement blocks until the first commits, exactly like an advisory lock. | | `"commit-time"` | An optimistic-concurrency engine — both acquirers proceed and the loser's COMMIT fails. Correctness comes from the retry owner replaying the whole unit, never from waiting, so `conflict: "commit-time"` derives the `"optimistic-retry"` execution tier (see [Backend Capabilities](#backend-capabilities) above) and requires it: declaring it on a non-interactive backend is refused the same way an out-of-place `drain` is. | `"engine-serialized"` and `"caller-serialized"` satisfy every drain site unconditionally — a writer slot and an in-process serialization promise are each already a stronger exclusion than any `drain` value could add, so attaching one to either mechanism is refused (see **Runtime validation** below) rather than silently ignored; attaching `conflict` to anything but `"row"` is refused the same way. `resolveWriteFencePlan` resolves one of five plans: - `{ kind: "lock", drain, sql }` — take the declared advisory lock (`sql`, the target's own spelling), and, when `drain === "table-lock"`, the table lock a drain site needs. - `{ kind: "row", drain, conflict, sql }` — take the SAME `sql.acquireKeyed` / `sql.acquireKeyedWithIsolation` a `"lock"` plan's site calls, spelled instead against the fences relation; `conflict` is the one fact a `"row"` site (and the execution tier) reads that a `"lock"` site never needs. - `{ kind: "engine-serialized" }` — no lock needed; the engine serializes writers by construction. - `{ kind: "caller-serialized" }` — no lock needed; the deployment's own promise excludes concurrent writers (see below). - `{ kind: "unfenced" }` — no declaration at all. Every fence that guards a read-then-write across statements refuses rather than running unfenced. Only a predicate carried *inside* the statement it guards degrades, since one statement cannot race itself. Resolution order: (1) the declared `writeFence` value, if present; (2) absent, AND the backend was built by `createSqliteBackend` / `createPostgresBackend` — derived from `dialect`, which is exactly what every lock site used to compute inline (this derivation never resolves `"row"`: it is the two bundled dialects' own `"advisory"`/`"engine-serialized"` split); (3) absent on anything else — `unfenced`, because an undeclared custom backend is by definition uncertified and inferring lock support from `dialect` alone is the unsound inference this capability replaces. The two bundled backends resolve exactly these declarations — copy the one matching your engine: - PostgreSQL: `writeFence: { mechanism: "advisory", drain: "table-lock" }` - SQLite: `writeFence: { mechanism: "engine-serialized" }` (no `drain`: the writer slot already excludes every drain site's writer, so a drain site under it always takes no statement — the same behavior `drain: "quiescent"` describes for `"advisory"`/`"row"`, without a `drain` field to spell it) A backend that declares `mechanism: "advisory"` also supplies `fenceSql`: `lockTables` (only needed when `drain: "table-lock"`) plus the two composable, no-`SELECT` forms `advisoryLockExpression` / `isolationFactExpression` a statement embeds inside a larger query it builds itself (see the schema-write-fence discussion above) — the complete `FenceSql` bag. `resolveWriteFencePlan`'s `lock` arm derives the standalone-statement forms every ordinary lock site actually calls — `acquireKeyed`; `acquireKeyedWithIsolation` (the lock plus the session's isolation-level fact, read in the same statement it locks in); and `isolationFact` — from those two expressions, so a backend author never spells both forms separately. `mechanism: "row"` needs no `advisoryLockExpression` at all: TypeGraph spells its own `acquireKeyed` / `acquireKeyedWithIsolation` against the fences relation (see below), and a `fenceSql.isolationFactExpression` — when supplied — rides the SAME acquisition statement, so a `"row"` target's isolation fact is read on the exact connection that took the row. The bundled PostgreSQL spelling is exported as `postgresFenceSql` from `@nicia-ai/typegraph/adapters/drizzle/postgres` — pass it straight through as `fenceSql` when wrapping that backend (under either `"advisory"` or `"row"`), or supply a custom `FenceSql` matching a different engine's lock syntax. A backend that declares `mechanism: "advisory"` with a `fenceSql` missing a member the resolved `mechanism`/`drain` combination needs is refused at construction with details code `WRITE_FENCE_SQL_UNAVAILABLE`, naming the missing member; `"row"` is refused the same way only for `drain: "table-lock"` without `lockTables` — its acquisition statement needs no author-supplied spelling at all, so a missing `tableNames.fences` instead refuses the first time a keyed site actually acquires the row, not at construction; `"engine-serialized"` and `"caller-serialized"` need no `fenceSql` to take a lock at all. #### The fences relation A `"row"`-mechanism backend needs one relation, `typegraph_fences(key TEXT PRIMARY KEY, generation BIGINT NOT NULL)` (`INTEGER NOT NULL` on SQLite) — part of TypeGraph's base schema on both bundled dialects, so a fresh install already has it and `generateSqliteMigrationSQL` / `generatePostgresMigrationSQL` add it to an existing database. It is **never dropped, never cleared by `clear()`, and never row-deleted** — the same durability contract `schema_versions` and `recorded_clock` carry. Every acquisition is one portable statement TypeGraph spells itself, never the profile: `INSERT INTO typegraph_fences (key, generation) VALUES (key, 1) ON CONFLICT (key) DO UPDATE SET generation = generation + 1 RETURNING generation`, keyed on `${namespace}:${key}` — the SAME advisory namespaces and per-position keys an `"advisory"`-mechanism backend locks on, verbatim, so the lock-order contract carries over unchanged to an engine using the fences relation instead of `pg_advisory_xact_lock`. A custom backend supplies the relation's physical name through `tableNames.fences` (defaulted to `typegraph_fences` by both bundled factories) exactly as it names every other TypeGraph-owned table. #### Runtime validation TypeScript's discriminated union only holds a caller who goes through the type checker — a plain JavaScript backend author, or a value round-tripped through JSON or a config file, can still supply an unrecognized `mechanism` string, an unrecognized `drain` or `conflict` string, a `drain` attached to `"engine-serialized"` / `"caller-serialized"`, or a `conflict` attached to anything but `"row"`. `resolveWriteFencePlan` validates every declaration — whether it came from `capabilities.writeFence` directly or from the first-party dialect fallback — before shaping a plan from it, and refuses with `ConfigurationError` details code `WRITE_FENCE_DECLARATION_INVALID`, naming the invalid `field` (`"mechanism"`, `"drain"`, or `"conflict"`) and, for an unrecognized value, the `accepted` list. An unrecognized `drain` never falls through to behaving like `"quiescent"` — it is refused outright, the same as an unrecognized `mechanism`. The same validator refuses `conflict: "commit-time"` outright when the target's own `capabilities.execution.interactiveTransactions` is `false`: that value is honored only by the `"optimistic-retry"` execution tier, which never derives without an interactive transaction to replay inside, so accepting the declaration there would silently drop it rather than apply it. #### `caller-serialized`: the promise split into two halves `caller-serialized` is for a deployment that knows its database has no other concurrent writer, but whose engine is neither an advisory-lock engine nor a single-writer one — a PostgreSQL-wire engine with no working `pg_advisory_xact_lock` / `LOCK TABLE`, for example. The promise has two halves, and TypeGraph only enforces the first: - **In process**, TypeGraph enforces it: every root member the backend classifies as mutating — collection writes, `store.transaction` / `transactionWithNative`, schema commits, identity and contribution maintenance, index materialization, table/DDL provisioning, `clearGraph`, import, and the raw-SQL members (`execute`, `executeRaw`, `executeStatement`, `executeTemporaryStatement`) that can carry an arbitrary write — runs through one per-backend serialized queue, so two concurrent calls through one pool cannot race each other. A root write awaited from inside a `store.transaction` callback is refused rather than left to deadlock behind the transaction's own queue slot. - **Outside the process**, the deployment enforces it: no other client writes to this database while this backend is open. TypeGraph cannot see or verify that half; declaring `caller-serialized` is asserting it. Adopting an externally owned transaction (`store.withTransaction(externalTx)`, backed by `adoptTransaction`) is refused outright on a `caller-serialized` backend, with `ConfigurationError` details code `CALLER_SERIALIZED_REFUSES_ADOPTION`: an adopted transaction's lifetime belongs to the caller, not to this backend's write-unit queue, so there is no honest way to hold a queue slot open for it — queuing it would block every other queued write until the caller's own transaction ends, and leaving it unqueued would let its writes interleave with the queue's own, silently breaking the promise `caller-serialized` makes. Open the transaction through this backend's own `transaction()` / `transactionWithNative()` instead, or do not declare `caller-serialized` on a backend that needs cross-store adoption. `createPostgresBackend` accepts `writeFence: { mechanism: "caller-serialized" }` — a claim about the deployment — while still refusing `mechanism: "engine-serialized"` outright, because that value is a claim about the *engine*, which this factory's own engine does not back. Constructing Operational Identity, or `history: true` / `revisionTracking: true`, against an `unfenced` backend is refused immediately at `createStore` — never mid-flush — with `ConfigurationError` details code `IDENTITY_REQUIRES_WRITE_FENCE` (identity) or `RECORDED_CLOCK_REQUIRES_WRITE_FENCE` (recorded-clock allocation), and the refusal message names the exact declaration line to add. A `lock` plan whose `drain` cannot back a site's `requires: "drain"` — `drain: "none"` — is refused with details code `WRITE_FENCE_UNAVAILABLE`, naming `details.operation` and the drain that could not be satisfied. `"engine-serialized"` and `"caller-serialized"` satisfy either `requires` value (`"keyed"` or `"drain"`) without consulting `drain`. The PostgreSQL schema fence refuses too, and it is worth knowing why it is not on the degradable side. The per-graph advisory lock plus `SELECT ... FOR UPDATE` a schema commit takes, and the `FOR SHARE` a managed write takes on that same row, each fence a read-then-write sequence that spans **statements**: `commitSchemaVersion` reads the active version and then writes the flip, and a managed write holds its `FOR SHARE` for the remainder of the transaction so the version it asserted stays true through the writes that follow. Skipping those locks would not give a slower-but-correct path; it would assert a version and then let the very change the assertion was checking for land before the write. So a PostgreSQL-dialect backend that resolves `unfenced` is refused at the schema commit with `WRITE_FENCE_UNAVAILABLE`, naming the operation. The one part that *does* degrade is the fence folded into a managed insert's own statement. That predicate is evaluated inside the INSERT that depends on it, and one statement cannot race itself: with no locking clause the fence subquery still yields no row when the expected version is no longer active, so the INSERT still writes nothing. This is how SQLite has always run the path, on the strength of its writer slot. This matters for a PostgreSQL-wire engine that implements neither `pg_advisory_xact_lock` nor the `FOR UPDATE` / `FOR SHARE` clauses, and whose engine merges concurrent transactions rather than serializing them. `writeFence` has no arm for "no exclusion mechanism at all" — every `mechanism` value claims something real. An engine with neither locks nor a writer slot nor a caller promise to make has nothing honest to declare, and `unfenced` is how that shows up downstream — but neither bundled factory will hand you that backend. `createPostgresBackend` and `createSqliteBackend` both build on `createSqlBackend`, which refuses at construction, with `ConfigurationError` details code `ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION`, when the resolved capabilities carry no `writeFence` at all — including `capabilities: { writeFence: undefined }` passed to either factory, which no longer builds a backend the way it once did: ```typescript createPostgresBackend(db, { capabilities: { writeFence: undefined }, }); // throws ConfigurationError: ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION ``` `unfenced` is reachable only outside that gate: a hand-assembled `GraphBackend` that never goes through `createSqlBackend`, or a custom `SqlEngineProfile` whose `declaredCapabilities.writeFence` some other override clears before it reaches a call site — never through either bundled factory. Getting there is a way of admitting the engine truly has nothing to declare; TypeGraph then refuses a schema-managed store built on it at `createStore`, naming the missing capability, rather than running a fence the engine cannot enforce. If the deployment instead knows it is the only writer of this database — a pool clamped to one connection, or a single-writer topology otherwise enforced outside TypeGraph — declare `writeFence: { mechanism: "caller-serialized" }` instead (see above): that is the honest way to spell a deployment convention. Do not reach for `mechanism: "engine-serialized"` for the same purpose — that declaration means the *engine* serializes writers by construction, and a deployment convention is not a construction. `createPostgresBackend` refuses that particular claim outright for this reason. ### Capability bundles A **capability bundle** groups a set of `GraphBackend` members that one operation family needs together, with one verdict resolver and one member accessor, so a caller never re-derives "does this backend support X" from a scattered `undefined` check. Seven pilot bundles ship in this release: | Bundle | Kind | Disposition | | ------------------------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `claims` | gated | Bidirectional cross-check between the `constraintClaims` declaration and the core members; disagreement in either direction refuses with `CONSTRAINT_CLAIM_SURFACE_MISMATCH` | | `statementExecution` | gated | Core `executeStatement` absent refuses with `IDENTITY_REQUIRES_STATEMENT_EXECUTION` | | `recordedRevisionOrigins` | gated | Core `ensureRevisionOriginsTable` absent refuses with the operation's own typed error | | `batchPointRead` | graduated | `getNodes` absent falls back to per-id `getNode`; `getEdges` absent falls back to per-id `getEdge` | | `uniqueSidecarBatch` | graduated | `insertUniqueBatch` absent falls back to `issueClaimsIndividually`; `checkUniqueBatch` absent falls back to a per-key loop; `hardDeleteUniquesByNodeIds` absent refuses with the operation's own typed error | | `contributionHealth` | graduated | `verifyContributions` / `repairContributions` / `rebuildContribution` absent each refuse with the operation's own typed error; `probeContributions` absent falls back to `{ entries: [] }` | | `endpointSetRead` | graduated | `findEdgesByEndpointSet` absent refuses set-oriented `bulkFindFrom` / `bulkFindTo` with `ENDPOINT_SET_READ_UNSUPPORTED`; singleton reads remain available | The port-mismatch rule that governs every bundle's member accessor is keyed to the disposition, not blanket: a `refuse`-disposition row whose backend object cannot actually reach the member throws that bundle's own `portSurfaceCode` (`CONSTRAINT_CLAIM_SURFACE_MISMATCH` for `claims`, `BUNDLE_PORT_SURFACE_MISMATCH` for the other six); a `fallback`-disposition row whose port cannot reach the member takes its declared fallback instead of throwing — the verdict said the member was there, the object it binds against says otherwise, and a fallback row is defined to degrade rather than assert. This bundle model ships for **seven of the twenty-one** member-bearing operation families measured in this workstream; the remaining fourteen are a named follow-up workstream, not a silent gap — their members keep working exactly as before, unbundled, with an access-count ceiling that prevents new scattered checks from accumulating ahead of that follow-up. A backend author does not need to do anything for these seven bundles today: both bundled backends already carry every core member each bundle's `dialects` scope requires. The atomic transport conformance runner is the foundation for certifying a **third-party** backend: the author supplies engine-specific statements, state observers, and exact-root provenance checks, while the runner asserts the shared transport contract. Bundle verdicts remain a separate check against the declared capabilities and the object the calls actually execute on. A backend must therefore make every declared capability (`constraintClaims`, `contributions`, and execution support) truthful about what the active backend object implements, not just which fields it sets. Run the conformance fixture in the custom backend's own test suite, then pair the earned declaration with the exact root transport in its factory: ```typescript import { decorateBackend, registerAtomicMutationPrograms, registerAtomicSqlProgram, runAtomicMutationProgramConformance, runAtomicTransportConformance, } from "@nicia-ai/typegraph/backend"; const backend = createCustomBackend({ execution: { interactiveTransactions: false, atomicBatch: "root", }, }); registerAtomicSqlProgram(backend, { executeAtomicBatch }); const authorCreatedWrapper = decorateBackend(backend, {}); await runAtomicTransportConformance({ ...transportCases, backend, derivedBackends: [authorCreatedWrapper], executeAtomicBatch, }); registerAtomicMutationPrograms(backend, mutationPrograms); const semanticCases = buildSemanticCases({ backend }); await runAtomicMutationProgramConformance({ backend, derivedBackends: [authorCreatedWrapper], equal: deepEqual, cases: semanticCases, }); ``` Transport registration is exact-resource evidence only: a derived backend does not inherit it, and a second registration on the same object is refused rather than replacing the function production uses. A bundled PostgreSQL transaction session earns a separate registration bound to its pinned client; it does not inherit the root's registration. Create wrappers with the exported `decorateBackend()` seam so the runner can verify their lineage back to the registered root instead of accepting an unrelated object as derivation evidence. The conformance fixture's mandatory provenance checks prove registration, lineage, derived isolation, and—when applicable—transaction isolation against the real objects supplied by the backend author. A non-interactive root reports the transaction-isolation check as skipped rather than claiming evidence it could not obtain. Transport registration certifies mechanics, not graph semantics, and therefore does not by itself opt a custom backend into any Store mutation program. The separate `registerAtomicMutationPrograms()` call is the semantic boundary: each member declares one complete TypeGraph mutation family implemented by that exact backend resource. Omitted families retain the portable path, and an empty profile or a profile registered before its atomic transport is refused with `ConfigurationError`. The semantic executors must preserve the same schema fence, validation, side-effect, refusal classification, rollback, postimage correlation, result ordering, and bind-ceiling contracts as the bundled implementation. Registering one family is not evidence for another. Derived and projected backends inherit neither registration. An exact transaction session must be registered independently before Store code can dispatch through it. `runAtomicMutationProgramConformance()` is the executable semantic boundary. For every reachable positive-limit variant in `mutationPrograms`, the fixture supplies three real Store-level cases: 1. an ordered success whose return value and independently read committed state both match their oracles; 2. a stale-schema-fence refusal that leaves the database unchanged; and 3. a family-specific typed refusal that either rolls back every sibling write after native dispatch or explicitly refuses before dispatch without writing. The runner resolves the profile from `backend`; it does not accept a detached profile description, caller-supplied provenance verdict, or fixture-owned dispatch counter. Before any fixture preparation can write, it validates the complete case inventory and probes the author's actual derived backend objects. It observes dispatch inside the exact registered executors and therefore refuses a success that silently used the portable fallback, a case bound to a different family claim, a missing or duplicate family case, and a case that claims an unregistered family. A zero entry limit is an honest opt-out and does not require an unreachable case. `mutateEdges` has separate `resolvedSet` and `durableConvergence` variants because proving one does not prove the other. Every semantic case identifies the exact `backend` its callbacks use. The runner checks that binding and the registered profile identity before any preparation, again between preparation and execution, and after execution, so a pre-dispatch refusal or a mid-run registry replacement cannot borrow another root's certificate. The fixture callbacks should invoke public Store methods and inspect committed rows through an independent database read. Supply at least one real wrapper or derived backend created with `decorateBackend()`; the runner does not manufacture a projection and mistake that tautology for author evidence. Do not instrument or replace the registered executors—the runner owns dispatch evidence. Run conformance with exclusive use of that exact root: unrelated same-variant writes during the observation window cannot be distinguished from fixture traffic. Mark each semantic refusal's `dispatch` as `"required"` or `"pre-dispatch"` according to the Store contract, and do not use executor return rows as the state oracle. Stale-fence cases always require native dispatch regardless of a fixture value supplied by untyped JavaScript. Match Store-level typed errors rather than raw driver sentinels. The runner is framework-agnostic, so the same fixture runs in the custom backend's own test suite. Pair it with the shared cross-backend Store integration suite; transport conformance alone cannot prove graph semantics. The profile is family-scoped: | Member | Store operations authorized | | ----------------------------- | ----------------------------------------------------------------------------------------- | | `createNodes` / `createEdges` | Eligible direct `bulkInsert()` and `bulkCreate()` programs | | `replaceNodes` | Eligible complete-document `nodes.bulkReplaceById()` programs | | `deleteNodes` / `deleteEdges` | Eligible direct `bulkDelete()` programs | | `updateNodes` / `updateEdges` | Eligible resolved update-only sets | | `mutateNodes` / `mutateEdges` | Eligible mixed create/update sets; the edge family also owns durable endpoint convergence | Executor limits such as `maxEntries`, `replaceNodes.maxEntries.plain`, `replaceNodes.maxEntries.claimed`, `createNodes.claimSupport.maxInputCostPerEntry`, and the two edge mutation ceilings are part of the registration contract and must be nonnegative integers; zero honestly declares that the backend's bind budget cannot admit one member of that shape. TypeGraph validates those declarations before publishing the exact-root profile. `claimSupport.families` explicitly advertises `uniqueness` and/or `disjointness`; an empty list with a zero bound honestly opts out of all claim work. The Store calls the exported `atomicNodeClaimInputCost()` owner for each member and refuses the native path when its complete normalized claim set exceeds the executor's declared bound. Custom executors must use that same helper instead of reproducing its dialect-reviewed bind formula. `deleteNodes.releasedClaimFamilies` similarly declares which owner-side claim cleanup the delete program proves. Bundled replacement executors also expose an `accepts(entries)` pre-dispatch proof. It packs prepared members with the same bind-weighted planner used by execution, so claimed batches are admitted by their actual work instead of an unrelated fixed 32-entry ceiling; `false` is an explicit no-SQL fallback verdict. Custom executors may provide the same exact admission seam when one claimed-member ceiling would be needlessly pessimistic. `replaceNodes.releasedClaimFamilies` declares which previous owner claims the replacement releases before acquiring its complete postimage claims; the Store does not infer that proof from `claimSupport`. Node create/update/mutation executors advertise derived-storage support separately through `projectionSupport.families`; omission or an empty list honestly opts out, and the Store never infers projection safety from transport registration alone. The supported families are `fulltext` and `embedding`. On a transactionless root, dedicated update-only and mixed mutation executors are independently reachable Store families even when their entry ceilings are equal, so each requires its own conformance evidence. On an interactive root, the collection-level read/partition/write unit moves into a transaction and exact-root registration does not follow; the root conformance inventory therefore excludes the mixed variants while continuing to require direct create, delete, update, and durable convergence evidence. Bundled PostgreSQL binds the same reviewed lowering to the exact transaction session and exercises the mixed node and edge variants against a real engine, including typed refusal rollback. The same session profile registers `replaceNodes`, so a caller-owned PostgreSQL transaction keeps blind replacement inside its savepoint-backed atomic program. An exact `atomicBatch: "session"` conformance fixture includes those mixed variants even though `interactiveTransactions` is true; nested-transaction isolation is reported as inapplicable because the fixture resource is already the open transaction. The transport runner accepts the same exact-session resource and certifies its ordered slots, parameter preservation, rollback, and empty-program behavior. Generic derived session objects still lose both the declaration and the identity-bound registrations. Backend authors implementing edge refusal paths use the exported `AtomicEdgeBatchEndpointRefusalError`, `AtomicEdgeBatchCardinalityRefusalError`, `AtomicEdgeConvergenceTombstoneRefusalError`, and `AtomicEdgeDeleteIdentityRefusalError` signals; restricted node deletion uses `AtomicNodeDeleteRestrictedRefusalError`. This preserves the Store's existing typed diagnostic classification rather than exposing driver-specific sentinel errors. Execution support is intentionally not collapsed into one ordered “tier.” An interactive transaction and an exact-root atomic batch are independent facts: a backend may provide either, both, or neither. Each Store operation selects the boundary its own semantics require instead of treating one mechanism as a universal substitute for the other. Endpoint-set reads have a small, independent conformance fixture for custom backends. Import `runEndpointSetReadConformance` from the `backend` entrypoint and provide the exact backend, one or more successful `FindEdgesByEndpointSetParams` cases, and at least one refusal case. The runner checks that the backend exposes `findEdgesByEndpointSet`, preserves the expected edge rows, and refuses invalid requests using the adapter's typed error. It does not create a schema or assume a driver, so the same fixture can run against any engine. A backend that omits the member remains valid for singleton reads; Store bulk endpoint reads refuse with `ENDPOINT_SET_READ_UNSUPPORTED`. ### Declared constraints require an interactive transaction A **constrained write** — one whose correctness rests on a check-then-write that no database key repeats at write time — runs its probe and its write under one per-graph mutual exclusion. That fence is a transaction-scoped construct on both dialects: SQLite's `BEGIN IMMEDIATE` writer slot, PostgreSQL's `pg_advisory_xact_lock` (which outside a transaction is taken and dropped inside its own implicit single-statement one, excluding nothing). A backend reporting `capabilities.execution.interactiveTransactions: false` can supply neither, so such a write is **refused** rather than run unfenced — a constraint enforced only when nothing races is the defect the fence exists to close. The refusal is a `ConfigurationError` with `details.code` `CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, and `details.constraint` naming which class needed the fence, because the way forward differs per class: | `details.constraint` | The write that needs the fence | Way forward without a transactional backend | | --- | --- | --- | | `edgeCardinality` | Creating or resurrecting an edge whose `cardinality` is `one`, `unique`, or `oneActive` | Declare the edge `cardinality: "many"` and enforce the limit in application code | | `edgeMatchKeyConvergence` | `getOrCreateByEndpoints` using an undeclared dynamic `matchOn` key | Declare the edge registration's durable `matchIdentity`, or use `create` with a caller-chosen id | | `nodeDisjointness` | Creating a node under a kind that participates in a `disjointWith` axiom | Drop the axiom and keep ids distinct across those kinds yourself | | `nodeUniquenessClaim` | **Updating or resurrecting** a node whose kind declares any unique constraint, of any scope — a transition reserves the new key *before* the row write it gates, and only a transaction can undo the pair together | Drop the constraint, or run updates on a transactional backend. Plain **creates** under a `scope: "kind"` unique are unaffected: their claim follows the row | | `nodeUniquenessScope` | Creating **or updating** a node under a `scope: "kindWithSubClasses"` unique that actually expands past the node's own kind | Scope the constraint to `"kind"`, which the uniques primary key enforces on its own | `importGraph` / `importGraphStream` is refused on the same backends whenever any node kind of the graph owes a claim ahead of its row — that is, declares **any** unique constraint or has a disjoint partner — or any edge kind is non-`many`. The import writes both creates and updates, so the widest of those placements is what decides it. This affects **Cloudflare D1**, **`drizzle-orm/neon-http`**, and any SQLite backend built with `transactionMode: "none"`. Durable Objects are unaffected — `do-sqlite` reports `capabilities.execution.interactiveTransactions: true` and fences normally. Unconstrained writes on those backends are untouched and keep working exactly as before: a `cardinality: "many"` edge created, updated and deleted; any node delete, including one whose kind participates in a disjointness axiom (a delete re-derives no cross-kind verdict); a node whose uniques are all `scope: "kind"`; and an undeclared `getOrCreateByEndpoints` that *finds* an existing edge in the default `ifExists: "return"` mode, or resurrects a `many` one — that resurrection is an id-keyed `UPDATE` that re-derives nothing. With `coalesceUnchangedUpserts` enabled, confirming that a single `ifExists: "update"` endpoint replay is unchanged requires the endpoint match-key convergence fence and therefore refuses on these backends. Outside the native durable-convergence envelope, the bulk `getOrCreateByEndpoints` form returns an all-live default-`"return"` batch from one set-oriented root read because that outcome writes nothing. Inside the native envelope, the authoritative upsert runs first; it preserves the logical `"found"` outcome in one exchange but may take incumbent-row locks and produce write amplification. If any member may write, the whole batch retains that refusal on transactionless roots unless it matches the narrow native durable-convergence envelope: schema-declared `matchIdentity`, `cardinality: "many"`, declared match fields, default `ifExists: "return"`, and no temporal mutation. That eligible form is one closed atomic exchange; dynamic match fields, update mode, constrained cardinality, temporal options, and all transaction-scoped or derived roots retain the refusal or fallback path required by their contracts. An otherwise eligible tombstoned winner cannot use the native path: the native attempt rolls back and transactionless convergence refuses with the typed `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` (`edgeMatchKeyConvergence`) error. Use a transaction-capable backend when schema-aware resurrection is required. ### Claim relations, and what they do not promise Underneath the lock, a declared constraint is also reserved in a **claim relation** whose primary key admits one live claimant per axis: `uniques` (for uniqueness scopes and `disjointWith` pairs) and `typegraph_edge_claims` (for `cardinality: "one" | "unique" | "oneActive"`). Both bundled backends carry them and report `capabilities.constraintClaims: true`. The claim is what makes those constraints hold for TypeGraph writers that hold no per-graph lock at all — `importGraph` is the one in the box. The protocol is application-maintained: raw SQL that writes only `nodes` or `edges` bypasses the corresponding claim write and can violate the declaration. An out-of-band writer is fenced only if it participates in the same claim protocol in the same transaction. Three properties of that mechanism are worth knowing before you rely on it: - **A claim row's lock is held to the end of the transaction, including on refusal.** A caller that catches a typed constraint error and keeps going — import's per-row recovery, or your own `try`/`catch` inside `store.transaction` — still holds the lock on the row it was refused at, and any other writer of that axis waits until the transaction ends. This is inherent to every row-lock fence, not specific to this one. - **Above READ COMMITTED, PostgreSQL reports a serialization failure instead of the typed error.** At `REPEATABLE READ` or `SERIALIZABLE`, `INSERT … ON CONFLICT DO UPDATE` raises `40001` rather than resolving the conflict, so the losing writer sees a serialization failure to retry rather than `UniquenessError`. SQLite has no such mode. This is unchanged from earlier versions, which already reserved single-kind uniqueness through the same statement. - **Pre-existing violations are neither repaired nor refused at boot.** A database that already held two live claimants of one axis before the claim relations existed keeps holding them; the next write that touches that axis is refused with the ordinary typed error naming the incumbent. `store.verifyConstraintFences()` is the read-only diagnostic that makes that state legible ahead of time: ```typescript for (const violation of await store.verifyConstraintFences()) { // violation.target names the claim row two claimants contend for console.warn(violation.family, violation.target.axis, violation.target.key); } ``` It reports one entry per contended axis — `nodeUniqueness` and `nodeDisjointness` carry the conflicting `owners` (each a `concrete_kind` / `node_id` pair, because ids are unique only per kind), `edgeCardinality` carries the conflicting `edgeIds`. It reads the nodes, edges and `uniques` relations, so it finds violations that predate the claim tables; it writes nothing, and it repairs nothing — choosing which claimant keeps the axis is a data-loss decision that stays with you. ### SQLite ↔ PostgreSQL parity The **query language is fully portable** between SQLite and PostgreSQL. Predicates (comparison, string/`ILIKE`, null, `between`, array, object, JSON-path), fixed and variable-length (recursive) traversals, bounded neighbor reads, per-edge-kind subgraph windows, one-statement query batches, aggregates (`count`/`sum`/`avg`/`min`/`max` with `groupBy`/`having`), set operations (`UNION`/`UNION ALL`/`INTERSECT`/`EXCEPT`, including traversal, subquery, `GROUP BY`/`HAVING`, and per-leaf `ORDER BY`/`LIMIT`/`OFFSET` leaves), ordering with `NULLS FIRST`/`LAST`, cursor pagination, temporal queries, and the fulltext query modes (`websearch`, `phrase`, `plain`, `raw`) all behave identically. A query you write against one backend compiles and runs the same way on the other. The remaining differences are **engine and runtime capability gaps** — they stem from what each database or hosted authorizer implements, not from TypeGraph choosing separate query semantics per backend: | Capability | SQLite | PostgreSQL | Behavior on the unsupported side | | ------------------------------------------------------ | ------------------------------------------------- | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------- | | Whole-graph temporary-table analytics | ✓ standard connections / ✗ D1 and Durable Objects | ✓ connection-based drivers / ✗ `neon-http` | Throws `UnsupportedBackendCapabilityError`; traversal algorithms with an inline engine fall back automatically | | Vector metric `inner_product` | ✗ | ✓ | Rejected at compile time on SQLite (`sqlite-vec`/`libsql-native` expose `cosine` + `l2`; `pgvector` adds `inner_product`) | | Vector index type `ivfflat` | ✗ | ✓ | Index declaration is **skipped** on SQLite (`indexTypes`: `hnsw`/`none` vs `hnsw`/`ivfflat`/`none`) | | Filtered approximate search **guarantees** a full page | ✓ `sqlite-vec` / ✗ `libsql-native` | ✗ (`pgvector` recovers, but is bounded) | Only `sqlite-vec` guarantees it; the others can return **fewer than `limit`** rows under heavy filtering — see below | | Per-query fulltext `language` override | ✗ | ✓ | Throws on SQLite — FTS5's tokenizer is fixed at table-create time; `tsvector` accepts a regconfig per query | | HNSW `efSearch` query tuning | ✗ | ✓ transactional HNSW drivers | Refused, never ignored: `UnsupportedBackendCapabilityError` with `details.capability` `vector.searchFrontierTuning` on **any** SQLite backend (vector and hybrid alike — neither `sqlite-vec`'s `vec0` KNN nor `libsql-native`'s DiskANN has a per-search frontier), and on transaction-less Postgres or a non-HNSW slot | | Bounded planner-statistics sampling | ✓ standard connections / ✗ D1 and Durable Objects | Native `ANALYZE` sampling | Restricted SQLite skips `analysis_limit` but still attempts scoped `ANALYZE`. Performance only — same results | | TypeGraph Identity Profile | ✓ transactional drivers | ✓ transactional drivers | Enabled graphs fail fast on non-atomic drivers; identity-disabled graphs retain their ordinary path | | Constraint claim relations (`capabilities.constraintClaims`) | ✓ | ✓ | Identical relations and identical statements on both dialects. A third-party backend that omits them declares `constraintClaims` absent and keeps the per-graph lock as its only fence | | Durable edge match identity (`capabilities.durableEdgeMatchIdentity`) | ✓ bundled adapters | ✓ bundled adapters | Both dialects persist the same canonical key and use a unique database arbiter. A custom backend must satisfy the full capability contract above or leave the capability absent | | Managed node projection fusion | ✓ registered atomic bulk programs; singleton fallback | ✓ registered atomic bulk programs; singleton create fusion | Eligible node bulk creates and resolved updates group fulltext/vector transitions into the same atomic program as their row mutations on both dialects. PostgreSQL additionally fuses an eligible singleton generated-ID create into one SQL statement when every active strategy supplies an inserted-node builder | | Managed node claim fusion (`capabilities.atomicNodeInsertClaims`) | ✗ portable transactional fallback | ✓ PostgreSQL/PGlite | SQLite keeps claim acquisition and insertion in the portable transaction. PostgreSQL transaction receivers fuse supported claim plans; a root non-transactional receiver is limited to exactly one generated-id, same-kind uniqueness claim with no other side effects | | Managed edge cardinality fusion | ✗ portable transactional fallback | ✓ PostgreSQL/PGlite transaction receivers | SQLite keeps its guarded claim and edge insert in the portable transaction. PostgreSQL can combine endpoint liveness, one cardinality claim, and the insert in one statement after any required graph lock | | Atomic SQL transport (`capabilities.execution.atomicBatch`) | ✓ on certified D1/libSQL roots; otherwise `none` | ✓ on bundled recognized PostgreSQL drivers, including neon-http | `root` means the exact backend owns the atomic boundary; `session` means the exact object is already bound to an open transaction and the outer transaction owns commit/rollback. Both require identity-keyed executor registration. Neon HTTP uses its native transaction batch; session-capable `pg`, postgres-js, neon-serverless, and PGlite drivers can execute programs on one pinned Drizzle transaction. Unrecognized drivers remain `none`. A custom backend must pass the framework-agnostic conformance runner before opting in; omitted support keeps the portable path | | Eligible registered managed writes | ✓ bundled SQLite roots, including D1 and libSQL | ✓ bundled PostgreSQL roots, including neon-http | Eligible singleton generated-ID nodes and `cardinality: "many"` edges use one authoritative create statement. Eligible node updates may carry fulltext/vector replacements; unconstrained non-durable-identity edge updates, direct edge deletes, and plain restricted node deletes use one authoritative read/gate plus one registered atomic mutation. Generated-, caller-, or mixed-ID node `bulkInsert`/`bulkCreate` batches compose supported multi-claim/cross-scope claim sets with projections in one schema-fenced native program; direct edge programs also maintain durable match identity and cardinality claims. Direct edge `bulkDelete` and plain restricted node `bulkDelete` use the same mutation profile. Eligible mixed `bulkUpsertById` sets, including node projections, use the profile on serverless roots and on exact bundled PostgreSQL transaction sessions; a generic derived backend still loses the evidence. A custom backend may opt in per family only after registering its exact transport and semantic executor. Unregistered or otherwise ineligible families, projected/identity-enabled node deletes, over-budget claimed members, cascade/disconnect deletes, and other managed writes retain the existing path | | Typed constraint error above READ COMMITTED | n/a (no such isolation mode) | ✗ at `REPEATABLE READ` / `SERIALIZABLE` | PostgreSQL raises `40001` from the claim's upsert instead of resolving the conflict, so the loser retries a serialization failure rather than reading `UniquenessError` | | Claim row lock released before end of transaction | ✗ | ✗ | Held to commit/rollback on both dialects, refusal included — a caller that catches a constraint error blocks other writers of that axis for the rest of its transaction | | Recursive traversal (`capabilities.recursiveTraversal`) | ✓ | ✓ | Identical on both bundled backends. A third-party backend declaring `{ supported: false, reason }` refuses the five recursion-dependent operations with `ConfigurationError` code `RECURSIVE_TRAVERSAL_UNSUPPORTED`; `weightedShortestPath` degrades to a predecessor walk instead — see above. Unweighted `shortestPath` is unaffected — it never emits a recursive CTE | | Write fence (`capabilities.writeFence`) | ✓ `engine-serialized` (single writer slot) | ✓ `lock` (advisory + table locks) | Identical guarantee, different mechanism. A custom backend that declares no `writeFence` resolves `unfenced` and is refused at construction for Operational Identity or TypeGraph-owned recorded-clock allocation | | Capability bundles (`CAPABILITY_BUNDLES`) | Identical | Identical | Both bundled backends implement every pilot bundle's core/extra members on both dialects it scopes to. A third-party backend with a port gap refuses (gated core, or a `refuse`-disposition extra) or degrades (a `fallback`-disposition extra) per that bundle's own registry row | | Engine-native lineage (`backend.lineage`) | ✗ (recorded-relations lineage via `history`) | ✗ (recorded-relations lineage via `history`) | Neither bundled profile declares its own `lineage`. `resolveLineage` derives it from the store's recorded relations whenever `history: true` is on, identically on both dialects, so a graph-merge diff against such a store is pruned the same way regardless of backend. Without `history`, no lineage source resolves and the diff is full; the anchor is the durable revision anchor when `revisionTracking: true`, otherwise the compatibility content fingerprint | | Engine-native recorded time (`backend.recordedTime`) | ✗ (TypeGraph-owned recorded relations via `history`) | ✗ (TypeGraph-owned recorded relations via `history`) | Neither bundled profile declares `recordedTime`, so `resolveRecordedTimeOwnership` derives `"typegraph-relations"` for both — `history: true` captures into TypeGraph's own recorded relations and clock, identically on both dialects, and every recorded-time integration suite and the parity snapshot run unchanged. A backend that supplies `recordedTime` (and the co-required `lineage`) reads and writes recorded time through its own engine instead: no TypeGraph capture, clock, or recorded relations, `RecordedInstant` anchors in the `e1:` form, and several TypeGraph-relation-specific surfaces refused — see [Engine-native recorded time](/queries/temporal#engine-native-recorded-time) and [Supplying `recordedTime`](/backend-authoring#supplying-recordedtime). No bundled backend implements this today; the capability is proven by a PostgreSQL-family simulation (shared by the always-running `tests/backends/postgres/pglite-engine-native-recorded-time.test.ts` and the `POSTGRES_URL`-gated `tests/backends/postgres/engine-native-recorded-time.test.ts`, both built on `engine-native-recorded-time-simulation.ts`) that dresses TypeGraph's own recorded relations as a temporal-table expression, labeled as a simulation rather than a real third engine | Identity support also has a **driver** dimension inside each dialect: | Driver | Atomic identity support | Behavior | | ------------------------------------------------------------ | ----------------------- | ------------------------------------------------------------------------------------------------------------------- | | Managed SQLite, libSQL, Durable Objects | ✓ | Full profile | | PostgreSQL `node-postgres`, `postgres-js`, neon-serverless, PGlite | ✓ | Full profile; identity-affecting writes serialize per graph, limiting each graph to one identity writer at a time | | Cloudflare D1 | ✗ | Enabled graphs fail at store construction with `ConfigurationError` details code `IDENTITY_REQUIRES_ATOMIC_BACKEND` | | `drizzle-orm/neon-http` | ✗ | Same fail-fast error; identity-disabled graphs retain the ordinary single-statement path | ### Filtered approximate search Every approximate (ANN) vector search carries at least one row filter: the liveness predicate that hides soft-deleted and out-of-validity rows. A `.where(...)` predicate narrows it further. Engines differ in where they apply that filter relative to the index traversal, which decides whether a page can come back short. Read it from `backend.capabilities.vector.filteredApproximateSearch`: ```typescript const filtered = backend.capabilities.vector?.filteredApproximateSearch; if (filtered?.guaranteesFullPage !== true) { // An approximate search here may return fewer than `limit` rows. } ``` **Check `guaranteesFullPage`, not `mode`.** `mode` names the mechanism the strategy asks the engine for; only `guaranteesFullPage` tells you whether a short page is possible. | `mode` | Strategy | `guaranteesFullPage` | Meaning | | ------------------- | --------------- | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `"filter-pushdown"` | `sqlite-vec` | `true` | The filter constrains the `vec0` KNN candidate set itself. `limit` matching rows come back whenever `limit` exist. | | `"iterative-scan"` | `pgvector` | `false` | The index is re-entered for more candidates (`hnsw.iterative_scan` / `ivfflat.iterative_scan`, applied automatically on pgvector ≥ 0.8). Much better recall than a post-filter, but **not a guarantee**: the scan stops at `hnsw.max_scan_tuples` / `ivfflat.max_probes`. And on **pgvector < 0.8** there is no iterative scan at all — the backend detects that, warns once, and the search stays `ef_search`-bounded. | | `"post-filter"` | `libsql-native` | `false` | DiskANN's `vector_top_k` is a table function with no filter pushdown and no way to re-enter the index. TypeGraph over-fetches `4 × (limit + offset)` neighbors and filters afterwards, so once more than that headroom is filtered out **the search silently returns fewer than `limit` rows while more matches exist**. | Heavy tombstone drift — routine in a temporal store — is what turns a bounded search from a theoretical caveat into a short page. When a full page matters, use an exact search (`approximate: false`), which scans and so applies the filter to every row; or declare the field's index as `"none"` so it is always brute-forced. Vector and fulltext capabilities are populated from the configured strategy, so the matrix above reflects the bundled strategies (`sqlite-vec`/`libsql-native`/`pgvector`, `fts5`/`tsvector`). A custom strategy advertising different `metrics`/`indexTypes`/`filteredApproximateSearch`/`searchFrontierTuning` shifts these rows accordingly — always check `backend.capabilities` at runtime rather than hard-coding the dialect. `searchFrontierTuning` is **required** on a vector strategy's capabilities, so a strategy must state whether it has a per-search ANN frontier knob rather than inheriting silence. It is a discriminated union: `{ tunable: true, parameter, indexType, requiresTransactionScope }` names the engine parameter `efSearch` maps to (`pgvector`: `hnsw.ef_search`, on an `hnsw` slot, needing a transaction to scope it), while `{ tunable: false, reason }` names why the engine has no such knob and is what makes `efSearch` a typed refusal there. A hand-written strategy that omits the field no longer compiles. Both bundled backends advertise `windowFunctions: true`. Relation `topPerPartition()` refuses execution with `UnsupportedBackendCapabilityError` when a custom backend sets `windowFunctions: false`. Vector, fulltext, and hybrid relevance-ranking queries use `ROW_NUMBER()` internally and throw `ConfigurationError` before SQL generation if a custom backend profile sets `windowFunctions: false` — there the window output *is* the result (the relevance k-cutoff / rank ordinal), so there is no correct fallback. `bulkFindByIndex({ limitPerInput })` also uses `ROW_NUMBER()` when available, but it does **not** throw on a windowless profile: the per-input cap is a transfer optimization with identical row semantics either way, so it degrades gracefully — fetching all matching ids and capping per group in application code. The unbounded `bulkFindByIndex` path needs no window and is always available. :::note[JSON is native on both backends] SQLite stores JSON as text and queries it with the built-in JSON functions (`json_extract`, `json_each`, …); PostgreSQL uses native `JSONB`. The dialect layer hides this difference, so JSON-path predicates and **B-tree expression indexes on scalar JSON properties** (`defineNodeIndex` / `defineEdgeIndex`) are at full parity. The one JSON-related difference is performance, not capability: PostgreSQL can use a single GIN index to accelerate array/object **containment** predicates (`contains()` / `containsAll()` / `hasKey()` / `pathEquals()`), whereas on SQLite those run as `json_each()` scans — correct results, just not index-accelerated. See [Indexes](/performance/indexes) for the full breakdown. ::: :::note[Transactions are driver-dependent, not backend-dependent] Both backends report `execution.interactiveTransactions: true` by default. The exception is symmetric and lives in specific drivers: Cloudflare D1 (SQLite) and `drizzle-orm/neon-http` (Postgres) are non-transactional, so they downgrade to `execution.interactiveTransactions: false`. Operations that require atomicity (`commitSchemaVersion`, `setActiveVersion`, Operational Identity) throw on those drivers regardless of backend. A schema-managed Store's write that cannot fuse its schema fence into its own statement fails closed the same way, because it has no other way to hold the transaction-scoped fence; `store.transaction()` refuses on those roots regardless. See [The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for which writes fuse. Eligible operations with a certified atomic SQL program remain available independently of this interactive transaction capability. ::: :::note[Aggregate set operations are a builder limitation, not a parity gap] `GROUP BY`/`HAVING` leaves are supported by the set-operation compiler on **both** backends, but the query builder does not expose `.union()`/`.intersect()`/`.except()` on `.aggregate()` queries. That limit applies equally to SQLite and PostgreSQL, so it is not a portability difference. ::: ## Connection Management Connection ownership follows the entrypoint: - **Managed Store factories** (`/sqlite/local` and `/postgres/pglite`) own the connection and provisioned resources. `await store.close()` releases them. - **Owned local backend factories** (`createLocalSqliteBackend` and `createLocalPgliteBackend`) also own their resources. A Store delegates `close()` to its backend, so `await store.close()` releases them. - **Bring-your-own adapter factories** (`createSqliteBackend`, `createPostgresBackend`, and `createLibsqlBackend`) leave connection ownership with the caller. Their Store's `close()` does not close the supplied client or pool. When you bring your own connection, you are responsible for: 1. **Creating connections** with appropriate configuration 2. **Connection pooling** for production use 3. **Closing connections** on shutdown ```typescript // You create the connection const sqlite = new Database("app.db"); const db = drizzle(sqlite); const backend = createSqliteBackend(db); const store = createStore(graph, backend); // You close the connection process.on("exit", () => { sqlite.close(); }); ``` Here `store.close()` leaves `sqlite` open because the application supplied the connection. Close the driver or pool through its own API. ### Serialized connections Some drivers run every statement through **one** connection. Two long-lived interchange streams cannot share such a connection — an export snapshot holds a read transaction for the whole stream while an import writes one per chunk — so TypeGraph refuses the second one with a typed error instead of letting it hang (see [Interchange serialized-connection guard codes](/errors#interchange-serialized-connection-guard-codes)). Recognizing a serialized connection means recognizing the *driver*, from the shape of the client object. That is deliberately conservative: a driver TypeGraph cannot positively identify is left unmarked, because refusing a pooled connection would refuse work that succeeds. | Driver / configuration | Detected | Notes | | --------------------------------------------------------------------------------------------- | ------------------ | ---------------------------------------------------------------------------------------------------------------- | | better-sqlite3, bun:sqlite, sql.js, local libSQL (`file:` / `:memory:`), Durable Object storage | ✓ automatic | One handle, one connection | | PGlite | ✓ automatic | One in-process WASM connection | | Bare `pg` / neon-serverless `Client`, a checked-out `PoolClient` | ✓ automatic | One owned socket | | `pg` `Pool` capped at one (`{ max: 1 }`, `{ max: "1" }`, `{ poolSize: "1" }`) | ✓ automatic | pg-pool does not coerce the cap, so the string forms are the same one-connection pool | | postgres-js capped at one (`{ max: 1 }`, `?max=1`, `PGMAX=1`) | ✓ automatic | Same reasoning on the postgres-js side | | Default-size pools, `neon-http`, D1, RDS Data API, remote libSQL (`http` / `ws`) | — deliberately not | Each statement gets an independent connection; refusing would refuse work that succeeds | | `expo-sqlite`, `op-sqlite`, `sqlite-proxy`, `pg-proxy`, a bespoke adapter | ✗ **declare it** | Serialized in fact, but the client exposes no shape TypeGraph can attribute to a known driver | | Bun `SQL` (Postgres) at `{ max: 1 }` | ✗ **declare it** | The cap is readable, but nothing identifies the driver, and a cap on an unknown client is not evidence | | postgres-js with a non-numeric string cap other than one, e.g. `?max=5` | ✗ **declare it** | Opens exactly one connection today only because postgres-js does not coerce the value — marking it would encode an upstream bug that will one day be fixed | For the rows marked **declare it**, tell TypeGraph what it cannot see. The option is on `createSqliteBackend` and `createPostgresBackend` — the two factories that resolve it. The batteries-included wrappers (`createLibsqlBackend`, `createLocalSqliteBackend`, `createLocalPgliteBackend`) do not take it, because each already detects its own connection. `{ mode: "shared", resource: pool }` is incorrect for a `pg.Pool` that can open multiple connections, even if several backends use that pool. Each transaction checks out its own connection; marking the pool as one resource makes independent snapshot exports and imports contend for a single lease and refuses concurrent operations that the pool can run. Leave the declaration absent for such a pool. ```typescript const sql = postgres(process.env.DATABASE_URL + "?max=5"); const backend = createPostgresBackend(drizzle(sql), { // This client really does run every statement on one connection. serializedResource: { mode: "shared", resource: sql }, }); ``` Two backends that name the **same** object are one serialized resource, exactly as two wrappers over a detected client are. Naming a *different* object than the one TypeGraph detected is refused with a `ConfigurationError` (`details.reason: "serialized-resource-conflict"`) rather than silently preferred: two wrappers over one connection given two different sentinels would stop being seen as a pair, which is the failure the guard exists to prevent. The refusal names each side by constructor (`details.declaredKind` / `details.detectedKind`) instead of carrying the two handles, because `details` is what `toLogString()` serializes and a driver handle there would log whatever that driver stores — a `pg.Pool` keeps its `connectionString`. The reverse declaration escapes a detection that is wrong for your topology: ```typescript const backend = createSqliteBackend(db, { serializedResource: { mode: "independent" }, }); ``` **Scope.** `{ mode: "independent" }` lifts the *shared-resource* refusal between two distinct backend objects. It does not lift the object-identity refusal, under which one SQLite backend exporting into **itself** is refused with `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`. That one is a fact about a single handle holding a single open snapshot transaction, not a claim about connection topology, so no declaration can make it false — pass a second backend instead. That surviving refusal is SQLite-only, so on PostgreSQL the declaration lifts the refusal for one backend exporting into itself as well: a client that hands out independent connections — which is exactly what the declaration claims — runs the snapshot and the writes it contends with on different ones. ## Database roles & least privilege `createStoreWithSchema()` and `createStore()` divide cleanly along DDL privilege, so a production deployment can run its application under a least-privilege, DML-only database role. - **`createStoreWithSchema(graph, backend)` runs DDL.** It bootstraps the base tables on a fresh database, applies safe auto-migrations, and adopts release-added deployment-wide base storage on pre-provisioned databases, even when the persisted graph schema is unchanged. The first adoption creates or repairs the graph-template and edge match-identity storage, then stamps a version marker; a warm base-schema check is one `SELECT` with no base-adoption DDL. It also durably materializes strategy-owned runtime storage — both fulltext and each `embedding()` field's per-`(kind, field)` vector table, plus a durable marker for each. It also brings TypeGraph's own base-relation **system indexes** up to the running library version: bootstrap DDL only runs on the very first boot, so an index shipped in a newer version reaches an already-initialized database through this step (built with `CREATE INDEX CONCURRENTLY` on PostgreSQL; a database whose indexes all exist settles from the catalog with no DDL). Contribution and system-index preparation have their own catalog checks and may still issue DDL, so the role it runs under **must hold `CREATE` / DDL privileges**. Run it once at startup, outside request handlers and transactions. (`store.evolve()` likewise provisions any embedding field it introduces, so it too needs DDL privileges.) Deployments that never run `createStoreWithSchema` (manual-DDL boot with a plain `createStore` attach) adopt new system indexes by calling `store.materializeSystemIndexes()` once under a DDL-capable role after upgrading; deployments that must not run index builds inline at boot (large tables behind a readiness probe) pass `systemIndexes: "skip"` to `createStoreWithSchema` and materialize out-of-band the same way. - **`createStore(graph, backend)` is a synchronous, zero-I/O attach.** It does not create tables, repair DDL, or record that runtime storage is materialized — it issues **no DDL ever**. Use it only to attach to a database a prior `createStoreWithSchema` boot already initialized. A fulltext operation or an **embedding write** against a database that was never initialized — a `create({ embedding })` or embedding update/delete — throws `StoreNotInitializedError` rather than silently emitting `CREATE TABLE` on the hot path. (Vector *reads* are not marker-gated: `store.search.vector`, `store.search.hybrid`, and a query-builder `.similarTo()` predicate compile to SQL against the per-field table directly, so on an un-provisioned database they surface the engine's own missing-relation error instead — `no such table: tg_vec_…` on SQLite, `relation … does not exist` on Postgres. Same cause, same fix; use `createVerifiedStore` to catch it at attach rather than at first query.) This is what lets a least-privilege role run vector ops: the table already exists. Graphs with no `searchable()` or `embedding()` fields are unaffected. The Store is also raw and unversioned: its writes do not participate in the schema-version fence. Direct backend writes have the same semantics. Quiesce those writers yourself before changing schemas. - **`createVerifiedStore(graph, backend)` is the same zero-DDL attach with a verification gate.** It reads the active schema row, folds the persisted graph extension, and refuses to construct the Store unless the database is at the same schema version as the code graph. Throws `BaseSchemaMigrationError` when deployment-wide base storage is missing, stale, or newer than the library, `MigrationError` on graph-schema drift (safe or breaking), `ConfigurationError` when no graph schema has been initialized, and `StoreNotInitializedError` when the schema is current but runtime-contribution markers are missing. The runtime-side counterpart of `createStoreWithSchema` for least-privilege deployments. If you only need the gate without building a Store (e.g. a readiness probe), call `assertSchemaCurrent`. Its managed writes require a transactional backend with the schema-write fence; non-transactional and unsupported custom backends can attach for reads but fail closed on the first write. The adapter equivalents (`createAdapterStoreWithSchema` and `createVerifiedAdapterStore`) carry the same managed metadata. So does `createAdapterStore(..., { reconciled })` with a cached reconciliation snapshot, and Stores returned by `evolve()` or rebound from an already-managed Store. Check `store.introspect().schemaVersion !== undefined` at runtime. Calling `store.clear()` deletes the schema rows and resets that Store to raw semantics; reopen it through a managed factory before resuming version-fenced writes. - **`store.verifyContributions()` diagnoses contribution storage; `store.repairContributions()` repairs safe findings under a privileged role.** Every gate above trusts the marker row without probing the catalog, so a database whose strategy-owned tables were dropped out of band opens clean and fails at the first dependent read or write. This method compares each contribution currently expected by the active graph and backend strategies with its marker and the catalog. It does not audit retired marker rows, and a never-attempted contribution with neither marker nor table is omitted, so an empty result is not initialization proof. It is read-only (`SELECT` only, no DDL) so the least-privilege role can run it, and it is deliberately not part of any open path. For a readiness check, construct the Store with `createVerifiedStore()` first and then run this diagnostic; otherwise use it as an operator check. The repair method re-audits current declarations, preserves data while repairing `missing-marker` and `failed-materialization`, and reports `stale` or `orphaned-marker` as `requires-rebuild`. Run repair through the DDL-capable migration role, not the least-privilege runtime role. Follow the per-state table in [The store opens clean but a fulltext or vector read fails](/troubleshooting#the-store-opens-clean-but-a-fulltext-or-vector-read-fails) rather than applying one repair to every entry. - **`store.probeContributions()` is the read-only readiness check; `store.rebuildContribution()` is the destructive last resort.** The two bracket `repairContributions()` into one escalation ladder: probe (writes nothing) → repair (non-destructive) → rebuild (destructive, but scoped to the calling graph). The probe reports one `ready` / `degraded` entry per search projection and is safe on a read path, on a replica, and under the least-privilege role — it shares the detection logic of the other two rather than reimplementing it, so it cannot disagree with the gate the hot path actually consults. The rebuild is the only repair for a `stale` contribution, whose table exists at a shape the current `createDdl` no longer produces; it deletes and refills only the calling graph's rows in the shared fulltext table, escalating to drop → recreate when that table holds no other graph's rows (under a database-scoped DDL advisory lock, since that DDL is database-global), and runs the whole sequence inside one transaction under the schema-write fence. It refuses with `ContributionRebuildUnsupportedError` for vector storage, whose embeddings exist only in the table it would drop (`reason: "vector-source-unavailable"`), and for a `stale` shape whose storage still holds other graphs' rows (`reason: "shared-storage-in-use"`, naming them in `details.otherGraphIds`). Run rebuilds through the DDL-capable migration role, in a maintenance window: the transaction is held for the whole refill, and on PostgreSQL a drop's `ACCESS EXCLUSIVE` lock blocks both searches and writes to any kind with `searchable()` fields until it commits. Reach it from a `createStore()` Store — the managed factory's boot step refuses to open while a contribution is `stale`. See [Contribution health: probe, repair, rebuild](/troubleshooting#contribution-health-probe-repair-rebuild). Strategy contributions declare an ownership `scope`: `"graph"` (the default for older custom strategies) provisions one physical contribution per graph, while `"deployment"` provisions shared physical storage once under TypeGraph's reserved deployment marker and records a separate graph-local activation marker. The built-in full-text strategies use deployment scope; vector slots remain graph-scoped. A subsequent graph open reads both attestations and performs no DDL, so it can run under a DML-only role without disabling full-text search. The reserved marker key is exported as `DEPLOYMENT_CONTRIBUTION_GRAPH_ID`; graph definitions must not use that id. ### Contribution capability parity `backend.capabilities.contributions` declares how far up the ladder a backend goes. Each rung is separate because a backend can genuinely stop at any of them, and a rung a backend cannot serve refuses with a typed error rather than returning something that looks like success. | Backend | `supported` | `probe` | `rebuild` | | --- | --- | --- | --- | | SQLite (better-sqlite3, bun:sqlite, libSQL, Durable Objects) | ✅ | ✅ | ✅ | | SQLite with `transactionMode: "none"` | ✅ | ✅ | ❌ no schema fence | | PostgreSQL (`pg`, `postgres-js`, PGlite, `neon-serverless`) | ✅ | ✅ | ✅ | | PostgreSQL over `neon-http` | ✅ | ✅ | ❌ no schema fence | | Custom fulltext strategy without `dropDdl` | ✅ | ✅ | ❌ no teardown DDL | | Fulltext disabled (`fulltext: false`) | ✅ | ✅ | ✅ with a schema fence | `rebuild` requires two things at once: a fulltext strategy that declares `dropDdl` on its contribution, and a transactional schema fence (`schemaWriteTransaction`) to run the sequence under. The HTTP-only PostgreSQL drivers cannot hold a session across statements, so they have no fence — the same reason they already report `capabilities.execution.interactiveTransactions === false`. A third-party strategy predating `dropDdl` keeps working for every other operation and is reported as not rebuildable rather than being dropped through a synthesized statement TypeGraph guessed at. Vector contributions are never rebuildable on any backend; that is a property of what TypeGraph stores, not of the engine. A backend built with `fulltext: false` has no fulltext contribution at all, so the first condition is vacuously satisfied and `rebuild` reduces to whether the backend has the transactional schema fence — the same value it would report if fulltext were still active on a driver with that fence. **`fulltext: false` stops creating and maintaining the fulltext table; it never drops one.** On a database that already carries fulltext rows, disabling fulltext leaves them in place and unmaintained: a hard delete performed while fulltext is off leaves an orphaned row behind in the fulltext table, because `hardDeleteNode`'s cascade has no active strategy to build a delete statement from. Re-enabling fulltext later therefore requires the destructive contribution rebuild — `store.rebuildContribution("fulltext")`, which drops and recreates the fulltext table — **not** `store.search.rebuildFulltext()`: that method pages live nodes to recompute their content, and a hard-deleted node has no row left in the node table for it to page, so it never revisits, and therefore never clears, the orphan. ### Recommended deployment shape Run schema/DDL changes as a **privileged, one-time migration step**, then run the application under a **least-privilege runtime role** that holds only `SELECT` / `INSERT` / `UPDATE` / `DELETE`: ```typescript // 1. Migration step — privileged role with DDL/CREATE. // // createStoreWithSchema is mandatory here: it bootstraps tables, // applies safe auto-migrations, commits the schema_versions row, // and writes the durable contribution markers. The runtime gate // checks all of those. const [/* store */] = await createStoreWithSchema(graph, adminBackend); // Optional prerequisite if you manage DDL externally with // drizzle-kit. Generated SQL creates the tables but does NOT // initialize the schema row or contribution markers — still run // createStoreWithSchema afterwards (it skips bootstrap when tables // already exist and commits the row + markers): // // import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; // await adminPool.query(generatePostgresMigrationSQL()); // await createStoreWithSchema(graph, adminBackend); ``` ```typescript // 2. Runtime — least-privilege, DML-only role. Zero DDL. // createVerifiedStore fails fast if the privileged migrator is behind. const runtimePool = new Pool({ connectionString: process.env.APP_DATABASE_URL }); const backend = createPostgresBackend(drizzle(runtimePool)); const [store] = await createVerifiedStore(graph, backend); ``` If the runtime role has no DDL privileges and you boot it with `createStoreWithSchema()` anyway, the first cold boot fails with a permission error on the bootstrap or contribution-marker DDL — see [Troubleshooting](/troubleshooting). ## Environment-Specific Setup ### Development ```typescript // In-memory for fast tests const { backend } = createLocalSqliteBackend(); // Or file-based for persistence during development const { backend } = createLocalSqliteBackend({ path: "./dev.db" }); ``` ### Testing ```typescript // Fresh in-memory database per test beforeEach(() => { const { backend } = createLocalSqliteBackend(); store = createStore(graph, backend); }); ``` ### Production Single-role setup — `createStoreWithSchema` bootstraps and migrates on boot, so the role needs DDL privileges. To run the application under a least-privilege, DML-only role instead, split the migration step out as described in [Database roles & least privilege](#database-roles--least-privilege). ```typescript // PostgreSQL with pooling const pool = new Pool({ connectionString: process.env.DATABASE_URL, max: 20, ssl: { rejectUnauthorized: false }, // For managed databases }); const db = drizzle(pool); const backend = createPostgresBackend(db); const [store] = await createStoreWithSchema(graph, backend); ``` ## Next Steps - [Schemas & Types](/core-concepts) - Define your graph schema - [Semantic Search](/semantic-search) - Vector embeddings and similarity search - [Limitations](/limitations) - Backend-specific constraints # Query Builder Overview > A fluent, type-safe API for querying your graph TypeGraph provides a fluent, type-safe query builder for traversing and filtering your graph. This page introduces the query categories and how they compose together. ## Query Categories Every query builder method falls into one of these categories: | Category | Purpose | Key Methods | |----------|---------|-------------| | [Source](/queries/source) | Entry point - where to start | `from()` | | [Filter](/queries/filter) | Reduce the result set | `whereNode()`, `whereEdge()` | | [Traverse](/queries/traverse) | Navigate relationships | `traverse()`, `optionalTraverse()`, `to()` | | [Recursive](/queries/recursive) | Variable-length paths | `recursive()` | | [Shape](/queries/shape) | Transform output structure | `select()`, `project()`, `map()`, `aggregate()` | | [Expressions](/queries/expressions) | Typed database calculations | `expr`, `project()`, expression callbacks | | [Aggregate](/queries/aggregate) | Summarize data | `groupBy()`, `count()`, `sum()`, `avg()` | | [Order](/queries/order) | Control result ordering/size | `orderBy()`, `limit()`, `offset()` | | [Temporal](/queries/temporal) | Time-based queries | `temporal()` | | [Compose](/queries/compose) | Reusable query parts | `pipe()`, `createFragment()` | | [Combine](/queries/combine) | Set operations | `union()`, `intersect()`, `except()` | | [Execute](/queries/execute) | Run and retrieve | `execute()`, `first()`, `count()`, `exists()`, `paginate()`, `stream()`, `batch()` | ## Query Flow A typical query follows this flow: ```text Source → Filter → Traverse → Filter → Shape → Order → Execute ↑__________________| (repeat as needed) ``` Each step is optional except Source and Execute. You can filter, traverse, and filter again as many times as needed before shaping and executing. ## Basic Example ```typescript const results = await store .query() .from("Person", "p") // Source .whereNode("p", (p) => p.status.eq("active")) // Filter .traverse("worksAt", "e") // Traverse .to("Company", "c") // Traverse (target) .whereNode("c", (c) => c.industry.eq("Tech")) // Filter .select((ctx) => ({ // Shape person: ctx.p.name, company: ctx.c.name, role: ctx.e.role, })) .orderBy("p", "name", "asc") // Order .limit(50) // Order .execute(); // Execute ``` ## Type Safety The query builder is fully typed. TypeScript infers result types based on your schema and selection: ```typescript // TypeScript infers: Array<{ name: string; email: string | undefined }> const results = await store .query() .from("Person", "p") .select((ctx) => ({ name: ctx.p.name, // string (required in schema) email: ctx.p.email, // string | undefined (optional in schema) })) .execute(); // Invalid property access is caught at compile time: .select((ctx) => ({ invalid: ctx.p.nonexistent, // TypeScript error! })) ``` For new database-side projections and calculations, use typed [database expressions](/queries/expressions). `project()` compiles its callback to SQL, while `map()` transforms decoded rows in JavaScript. Existing `select()` callbacks retain their compatibility behavior. ## When to Use Queries vs Store API **Use the query builder** when you need: - Filtering based on node properties - Traversing relationships between nodes - Aggregating data across multiple nodes - Complex predicates with AND/OR logic **Use the [Store API](/schemas-stores#store-api)** for simple operations: - Get a node by ID - Create a new node - Update a node's properties - Delete a node ## Predicates Reference Predicates are the building blocks for filtering. Each data type has its own set of predicates: | Type | Documentation | |------|--------------| | String | [String Predicates](/queries/predicates/#string) | | Number | [Number Predicates](/queries/predicates/#number) | | Date | [Date Predicates](/queries/predicates/#date) | | Array | [Array Predicates](/queries/predicates/#array) | | Object | [Object Predicates](/queries/predicates/#object) | | Embedding | [Embedding Predicates](/queries/predicates/#embedding) | ## Performance Tips ### Filter Early Apply predicates as early as possible to reduce the working set: ```typescript // Good: Filter at source store .query() .from("Person", "p") .whereNode("p", (p) => p.active.eq(true)) .traverse("worksAt", "e") .to("Company", "c"); // Less efficient: Filter after traversal store .query() .from("Person", "p") .traverse("worksAt", "e") .to("Company", "c") .whereNode("p", (p) => p.active.eq(true)); ``` ### Be Specific with Kinds Unless you need subclass expansion, use exact kinds: ```typescript // More efficient: Exact kind .from("Podcast", "p") // Less efficient: Includes all subclasses .from("Media", "m", { includeSubClasses: true }) ``` ### Always Paginate Large Results ```typescript const page = await store .query() .from("Event", "e") .orderBy("e", "date", "desc") .limit(100) .select((ctx) => ctx.e) .execute(); ``` ## Next Steps Start with the fundamentals: 1. [Source](/queries/source) - Starting queries with `from()` 2. [Filter](/queries/filter) - Reducing results with predicates 3. [Traverse](/queries/traverse) - Navigating relationships 4. [Shape](/queries/shape) - Transforming output with `select()` # Temporal > Time-based queries with temporal() TypeGraph tracks temporal validity for all nodes and edges. Use temporal queries to view the graph at a point in time, audit changes, or access historical data. ## Temporal Modes The `temporal()` method controls which versions of data are returned: | Mode | Description | |------|-------------| | `"current"` | Only currently valid data (default behavior) | | `"asOf"` | Data as it existed at a specific timestamp | | `"includeEnded"` | All versions, including historical | | `"includeTombstones"` | All versions, including soft-deleted | ## Current State (Default) By default, queries return only currently valid, non-deleted data: ```typescript // Returns only current, non-deleted nodes const currentPeople = await store .query() .from("Person", "p") .select((ctx) => ctx.p) .execute(); ``` This is equivalent to: ```typescript .temporal("current") ``` ## Point-in-Time Queries (asOf) Query the graph as it existed at a specific moment: ```typescript const yesterday = new Date(Date.now() - 24 * 60 * 60 * 1000).toISOString(); const pastState = await store .query() .from("Article", "a") .temporal("asOf", yesterday) .whereNode("a", (a) => a.id.eq(articleId)) .select((ctx) => ctx.a) .execute(); ``` This returns nodes and edges that were valid at the specified timestamp, even if they've since been updated or deleted. ### Use Cases for asOf - **Auditing**: See what data looked like at a specific time - **Debugging**: Reproduce issues by querying historical state - **Compliance**: Generate point-in-time reports - **Recovery**: Find old values before an erroneous update ```typescript // What did the user's profile look like last week? const lastWeek = new Date(Date.now() - 7 * 24 * 60 * 60 * 1000).toISOString(); const historicalProfile = await store .query() .from("User", "u") .temporal("asOf", lastWeek) .whereNode("u", (u) => u.id.eq(userId)) .select((ctx) => ctx.u) .first(); ``` ## Shared-Coordinate Views (store.asOf) `.temporal("asOf", T)` pins a single query. When several reads should share one temporal coordinate, pin it once with `store.asOf(T)` and reuse the returned **read-only view** — TypeGraph's as-of database value, in the style of Datomic `(d/as-of db t)` and SQL:2011 `FOR SYSTEM_TIME AS OF`. ```typescript const past = store.asOf("2024-01-01T00:00:00.000Z"); // Every read on `past` observes the graph as it was valid at that instant. const alice = await past.nodes.Person.getById(aliceId); const jobs = await past.edges.worksAt.findFrom(alice); const peers = await past.reachable(aliceId, { edges: ["knows"] }); const team = await past.subgraph(aliceId, { edges: ["reportsTo"] }); const names = await past .query() .from("Person", "p") .whereNode("p", (p) => p.department.eq("Engineering")) .select((ctx) => ctx.p.name) .execute(); ``` The view pins the `nodes` / `edges` collections (`getById`, `getByIds`, `find`, `count`, `findFrom`, `findTo`), `query()`, `subgraph()`, and the graph algorithms (`reachable`, `canReach`, `shortestPath`, `neighbors`, `degree`). For the other modes, use `store.view({ mode, asOf })`: ```typescript // A view over every version, including soft-deleted ones. const audit = store.view({ mode: "includeTombstones" }); const everyVersion = await audit.nodes.Document.find(); ``` A view is **read-only**: writes stay on the live `store`, and a view collection rejects `create` / `update` / `delete` with a `ConfigurationError`. `search` is refused on a non-`"current"` view (the fulltext / vector index reflects current state only). `asOf` must be a canonical UTC ISO-8601 timestamp (`YYYY-MM-DDTHH:mm:ss.sssZ`). See the [`store.asOf` / `store.view` reference](/schemas-stores#temporal-views-storeasof-and-storeview) for the full surface. ## Recorded Time (Bitemporal) The modes above query **valid time** — *when a fact was true in the world* (`validFrom` / `validTo`). Recorded time (also called **system time**) is the second axis — *when a fact was recorded by TypeGraph*. With the built-in captured relation, TypeGraph can run **bitemporal graph reads** for TypeGraph-managed writes: you can ask "what did TypeGraph reconstruct as true, as of a captured commit instant?" — including seeing values that were later corrected. Recorded-time capture is **opt-in** per store, because it writes a history row for every committed TypeGraph collection change: ```typescript const store = createStore(graph, backend, { history: true }); ``` With `history: true`, every committed TypeGraph node/edge write is captured into recorded-time relations (`typegraph_recorded_nodes` / `typegraph_recorded_edges`) stamped with a per-graph monotonic commit instant. Enable it on a **fresh graph**: there is no backfill, so an entity that already exists is first recorded the next time it is written through TypeGraph. Capture requires a transactional backend with statement execution (the built-in SQLite / PostgreSQL backends). Advanced hosts can bind an already-populated recorded relation for reads without using TypeGraph's writer wrapper: ```typescript import { createSqlSchema, recordedRelation } from "@nicia-ai/typegraph"; const recordedRead = recordedRelation({ schema: createSqlSchema({ recordedNodes: "audit_nodes", recordedEdges: "audit_edges", }), }); const store = createStore(graph, backend, { recordedRead }); ``` That option only supplies the read source for `asOfRecorded(T)` reconstruction. It does not capture writes, advance TypeGraph's recorded clock, or make `store.recordedNow()` available. If TypeGraph should own capture, use `history: true`. `recordedRead` must be created by `recordedRelation({ schema })` with a `createSqlSchema(...)` schema; the store validates those factory descriptors at runtime and rejects combining them with `history: true`. ### Reading at a recorded instant `store.asOfRecorded(T)` reconstructs the graph as TypeGraph recorded it at instant `T`. `T` is a `RecordedInstant`: a branded, versioned string containing both a per-graph logical revision and a physical wall-time high-water mark. It originates from `store.recordedNow()` (below), or from `asRecordedInstant(...)` when an anchor previously returned by TypeGraph has round-tripped through untyped storage: ```typescript import { asRecordedInstant, recordedInstantWallTime, } from "@nicia-ai/typegraph"; const recorded = store.asOfRecorded( asRecordedInstant("r1:0000000000000042:2024-06-01T12:00:00.000Z"), ); const doc = await recorded.nodes.Document.getById(docId); const cited = await recorded.edges.cites.getByIds(citationIds); const reachable = await recorded.reachable(docId, { edges: ["cites"] }); ``` Recorded collections also expose `scan()` for complete snapshot reconstruction. Each call returns at most 1,000 entities in canonical `id` order; use the opaque `nextCursor` to continue without retaining a separate identity inventory: ```typescript const first = await recorded.nodes.Document.scan({ limit: 500 }); const second = first.nextCursor === undefined ? undefined : await recorded.nodes.Document.scan({ limit: 500, after: first.nextCursor, }); const citations = await recorded.edges.cites.scan({ limit: 500 }); ``` Scan cursors are forward-only and bound to the graph, entity kind, and both temporal coordinates. Passing a cursor to another collection or recorded-time view throws a `ValidationError` instead of silently skipping data. Iterate each declared node and edge kind to reconstruct a complete historical graph snapshot. A raw wall-clock string — `store.asOfRecorded(new Date().toISOString())` — does **not** type-check, by design. Wall time does not identify which commit to read when several commits share a millisecond. The anchor's logical revision provides that order; its ISO component records a non-decreasing physical wall-time high-water mark. To pin "as things stand right now" deterministically, use `store.recordedNow()` (the recorded high-water mark), then guard the `undefined` case before passing it to `store.asOfRecorded()`. ```typescript await store.nodes.Document.update(docId, { title: "Revised" }); const checkpoint = await store.recordedNow(); // a stable anchor for this state if (checkpoint === undefined) throw new Error("expected a recorded checkpoint"); console.log(recordedInstantWallTime(checkpoint)); // canonical UTC wall time // ...later, however much the graph has changed: const asOfCheckpoint = store.asOfRecorded(checkpoint); ``` `recordedNow()` is **graph-global**, not scoped to any one caller or write. It is the single high-water mark for the whole graph, advanced by *every* committed capture from *any* writer. So a change in `recordedNow()` across two reads means "something committed to this graph in between" — **not** "the write I just made landed." Do not use a `recordedNow()` advance as a per-writer "did my write succeed?" signal: under any concurrent writer to the same graph it both misses dropped writes (another writer moved the clock) and misfires on no-op writes. To confirm a specific write committed, observe the write itself (e.g. its return value, or run it inside `store.transaction(...)` and act on success), not the global clock. #### Logical revision and physical time The canonical encoding is `r1:<16-digit revision>:`. Revisions are strict and monotonic within one graph. The physical component is sampled from the application clock and clamped to the previous anchor only when that clock moves backward. It may repeat, but never decreases. TypeGraph does not add one millisecond per commit, so throughput cannot push recorded wall time beyond the greatest wall time the graph has actually observed. After a backward clock correction, the component remains at its prior high-water mark until wall time catches up. This non-decreasing physical component preserves cumulative diagonal replay for default validity timestamps: a later recorded anchor cannot pin valid time before an earlier commit's default `valid_from`. The fixed-width revision prefix makes anchors lexicographically sortable within a graph and gives each captured transaction a distinct addressable state. Use `compareRecordedInstants(a, b)` rather than manually comparing strings, and only compare anchors from the same graph. Recorded relations store the revision as an integer, so their open interval ceiling is independent of the `r1` API encoding and PostgreSQL range scans do not depend on text collation. Recorded clocks remain per graph, and TypeGraph does not provide one cross-graph recorded anchor. Batch related writes in `store.transaction(...)`: one transaction allocates one recorded instant. For event logs, align transactions with durable replay or checkpoint boundaries, and cap transaction size separately so an initial sync does not hold a write lock or capture buffer without bound. Direct `store.asOfRecorded(T)` is **diagonal** bitemporal sugar: it uses the anchor's logical revision for the recorded-time axis and its physical wall-time component for the valid-time axis. To pin the two axes independently — *what was valid at one instant, as TypeGraph captured it at another* — chain from a valid-time view: ```typescript // The state valid on Jan 1, as TypeGraph recorded it on Jun 1 // (e.g. after a correction was entered later). const corrected = store .asOf("2024-01-01T00:00:00.000Z") .asOfRecorded( asRecordedInstant("r1:0000000000000042:2024-06-01T12:00:00.000Z"), ); const asKnownThen = await corrected.nodes.Invoice.getById(invoiceId); ``` Use `recordedInstantRevision(T)` for diagnostics and `recordedInstantWallTime(T)` for display or logging. Do not split the versioned anchor string manually. `store.view({ mode }).asOfRecorded(T)` composes recorded time with any valid-time mode — e.g. `includeTombstones` to reconstruct soft-deleted rows at a recorded instant. ### The recorded view surface A `RecordedStoreView` is a **narrow, reconstructing** read lens. It exposes only reads that can be faithfully rebuilt from the recorded relations: - **Point reads** — `nodes..getById` / `getByIds`, and the edge equivalents - **`query()`** — a sealed query builder over the recorded relations - **`subgraph()`** and the graph algorithms — `reachable`, `canReach`, `shortestPath`, `degree` Broad collection reads (`find` / `count` / `findFrom` / …), `search`, and fulltext / vector predicates are **refused** with a `ConfigurationError` / `UnsupportedPredicateError`: the fulltext and vector indexes reflect *current* state only, so they cannot answer a recorded-time question. `T` must use the canonical versioned RecordedInstant encoding; a plain ISO timestamp is rejected. :::caution[Preview-schema migration] Timestamp-only anchors and recorded tables created by the initial preview need an explicit offline migration. Run `migrateLegacyRecordedTime({ backend })` before opening the upgraded store, then translate externally persisted checkpoints with `migrateRecordedAnchor({ backend, graphId, anchor })`. See [Migrating preview recorded time](/schema-management#migrating-preview-recorded-time). ::: ### Engine-native recorded time Everything above describes **TypeGraph-owned** recorded time: `history: true` captures into TypeGraph's own recorded relations and clock. A backend can instead track recorded time itself — declare `EngineProvisioning.recordedTime` on it — and `history: true` then reads and writes through the engine's own temporal storage; TypeGraph's capture relations, clock, and write-fence-gated clock allocation are never engaged. Which ownership a store reads under is **derived**, never an option you set: it is `"engine-native"` exactly when the backend declares `recordedTime`, `"typegraph-relations"` otherwise. Neither bundled SQLite nor PostgreSQL profile declares it, so every example on this page runs under `"typegraph-relations"` as shown; see [Supplying `recordedTime`](/backend-authoring#supplying-recordedtime) for what a third-party engine implements to opt in. Under engine-native ownership: - `store.recordedNow()`, `store.revisionNow()`, and `TransactionReceipt.recorded` all come from the engine's own revision instead of TypeGraph's clock — one call per transaction, not per graph. `TransactionReceipt.recorded` is stamped only when a graph node/edge/identity write inside the transaction actually changed a row — a delete of a missing id, an `insertNodeIfAbsent` that found the row, and a coalesced no-op upsert all leave it `undefined`, matching a read-only transaction. A transaction whose only effect is a raw `tx.sql` statement also leaves it `undefined` even though the engine's revision advances underneath it; use a graph collection write when you need `receipt.recorded` to reflect the change. To observe those writes, a receipted engine-native transaction routes every write through an observing wrapper, so `transactionWithReceipt` does not use session-scoped atomic batching where a plain `transaction` on the same store would. - `RecordedInstant` anchors use the engine form `e1::` rather than `r1:<16-digit revision>:`. The revision is an opaque, engine-assigned token, never parsed as a number, so ordering two `e1:` anchors (`compareRecordedInstants`) falls back to the timestamp component only — two engine revisions minted within the same millisecond compare equal even though they are distinct commits, unlike a `r1:` anchor's strict per-commit counter. `recordedInstantWallTime(instant)` works for either form; `recordedInstantRevision(instant)` throws for an `e1:` anchor, since there is no TypeGraph numeric revision to return. - `store.asOfRecorded(instant)` requires an instant minted under the SAME store's own ownership form. An engine-native store refuses an `r1:` instant, and a TypeGraph-owned store refuses an `e1:` instant, both with a `ConfigurationError` (`RECORDED_INSTANT_OWNERSHIP_MISMATCH`) — an anchor from one ownership form is never valid against the other, even against a different store over the same data. - No recorded relation is read or written. (A profile built on the bundled schema factories still creates the recorded tables as part of its base DDL — they just stay empty.) `recordedRead: recordedRelation({ schema })` (above) and `migrateLegacyRecordedTime` are both refused: neither has a TypeGraph-owned recorded relation to bind or migrate. - `revisionTracking: true` is refused whether or not `history: true` is also requested — there is no TypeGraph clock for it to advance; the engine's own revision is the only tracking engine-native has, and it is available only under `history: true`. - Reconstructing identity at a recorded coordinate — `store.identityAtCoordinate` at a past instant, and any query that reaches the historical identity traversal — is refused: identity history reads TypeGraph's own recorded relations directly, which an engine-native backend does not populate. Read identity at the current coordinate instead, or use a TypeGraph-owned store for historical identity reconstruction. Everything else on this page — `asOfRecorded`'s diagonal composition with `asOf`, the recorded view surface's read shape, `includeTombstones` composition — behaves the same under either ownership form; only the anchor grammar, the write mechanics, and the refusals above differ. See [Lineage and pruned diffs](/graph-merge#lineage-and-pruned-diffs) for how graph-merge derives a change delta under engine-native ownership — from the engine's own `lineage`, never from recorded relations, since none exist to derive one from. ### Writing with history enabled Capture flushes at transaction commit, so writes must go through the store's typed collections — use `store.transaction(...)` as usual: ```typescript await store.transaction(async (tx) => { await tx.nodes.Document.create({ title: "Draft" }); }); ``` #### Raw SQL under history capture The portable `HistoryStore` exposes neither raw SQL nor caller-owned transaction adoption. If the store was deliberately created through `createAdapterStore(..., { history: true })`, raw `tx.sql` is still disabled (it would bypass capture), and `store.withTransaction(externalTx)` is replaced by the callback form `store.withRecordedTransaction(externalTx, async (tx) => { ... })`, which gives capture a flush point before your transaction commits. Out-of-band database writes and row-returning raw SQL paths are not audited by the built-in capture wrapper; use TypeGraph collection writes when the recorded relation is the source of truth. The adapter history store's `.backend` is a runtime and type-level `HistoryStoreBackend` projection. Capture-wrapped graph reads and writes remain available. `executeRaw`, `executeStatement`, `executeDdl`, `trustedImport`, `clearGraph`, and nested `transaction` are absent because each can mutate live rows without a corresponding capture flush. The full guarded backend remains internal to TypeGraph's query and transaction implementation. `store.withTransaction` on a history-enabled store is a **compile error** (the `externalTx` argument is rejected with a message naming `withRecordedTransaction`); the runtime guard still throws `ConfigurationError` if suppressed. Inside an `AdapterHistoryStore.transaction(...)`, the typed context omits `tx.sql`, and `tx.sqlAvailability` reports `"history"` (or `"revisionTracking"`) so portable code can branch without touching the runtime guard. Suppressed JavaScript or TypeScript access still throws — see the `tx.sqlAvailability` guidance in [Cross-Store Transactions](/recipes/). Both guards carry a branchable `details.code`; see [Recorded-capture guard codes](/errors/#recorded-capture-guard-codes). To write your own relational tables atomically with graph writes on a history store, pass your transaction handle to `withRecordedTransaction` and write your tables through **that** handle (not `tx.sql`): ```typescript await db.transaction(async (pgTx) => { const { receipt } = await store.withRecordedTransaction(pgTx, async (tx) => { await tx.nodes.Document.update(documentId, props); // graph write }); await pgTx.insert(streamCursors).values(cursorRow); // your own table }); // one COMMIT / ROLLBACK across both layers ``` `withRecordedTransaction` returns a [`TransactionOutcome`](/schemas-stores/#transaction-receipts): destructure `{ result, receipt }`. `receipt.writes` counts the graph writes (drop detection) and `receipt.recorded` is this transaction's recorded commit instant — the per-transaction replay anchor. When the callback runs user code that also bookkeeps, scope a sub-receipt with `tx.measure((scoped) => ...)`: writes through the `scoped` context are attributed to the sub-receipt, while the surrounding bookkeeping written through `tx` is not. This is separate from `recordedRead`: a store created with a `recordedRead` binding can reconstruct from a relation populated by another system, but TypeGraph is not responsible for making that relation complete or atomic with live writes. #### Write cost: batch under `history: true` Each **un-batched** write under `history: true` becomes its own transaction — it allocates a recorded commit instant under a per-graph clock lock and flushes one history row at commit. So a tight loop of single `create`/`update`/`delete` calls pays that fixed cost once per call. Wrapping the same writes in one `store.transaction(...)` allocates **one** recorded instant for the whole batch and amortizes the overhead to roughly nothing. Measured per-op latency, identical workload with capture off vs on (history off → on; N = 400; reproduce with `pnpm --filter @nicia-ai/typegraph-benchmarks bench:recorded-write`): | Workload | SQLite | PostgreSQL | | ------------------------------ | -----: | ---------: | | create — un-batched (per op) | ~2.5× | ~5.5× | | create — **batched in one txn** | ~1.5× | ~1.0× | | update — un-batched (per op) | ~2.8× | ~6× | | soft delete — un-batched | ~1.7× | ~1.9× | The takeaway: capture is opt-in and cheap when you batch. Under `history: true`, prefer `store.transaction(...)` for bulk writes; a loop of individual collection writes is the one pattern that pays the per-write multiple. (Stores created without `history: true` are unaffected — graph writes never touch the capture path.) Batching also reduces recorded-clock consumption: one captured transaction advances the per-graph clock once, even when it contains many writes. See [Logical revision and physical time](#logical-revision-and-physical-time) for the anchor format. > **Performance.** Recorded reads reconstruct from the history relations rather > than the live tables, so they are slower than current-state reads — most > noticeably for full-graph `subgraph` / algorithm reconstructions on > PostgreSQL. Reach for `asOfRecorded` for audit and point-in-time > reconstruction, not hot-path reads. ## Including Historical Data (includeEnded) View all versions, including superseded records: ```typescript const history = await store .query() .from("Article", "a") .temporal("includeEnded") .whereNode("a", (a) => a.id.eq(articleId)) .orderBy((ctx) => ctx.a.validFrom, "desc") .select((ctx) => ({ title: ctx.a.title, validFrom: ctx.a.validFrom, validTo: ctx.a.validTo, version: ctx.a.version, })) .execute(); // Result shows all versions: // [ // { title: "Final Title", validFrom: "2024-03-01", validTo: undefined, version: 3 }, // { title: "Draft v2", validFrom: "2024-02-15", validTo: "2024-03-01", version: 2 }, // { title: "Initial Draft", validFrom: "2024-02-01", validTo: "2024-02-15", version: 1 }, // ] ``` ### Audit Trail Build a complete change history: ```typescript async function getAuditTrail(nodeId: string) { return store .query() .from("Document", "d") .temporal("includeEnded") .whereNode("d", (d) => d.id.eq(nodeId)) .select((ctx) => ({ version: ctx.d.version, title: ctx.d.title, status: ctx.d.status, validFrom: ctx.d.validFrom, validTo: ctx.d.validTo, updatedAt: ctx.d.updatedAt, })) .orderBy("d", "version", "asc") .execute(); } ``` ## Including Soft-Deleted Data (includeTombstones) Include records that have been soft-deleted: ```typescript const allIncludingDeleted = await store .query() .from("User", "u") .temporal("includeTombstones") .select((ctx) => ({ id: ctx.u.id, name: ctx.u.name, deletedAt: ctx.u.deletedAt, // Will have a value for deleted records })) .execute(); ``` ### Filtering Deleted Records ```typescript // Find only deleted records const deletedUsers = await store .query() .from("User", "u") .temporal("includeTombstones") .whereNode("u", (u) => u.deletedAt.isNotNull()) .select((ctx) => ({ id: ctx.u.id, name: ctx.u.name, deletedAt: ctx.u.deletedAt, })) .execute(); ``` ## Temporal Metadata Fields When querying with temporal context, these fields are available: | Field | Type | Description | |-------|------|-------------| | `validFrom` | `string \| undefined` | When this version became valid (`undefined` on an **open-left** row — see below) | | `validTo` | `string \| undefined` | When this version was superseded (undefined if current) | | `createdAt` | `string` | When the node was first created | | `updatedAt` | `string` | When this version was written | | `deletedAt` | `string \| undefined` | Soft-delete timestamp (undefined if not deleted) | | `version` | `number` | Optimistic concurrency version number | ### Open-left rows (`validFrom` is `undefined`) A row may have **no lower bound at all**, which means "valid since forever, as far as this store knows". `asOf` and `current` treat such a row as valid at every instant strictly before its `validTo`, or every instant if it has no end. These writes produce one: - a Store create or resurrecting upsert stating `validFrom: null`; - an interchange record stating `validFrom: null` — a source row confirmed to have no lower bound, round-tripped rather than re-stamped; - a **born-already-ended** write: one that CREATES a row, or RESETS its window, while stating a `validTo` at or before its own instant and no `validFrom`. The row's start is unknown rather than after its end, so no bound is stored and the row reads back at every `asOf` before that end. A `validTo` in the *future* is unaffected — it still stamps the write instant, so the row stays invisible at instants before it existed. Every **node** path that resets the window qualifies, and reaches the same stored shape: a create on a fresh id, a create on a tombstoned one, and a resurrecting `upsertById` / `bulkUpsertById`. An **edge** never does: an edge create cannot land on a tombstone (a taken id raises `Edge already exists`), and the two paths that resurrect one — `bulkUpsertById` and `getOrCreateByEndpoints` — RETAIN the bound the row carries and judge the stated `validTo` against it. #### Rows written by older versions Before that rule existed, a born-already-ended write stored the write instant as `valid_from`, leaving a window that runs backwards — a row readable at **no** coordinate at all. Upgrading does not rewrite such rows; they keep their window and stay invisible until an operator repairs them explicitly with `repairInvertedValidityWindows`, which normalizes them to the open-left shape above. Prefer `relations: "live-and-recorded"`: repairing only the live axis leaves the recorded twin inverted, so `asOfRecorded` reads keep returning the invisible shape. See [Repairing inverted validity windows](/schema-management#repairing-inverted-validity-windows) for the operator checklist — run it with writers stopped, and re-baseline merge branches afterwards. ```typescript .select((ctx) => ({ ...ctx.a, // All node properties validFrom: ctx.a.validFrom, validTo: ctx.a.validTo, createdAt: ctx.a.createdAt, updatedAt: ctx.a.updatedAt, deletedAt: ctx.a.deletedAt, version: ctx.a.version, })) ``` ## Temporal Traversals Temporal modes apply to traversals as well: ```typescript // See who worked at a company last year const lastYear = new Date("2023-01-01").toISOString(); const pastEmployees = await store .query() .from("Company", "c") .temporal("asOf", lastYear) .whereNode("c", (c) => c.name.eq("Acme Corp")) .traverse("worksAt", "e", { direction: "in" }) .to("Person", "p") .select((ctx) => ({ name: ctx.p.name, role: ctx.e.role, })) .execute(); ``` `store.subgraph()` and `store.algorithms.*` accept the same `temporalMode` and `asOf` options, defaulting to `graph.defaults.temporalMode`. See [Temporal Behavior](/graph-algorithms#temporal-behavior) for the algorithm surface and [`store.subgraph()` options](/schemas-stores#storesubgraphrootid-options) for subgraph. ## Real-World Examples ### Version Comparison Compare two versions of a document: ```typescript async function compareVersions(docId: string, v1: number, v2: number) { const versions = await store .query() .from("Document", "d") .temporal("includeEnded") .whereNode("d", (d) => d.id.eq(docId)) .select((ctx) => ctx.d) .execute(); const version1 = versions.find((v) => v.version === v1); const version2 = versions.find((v) => v.version === v2); return { version1, version2 }; } ``` ### Compliance Reporting Generate a report as of a specific date: ```typescript async function generateQuarterlyReport(quarterEnd: string) { const activeContracts = await store .query() .from("Contract", "c") .temporal("asOf", quarterEnd) .whereNode("c", (c) => c.status.eq("active")) .traverse("belongsTo", "e") .to("Customer", "cust") .select((ctx) => ({ contractId: ctx.c.id, value: ctx.c.value, customer: ctx.cust.name, })) .execute(); return { asOf: quarterEnd, totalContracts: activeContracts.length, totalValue: activeContracts.reduce((sum, c) => sum + c.value, 0), contracts: activeContracts, }; } ``` ### Undo/Recovery Find the previous value before an update: ```typescript async function getPreviousVersion(nodeId: string) { const versions = await store .query() .from("Document", "d") .temporal("includeEnded") .whereNode("d", (d) => d.id.eq(nodeId)) .select((ctx) => ctx.d) .orderBy("d", "version", "desc") .limit(2) .execute(); return { current: versions[0], previous: versions[1], }; } ``` ## Next Steps - [Filter](/queries/filter) - Filtering with predicates - [Traverse](/queries/traverse) - Graph traversals - [Execute](/queries/execute) - Running queries - [Bitemporal Time Travel](/examples/bitemporal-time-travel) - Valid time plus recorded time in one runnable example - [Agent Decision Replay](/examples/agent-decision-replay) - Reconstruct the exact graph an agent saw - [Breach Forensics](/examples/breach-forensics) - Traverse a reconstructed access graph at the breach instant # Troubleshooting > Solutions to common issues and frequently asked questions This guide covers common issues and their solutions when working with TypeGraph. ## Installation Issues ### "Cannot find module '@nicia-ai/typegraph'" **Cause:** Package not installed or using wrong package name. **Solution:** ```bash npm install @nicia-ai/typegraph zod drizzle-orm ``` ### "better-sqlite3 compilation failed" **Cause:** Native module compilation requires build tools. **Solutions:** **macOS:** ```bash xcode-select --install ``` **Ubuntu/Debian:** ```bash sudo apt-get install build-essential python3 ``` **Windows:** ```bash npm install --global windows-build-tools ``` **Alternative:** Use `sql.js` for pure JavaScript SQLite (no compilation needed). ### Missing optional `drizzle-orm` peer **Cause:** The managed SQLite or PGlite Store entrypoint was called without the optional `drizzle-orm` peer installed. **Solution:** Install the peer in the application that uses the managed entrypoint: ```bash npm install drizzle-orm ``` The root package and other portable entrypoints do not require Drizzle. Explicit `@nicia-ai/typegraph/adapters/drizzle/...` entrypoints load Drizzle when the module is evaluated, so a missing peer there appears as the runtime's raw module-resolution error instead of `MISSING_PEER_DEPENDENCY`. See [Managed Store Entrypoints](/backend-setup#managed-store-entrypoints). ### "Module not found: drizzle-orm/better-sqlite3" **Cause:** An explicit Drizzle adapter import is missing `drizzle-orm`, or the application imported the wrong Drizzle subpath. **Solution:** First install `drizzle-orm`, then ensure the import matches the driver: ```bash npm install drizzle-orm ``` ```typescript // Correct import { drizzle } from "drizzle-orm/better-sqlite3"; // Incorrect import { drizzle } from "drizzle-orm"; ``` ## Schema Definition Errors ### "Node schema contains reserved property names" **Cause:** Using reserved keys (`id`, `kind`, `meta`) in your Zod schema. **Solution:** Rename your properties: ```typescript // Bad - 'id' is reserved const User = defineNode("User", { schema: z.object({ id: z.string(), // Error! name: z.string(), }), }); // Good - use a different name const User = defineNode("User", { schema: z.object({ externalId: z.string(), name: z.string(), }), }); ``` TypeGraph automatically provides `id`, `kind`, and `meta` on all nodes. ### "Edge type already has constraints defined" **Cause:** Defining `from`/`to` constraints on both the edge type and graph registration. **Solution:** Define constraints in one place only: ```typescript // Option 1: On the edge type (reusable across graphs) const worksAt = defineEdge("worksAt", { from: [Person], to: [Company], }); const graph = defineGraph({ edges: { worksAt: { type: worksAt }, // No from/to here }, }); // Option 2: On the graph (flexible per-graph) const worksAt = defineEdge("worksAt"); const graph = defineGraph({ edges: { worksAt: { type: worksAt, from: [Person], to: [Company] }, }, }); ``` ## Runtime Errors ### ValidationError: "Invalid input" **Cause:** Data doesn't match the Zod schema. **Solution:** Check the error details for specific issues: ```typescript try { await store.nodes.Person.create({ name: "" }); } catch (error) { if (error instanceof ValidationError) { console.log(error.details.issues); // Zod issues array } } ``` ### NodeNotFoundError **Cause:** Attempting to read/update/delete a non-existent node. **Solution:** Check if the node exists first or handle the error: ```typescript const node = await store.nodes.Person.getById(someId); if (!node) { // Handle missing node } // Or use error handling try { await store.nodes.Person.update(someId, { name: "New" }); } catch (error) { if (error instanceof NodeNotFoundError) { console.log(`Node ${error.details.id} not found`); } } ``` ### RestrictedDeleteError **Cause:** Attempting to delete a node that has edges, with `onDelete: "restrict"` (the default). **Solution:** Either delete the edges first or use a different delete behavior: ```typescript // Option 1: Delete edges first. Include ended-but-not-deleted edges if you // are cleaning up historical validity windows too. const edges = await store.edges.worksAt.findFrom(person, { temporalMode: "includeEnded", }); for (const edge of edges) { await store.edges.worksAt.delete(edge.id); } await store.nodes.Person.delete(person.id); // Option 2: Use cascade delete in schema const graph = defineGraph({ nodes: { Person: { type: Person, onDelete: "cascade" }, }, }); ``` ### DisjointError **Cause:** Creating a node with an ID that's already used by a disjoint type. **Solution:** Ensure IDs are unique across disjoint types or don't use explicit IDs: ```typescript // If Person and Organization are disjoint: // Bad - same ID for different types await store.nodes.Person.create({ name: "Alice" }, { id: "entity-1" }); await store.nodes.Organization.create({ name: "Acme" }, { id: "entity-1" }); // Error! // Good - let TypeGraph generate unique IDs await store.nodes.Person.create({ name: "Alice" }); await store.nodes.Organization.create({ name: "Acme" }); ``` ## Query Issues ### "Alias 'x' is already in use" **Cause:** Using the same alias twice in a query. **Solution:** Use unique aliases: ```typescript // Bad store.query().from("Person", "p").traverse("knows", "e").to("Person", "p"); // Error! 'p' already used // Good store.query().from("Person", "p1").traverse("knows", "e").to("Person", "p2"); ``` ### Empty results when expecting data **Causes and solutions:** 1. **Type mismatch:** Ensure you're querying the correct node type ```typescript // Check the node type name matches exactly .from("Person", "p") // Must match defineNode("Person", ...) ``` 2. **Missing includeSubClasses:** When querying a superclass ```typescript .from("Content", "c", { includeSubClasses: true }) ``` 3. **Strict predicate:** Check your filters aren't too restrictive ```typescript // Debug by removing filters temporarily const all = await store .query() .from("Person", "p") .select((c) => c.p) .execute(); console.log(all.length); // How many total? ``` ### Slow queries **Solutions:** 1. **Use the query profiler:** ```typescript import { QueryProfiler } from "@nicia-ai/typegraph/profiler"; const profiler = new QueryProfiler(); profiler.attachToStore(store); // Run your queries... const report = profiler.getReport(); console.log(report.recommendations); ``` 2. **Add indexes** based on profiler recommendations: ```typescript import { defineNodeIndex } from "@nicia-ai/typegraph/indexes"; const nameIndex = defineNodeIndex(Person, { fields: ["name"] }); ``` 3. **Limit results:** ```typescript .limit(100) // Or use pagination .paginate({ first: 20 }) ``` ## Database Connection Issues ### "Database is locked" (SQLite) **Cause:** Multiple processes accessing the same SQLite file without WAL mode. **Solution:** Enable WAL mode: ```typescript const sqlite = new Database("myapp.db"); sqlite.pragma("journal_mode = WAL"); ``` ### Connection pool exhausted (PostgreSQL) **Cause:** Too many concurrent connections. **Solution:** Configure pool limits: ```typescript import { Pool } from "pg"; const pool = new Pool({ connectionString: process.env.DATABASE_URL, max: 20, // Adjust based on your needs idleTimeoutMillis: 30000, }); ``` ### "relation 'typegraph_nodes' does not exist" **Cause:** Migration not run. **Solution:** Run the migration SQL: ```typescript // PostgreSQL import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; await pool.query(generatePostgresMigrationSQL()); // SQLite import { generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; sqlite.exec(generateSqliteMigrationSQL()); ``` ### "permission denied" / cannot create relation on boot **Cause:** `createStoreWithSchema()` is a privileged entry point. Its warm base-schema check is one marker read with no base-adoption DDL, but bootstrap, pending base adoption, graph migrations, contribution preparation, or system index materialization can issue DDL. A DML-only role cannot safely own it. **Solution:** Run schema/DDL changes as a privileged one-time migration step, then attach at runtime with the zero-DDL `createVerifiedStore()` (or `createStore()`) under the least-privilege role. See [Database roles & least privilege](/backend-setup#database-roles--least-privilege). ### `BaseSchemaMigrationError` from a zero-DDL runtime path **Cause:** `createVerifiedStore`, `assertSchemaCurrent`, or graph-template registration/instantiation found a missing, stale, or newer deployment-wide base-schema marker. These paths deliberately do not repair physical storage. **Solution:** For a missing or stale marker, run `createStoreWithSchema(graph, adminBackend)` once under a DDL-capable role, or apply the published external base-schema migration and stamp its marker last. For a newer marker, deploy a TypeGraph release that supports that version. The error details include `installedVersion`, `requiredVersion`, and `reason`. ### `MigrationError` from `createVerifiedStore` / `assertSchemaCurrent` **Cause:** The runtime is using a code graph whose schema is ahead of the database. The least-privilege runtime cannot migrate — by design, it fails fast so requests don't run against a stale schema. **Solution:** Run `createStoreWithSchema(graph, adminBackend)` under the privileged role before promoting the new runtime build (apply any generated migration SQL first if you manage DDL externally), then restart the runtime. The thrown `MigrationError.message` includes the diff summary and migration actions to apply. ### `ConfigurationError`: "no schema has been initialized" **Cause:** A verifying attach (`createVerifiedStore` / `assertSchemaCurrent`) ran before any privileged `createStoreWithSchema()` boot — the database has no `schema_versions` row (or no typegraph tables at all). The runtime deliberately refuses to bootstrap under a least-privilege role. **Note:** running only the generated migration SQL is not sufficient — it creates the tables but does not write the schema row or contribution markers. **Solution:** Run `createStoreWithSchema(graph, adminBackend)` once under the privileged role. If you manage DDL externally with drizzle-kit / `generatePostgresMigrationSQL()` / `generateSqliteMigrationSQL()`, apply that first, then still run `createStoreWithSchema()` to commit the schema row and contribution markers. See [Database roles & least privilege](/backend-setup#database-roles--least-privilege). ### `StoreNotInitializedError` on the first operation **Cause:** The store was created with `createStore()` (a zero-I/O attach that never materializes runtime storage) against a database that no `createStoreWithSchema()` boot has initialized — commonly the runtime started before the privileged migration step ran, or the wrong role/ database is configured. This covers fulltext operations and **embedding writes**: a `store.nodes.*.create({ embedding })` (or embedding update/delete) against an un-provisioned per-`(kind, field)` table throws here rather than lazily issuing `CREATE TABLE` on the hot path. Vector *reads* — `store.search.vector`, `store.search.hybrid`, and a query-builder `.similarTo()` predicate — compile straight to SQL against the per-field table, so they surface the engine's own missing-relation error instead (`no such table: tg_vec_…` on SQLite, `relation … does not exist` on Postgres) — same cause, same solution. `createVerifiedStore()` catches every one of these cases at boot rather than at the first hot-path operation. A **`stale`** variant of this error on a vector field means something different: the storage exists but was provisioned at a different shape — typically the field's declared dimension changed after the table was created. Boot deliberately leaves such a slot untouched (with a console warning); run `store.reembedVectorField(kind, fieldPath)` to recreate the storage at the new shape and re-embed. **`ContributionUnavailableError`** with `state: "physical-storage-missing"` means the physical fulltext table disappeared after initialization. Gated fulltext operations preserve the driver error as `cause`, and transactional backends roll back failed searchable writes. Query-builder fulltext predicates compile directly to SQL and can still surface the engine's missing-relation error. Run `store.rebuildContribution("fulltext")` to recreate the table and repopulate it from the graph's nodes. A verified attach checks markers rather than the physical catalog; use `probeContributions()` when startup must detect out-of-band table loss. **Solution:** Run `createStoreWithSchema(graph, adminBackend)` once under the privileged role before the runtime attaches (it writes the contribution markers that `createStore` / `createVerifiedStore` only check), and prefer `createVerifiedStore()` over bare `createStore()` so drift fails fast. See [Database roles & least privilege](/backend-setup#database-roles--least-privilege). A plain `createStore()` performs no reads and therefore cannot check the base-schema marker at attach. If an edge write reaches legacy storage without the match-identity columns, it throws `ConfigurationError` with `details.code === "EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE"` instead of leaking the database driver's missing-column error. The remedy is the same privileged base-schema adoption. ### `SCHEMA_WRITE_FENCE_UNSUPPORTED` on the first managed write **Cause:** The Store carries committed schema metadata (for example, it came from `createStoreWithSchema`, `createVerifiedStore`, an adapter equivalent, or a cached `{ reconciled }` snapshot), but its backend cannot run transactions or does not implement the schema-write fence. Common examples are Cloudflare D1, `drizzle-orm/neon-http`, and incomplete custom backends. The attach can still succeed for reads; writes fail closed rather than racing a schema change. **Solution:** Use a transactional backend (`neon-serverless`, regular PostgreSQL, SQLite with transactions, or Durable Objects SQLite). If raw, unfenced writes are an explicit application decision, construct the Store with `createStore()` / `createAdapterStore()` without `{ reconciled }` and quiesce writers yourself around schema changes. ### A convergence or claim write refuses on a non-transactional backend This is expected when the operation needs an interactive transaction. Static adapter batches are not a public transaction, and a sequence of independent requests cannot safely implement Operational Identity, claim/cardinality checks, or undeclared dynamic `matchOn` convergence. Use a backend with `capabilities.execution.interactiveTransactions === true` and call `store.transaction(...)` for those operations. A declared edge `matchIdentity` is the exception for the eligible root `getOrCreateByEndpoints` path: its persisted canonical key is backed by a database unique arbiter, so the authoritative one-statement command can return `created` or `found` without an interactive transaction. This exception does not cover claims, history/revision sidecars, or other writes in the same application workflow. ### The store opens clean but a fulltext or vector read fails **Cause:** The durable physical contribution marker still says `initialized` while the table it names is gone — a partial restore, a hand-run `DROP`, or a schema-scoped restore that missed the strategy-owned tables. Nothing on the open path probes the catalog: boot and the runtime asserts short-circuit on a per-instance signature cache and then on durable marker rows alone, which keeps the hot path free of catalog round trips. The cost is that this database opens completely clean and fails at the first read or write that depends on the affected slot. For deployment-scoped storage such as the shared fulltext table, readiness is the conjunction of that physical marker and a graph-local activation marker. Vector tables remain graph-scoped and use one graph-local physical marker. **Diagnosis:** `store.verifyContributions()` reports detected drift or a recorded failed attempt among contributions currently expected by the active graph and backend strategies. Each entry carries `owner`, `logicalName`, `physicalName`, a `state`, and — for vector slots — `kind` and `fieldPath`. When the marker recorded an error against its last attempt, `lastError` carries it: `state` tells you which repair to run, `lastError` tells you why it broke, which is often a different question. The call is read-only — one existence query per contribution table, no DDL and no writes — so it is safe on a live store under a least-privilege role. It is deliberately **not** a boot step; call it from a health check or an operator script. **Solution:** Run `repairContributions()` on a Store backed by the privileged DDL-capable connection: ```typescript const result = await adminStore.repairContributions(); for (const entry of result.results) { if (entry.status === "failed") { console.error(entry.diagnostic, entry.error); } if (entry.status === "requires-rebuild") { console.warn("manual rebuild required", entry.diagnostic); } } ``` The method performs its own fresh audit and resolves contribution declarations from the committed graph and the active backend strategies. It never accepts a diagnostic, physical table name, or DDL from the caller. A Store opened before another writer evolved the graph catches up before enumerating vector slots instead of repairing from its stale in-memory graph snapshot. | `state` | Result | Data behavior | | --- | --- | --- | | `missing-marker` | `repaired` or `failed` | Runs idempotent DDL and re-stamps the marker; existing rows are preserved | | `failed-materialization` | `repaired` or `failed` | Retries the current idempotent contribution DDL | | `orphaned-marker` | `requires-rebuild` | The table and its data are already gone | | `stale` | `requires-rebuild` | The stored physical shape does not match the current declaration | `remaining` is a fresh post-repair diagnostic pass. An empty `remaining` array means no current declaration remains unhealthy after the pass. Once it is empty, a second call is idempotent and returns no results. For a vector `requires-rebuild` entry, use `reembedVectorField(kind, fieldPath, { embed })`. It drops and recreates the slot, so pass an `embed` callback or the field comes back with zero embeddings. For a fulltext `requires-rebuild` entry, use `rebuildContribution("fulltext")` — the third rung of the ladder, described below. Do not hand-edit the marker or run backend-owned DDL directly. `repairContributions()` intentionally does not use the public diagnostic as an instruction list and does not force marker writes. A warm backend re-reads the marker, and the normal signature guard still refuses to bless stale storage. **An empty result does not mean everything was checked.** The diagnostic enumerates only current declarations. It ignores retired marker rows and treats an expected contribution with neither marker nor table as never attempted, so `[]` is not proof of initialization. A backend that cannot probe its own catalog throws `ConfigurationError` rather than reporting a clean bill of health, but vector slots on a backend without vector support are skipped silently and correctly — that backend never materialized them, so reporting them would be a false positive on every store it opens. For a readiness check, first attach with `createVerifiedStore()` to establish schema and marker initialization, then run this diagnostic. Also assert `backend.capabilities.vector?.supported` when embedding storage is required rather than treating an empty array as proof that it is intact. ### Contribution health: probe, repair, rebuild The three contribution maintenance operations form one escalation ladder. Each rung does strictly more, and costs strictly more, than the one below it. Start at the top of this table and stop as soon as the projection is `ready`. | Rung | Call | Writes | Use when | | --- | --- | --- | --- | | 1. Probe | `store.probeContributions()` | Nothing | You want to know whether search is coherent right now. Safe on a read path, on a replica, and under a least-privilege role | | 2. Repair | `store.repairContributions()` | Marker rows and idempotent `CREATE ... IF NOT EXISTS` | The probe reports `degraded` and `verifyContributions()` says `missing-marker` or `failed-materialization` — storage is intact and only the bookkeeping is wrong | | 3. Rebuild | `store.rebuildContribution("fulltext")` | **Deletes and refills this graph's rows**; drops and recreates the shared storage only when no other graph has rows in it | `verifyContributions()` says `stale` or `orphaned-marker`, which repair reports as `requires-rebuild` | **Rung 1 — the read-only probe.** One entry per search projection the graph declares, so a caller can decide whether to issue a query without running a write operation first: ```typescript const health = await store.probeContributions(); for (const entry of health.entries) { if (entry.state !== "ready") { console.warn(`${entry.contribution} search is ${entry.state}`, entry.detail); } } ``` `entries` is empty when there is nothing to assess — a graph with no `searchable()` or `embedding()` fields, or a backend with no contribution machinery. It is never empty because a check was skipped: a backend that provisions contributions but cannot probe its catalog throws `ConfigurationError`, and declares the gap as `capabilities.contributions.probe === false`. Route on `state`; `detail` is a human-readable summary and not a stable format, so call `verifyContributions()` for the structured per-table findings behind it. `graphRevision` stamps the durable revision the assessment was taken at, placing the probe in the graph's committed history. It is graph-global like the clock it reads: an advance between two probes means something committed in between, not that a particular caller's write landed. It is absent unless the Store is revision-tracked (`revisionTracking: true` or `history: true`) and a tracked write has already anchored the clock. A store with no revision clock has no revision to stamp, and substituting a wall-clock timestamp or the schema version would be a weaker guarantee wearing the name of a stronger one — the schema version in particular does not advance on data writes, so it could not order anything. `state: "building"` is reserved. No shipped path publishes it; the destructive rebuild is atomic, so a concurrent probe observes the state before or after it and never a partial one. Treat it as "not `ready`". **Rung 3 — the destructive rebuild.** A `stale` fulltext contribution means the table exists at the shape a *previous* `createDdl` produced. The ordinary ensure path cannot fix it: its `CREATE ... IF NOT EXISTS` no-ops against the existing table, and re-stamping the marker there would leave it blessing storage whose shape is wrong — precisely what the drift guard exists to prevent. Only a drop makes the recreate meaningful, so the drop is its own named operation rather than a flag on the ensure path: ```typescript const result = await adminStore.rebuildContribution("fulltext"); // { rebuilt: ["typegraph_node_fulltext"], processed, repopulated, skipped } ``` **Opening a Store to run it.** A `stale` contribution makes `createStoreWithSchema()` refuse: its boot step materializes runtime contributions, and the drift guard will not run the current DDL against a table provisioned at another shape. That refusal is deliberate and persistent — it repeats on every restart until the shape is fixed, and it leaves the `stale` verdict intact rather than downgrading it to a state whose repair would bless the wrong shape. Reach the rebuild from a Store opened without that boot step, which `createStore()` and `createVerifiedStore()` are (they run no DDL by contract): ```typescript const adminStore = createStore(graph, backend); await adminStore.probeContributions(); // degraded, detail names `stale` await adminStore.rebuildContribution("fulltext"); // createStoreWithSchema() now opens normally again. ``` **The rebuild is scoped to the graph you call it on.** The fulltext projection is one physical table holding every graph's rows keyed by `graph_id`, so the default teardown is the same `DELETE ... WHERE graph_id` that `clear()` issues — this graph's index content and nothing else — followed by the current `createDdl`, a refill from this graph's node rows, and the marker stamp, all in one transaction under the same per-graph fence as a schema commit. That path takes **no table lock at all**: the delete is transactional and touches only rows this graph owns, so it never makes another graph's writers wait. An interrupted rebuild rolls back to the state it started from rather than leaving storage attested but empty. It escalates to dropping and recreating that shared table only when the table holds no other graph's rows — the case where the drop takes nothing with it. That escalation is the one repair for storage provisioned at a shape the current DDL no longer produces, and because the DDL it issues is database-global it runs under a database-scoped advisory lock (`typegraph:contribution-ddl`; a no-op on SQLite, whose fence already holds the single writer slot) rather than only the per-graph fence. **When the shared table is in use by another graph, a `stale` rebuild refuses.** Only recreating the storage repairs a `stale` shape, so if that storage still holds rows belonging to other graphs the call throws `ContributionRebuildUnsupportedError` with `reason: "shared-storage-in-use"` rather than destroying content it cannot reconstruct — those rows are derived from other graphs' nodes through their own schemas — or re-stamping this graph's marker over a physical shape nothing verified. `details.otherGraphIds` names the graphs that are in the way. The sanctioned repair is a maintenance window with every graph on that database offline: drop the table out of band, then run `store.rebuildContribution("fulltext")` once per graph, each run recreating the table from the current DDL and refilling that graph's own rows. What the *recreate* path costs, and why it is still the right trade: the transaction is held for the whole refill, and on PostgreSQL the rebuild takes `LOCK TABLE ... IN ACCESS EXCLUSIVE MODE` on the shared table and keeps it until commit. It takes that lock **before** deciding to drop, not merely as a side effect of the `DROP TABLE` — the verdict "no other graph has rows here, so dropping this destroys nothing" is only as good as the exclusion it was computed under. Ordinary fulltext writes take no advisory lock, so the contribution DDL lock excludes other *rebuilds* and nothing else: a neighbouring graph's `INSERT` could commit between an unlocked probe and the drop, and be destroyed by a rebuild that had already decided it was alone. The sequence is therefore probe → `ACCESS EXCLUSIVE` → re-probe, and only the re-probe's verdict authorizes a drop. A verdict that flips under the lock loses the drop and keeps the lock (PostgreSQL holds locks until commit), which is the rare and safe direction to be wrong in. The cheap unlocked probe ahead of it exists only to keep the graph-scoped path off the relation lock, and can only err toward keeping the table. The window blocks more than searches — every write to a kind with `searchable()` fields maintains the same table, so those block too, for every graph on the database. On SQLite the rebuild holds the write lock for the same span, so concurrent writers wait out their busy timeout and then fail. Run it in a maintenance window on a large graph. When the storage *shape* is fine and only the content is stale — a field gained `searchable()` after data was written, or a `language` changed — `store.search.rebuildFulltext()` is the incremental, resumable pass that transacts per page instead. Nothing is permanently lost for the graph you rebuild: its searchable text is derived from node properties TypeGraph already stores. Other graphs on the same database are not in reach either — their rows are kept by the graph-scoped delete, and the drop that would take them never runs (a `stale` shape that could only be repaired by that drop refuses instead). Nodes whose stored `props` cannot be read as an object are counted in `skipped` and are absent from the rebuilt index; `store.search.rebuildFulltext()` reports their ids individually. **Vector contributions cannot be rebuilt, and the call refuses rather than trying.** `rebuildContribution("vector")` always throws `ContributionRebuildUnsupportedError` with `reason: "vector-source-unavailable"`. TypeGraph stores the vectors callers supply and never the inputs that produced them, so the embeddings exist only in the storage a rebuild would drop — dropping anyway would destroy them and hand back storage that looks healthy and returns nothing. `reembedVectorField(kind, fieldPath, { embed })` is the sanctioned destructive path for vector storage precisely because it takes the callback that can regenerate what the drop discards. The same typed error covers two wiring gaps, and both refuse before anything is dropped: `reason: "no-drop-ddl"` when the active fulltext strategy declares no `dropDdl` on its contribution, and `reason: "no-schema-fence"` when the backend exposes no `schemaWriteTransaction` to make the sequence atomic (the HTTP-only PostgreSQL drivers, and SQLite with transactions disabled). Both are declared ahead of time as `capabilities.contributions.rebuild === false`. ## Semantic Search Issues ### "Extension not found" / "vector type not available" **Cause:** Vector extension not installed. Only applies to PostgreSQL (pgvector) and SQLite (sqlite-vec). libSQL / Turso has a built-in native vector engine — there is nothing to load and it is wired automatically by `createLibsqlBackend`. **PostgreSQL:** ```sql CREATE EXTENSION IF NOT EXISTS vector; ``` **SQLite:** ```typescript import * as sqliteVec from "sqlite-vec"; sqliteVec.load(sqlite); // Must be called before creating backend ``` ### "Dimension mismatch" **Cause:** Query embedding has different dimension than stored embeddings. **Solution:** Use consistent embedding dimensions: ```typescript // Schema defines 1536 dimensions const Document = defineNode("Document", { schema: z.object({ embedding: embedding(1536), }), }); // Query embedding must also be 1536 const queryEmbedding = await generateEmbedding(text); console.log(queryEmbedding.length); // Should be 1536 ``` ### "Inner product not supported" (SQLite / libSQL) **Cause:** `inner_product` is PostgreSQL-only. Neither sqlite-vec nor libSQL support the inner product metric (cosine and l2 only). Check `backend.capabilities.vector.metrics` for the active backend. **Solution:** Use cosine or L2: ```typescript // Instead of: d.embedding.similarTo(query, 10, { metric: "inner_product" }); // Use: d.embedding.similarTo(query, 10, { metric: "cosine" }); ``` ## TypeScript Issues ### "Property 'x' does not exist on type" **Cause:** Accessing a property not defined in your schema. **Solution:** Ensure the property is in your Zod schema: ```typescript const Person = defineNode("Person", { schema: z.object({ name: z.string(), email: z.string().optional(), }), }); // Now both properties are available with correct types const person = await store.nodes.Person.getById(id); person?.name; // string person?.email; // string | undefined ``` ### Type inference not working in select **Cause:** Complex generic inference limitations. **Solution:** Use explicit typing or simplify: ```typescript // If inference fails, be explicit .select((ctx) => ({ name: ctx.p.name as string, company: ctx.c.name as string, })) ``` ## Still Having Issues? 1. **Check the [Limitations](/limitations)** page for known constraints 2. **Review [Architecture](/architecture)** to understand how TypeGraph works 3. **Search [GitHub Issues](https://github.com/nicia-ai/typegraph/issues)** for similar problems 4. **Open a new issue** with a minimal reproduction case # Errors > Error types and handling in TypeGraph TypeGraph uses typed errors to communicate specific failure conditions. All errors extend the base `TypeGraphError` class and include categorization, contextual details, and actionable suggestions. ## Error Categories Every error is categorized to help determine the appropriate response: | Category | Description | Typical Response | |----------|-------------|------------------| | `user` | Invalid input or misuse of API | Fix the input and retry | | `constraint` | Graph constraint violated | Handle as business logic violation | | `system` | Internal or infrastructure error | Log, alert, potentially retry | ```typescript import { isUserRecoverable, isConstraintError, isSystemError } from "@nicia-ai/typegraph"; try { await store.nodes.Person.create(data); } catch (error) { if (isUserRecoverable(error)) { // Show validation errors to user return { error: error.toUserMessage() }; } if (isConstraintError(error)) { // Handle business rule violation return { error: "This operation violates a constraint" }; } if (isSystemError(error)) { // Log and alert console.error(error.toLogString()); throw error; } } ``` ## Base Error ### `TypeGraphError` Base error class for all TypeGraph errors. ```typescript class TypeGraphError extends Error { readonly code: string; readonly category: ErrorCategory; readonly details: Readonly>; readonly suggestion?: string; // Format error for end users (includes suggestion if available) toUserMessage(): string; // Format error for logging (includes code, category, and details) toLogString(): string; } type ErrorCategory = "user" | "constraint" | "system"; ``` **Properties:** | Property | Type | Description | |----------|------|-------------| | `code` | `string` | Machine-readable error code | | `category` | `ErrorCategory` | Error classification for handling | | `details` | `Record` | Additional context about the error | | `suggestion` | `string \| undefined` | Actionable guidance for resolution | **Methods:** | Method | Returns | Description | |--------|---------|-------------| | `toUserMessage()` | `string` | Human-readable message with suggestion | | `toLogString()` | `string` | Detailed string for logging/debugging | ## Validation Errors ### `ValidationError` Thrown when schema validation fails during node or edge creation/update. Includes structured issue details with context about which entity failed. ```typescript interface ValidationErrorDetails { readonly issues: readonly ValidationIssue[]; readonly entityType?: "node" | "edge"; readonly kind?: string; readonly operation?: "create" | "update"; readonly id?: string; } interface ValidationIssue { readonly path: string; readonly message: string; readonly code?: string; } ``` **Example:** ```typescript try { await store.nodes.Person.create({ name: "" }); // Empty name fails min(1) } catch (error) { if (error instanceof ValidationError) { console.log(error.category); // "user" console.log(error.details.kind); // "Person" console.log(error.details.operation); // "create" console.log(error.details.issues); // [{ path: "name", message: "String must contain at least 1 character(s)" }] console.log(error.toUserMessage()); // "Validation failed for Person create: name - String must contain at least 1 character(s) // // Suggestion: Check the data you're providing matches the schema..." } } ``` #### `INVERTED_VALIDITY_WINDOW` A `ValidationError` whose issue carries the exported code `INVERTED_VALIDITY_WINDOW` refused a valid-time window of negative width: the write's `validTo` precedes the row's effective `validFrom`, so the row would have stopped being true before it started and no `asOf` coordinate could observe it. Branch on the code rather than on the message. ```typescript import { INVERTED_VALIDITY_WINDOW_CODE, ValidationError } from "@nicia-ai/typegraph"; try { // The stored validFrom is later than this end. await store.edges.worksAt.update(edgeId, {}, { validTo: "2020-01-01T00:00:00.000Z" }); } catch (error) { if ( error instanceof ValidationError && error.details.issues.some((issue) => issue.code === INVERTED_VALIDITY_WINDOW_CODE) ) { // Supply an explicit validFrom for a historical window, or drop validTo. } } ``` Interchange import records the same refusal as a per-row error prefixed with the code, so one bad row does not abort the import; trusted import refuses the whole stream with `TrustedImportError` reason `invalid_stream`. A zero-width window (`validTo === validFrom`) is legal and never raises this, and neither is a write that STAMPS its own start while carrying only a historical `validTo`: any create, and a node resurrection through `upsertById` / `bulkUpsertById`. Both store no lower bound instead. An edge resurrection RETAINS the bound the row already holds, so a `validTo` before that bound still raises this. #### `IMMUTABLE_VALIDITY_LOWER_BOUND` A `ValidationError` whose issue carries the exported code `IMMUTABLE_VALIDITY_LOWER_BOUND` refused a `validFrom` the write could not apply. A live row's lower bound is history: an in-place update never rewrites `valid_from`, so a bound naming a different instant is refused rather than accepted and silently dropped. The message names both instants — the one stated and the one the row stores — so you can restate the stored bound without a second read. ```typescript import { IMMUTABLE_VALIDITY_LOWER_BOUND_CODE, ValidationError } from "@nicia-ai/typegraph"; try { // The row is live and started at some other instant. await store.nodes.Person.upsertById(id, props, { validFrom: "2020-01-01T00:00:00.000Z" }); } catch (error) { if ( error instanceof ValidationError && error.details.issues.some( (issue) => issue.code === IMMUTABLE_VALIDITY_LOWER_BOUND_CODE, ) ) { // Omit validFrom, or restate the bound the row already holds. } } ``` What deliberately does not raise it: - **Restating the stored bound.** Naming the instant the row already holds is accepted; there is nothing to apply and nothing being ignored. - **A create, or a resurrection.** Both write a fresh window, so a stated `validFrom` is stored — that is the way to give a row a different lower bound. - **`getOrCreateByEndpoints` returning an existing edge.** That branch performs no write, so `validFrom` / `validTo` describe the row to create if none is found. `clearValidTo` is refused on a live return-mode match because it names a mutation; use `ifExists: "update"`. - **A node upsert or endpoint-matched edge update with `onImmutableLowerBound: "preserve"`.** This explicitly treats `validFrom` as create/resurrection-only input. A live-row update keeps its stored lower bound while still applying props and `validTo`; the default remains `"refuse"` so an unqualified bound is never silently dropped. Edge updates use the policy with `ifExists: "update"`; the bulk edge form sets it per item. Under the default `"refuse"` policy, it reaches every path that accepts `validFrom` against a live row: `upsertById`, `bulkUpsertById` (including a repeated id in one batch, judged against the row the batch just queued), `getOrCreateByEndpoints` / `bulkGetOrCreateByEndpoints` with `ifExists: "update"`, and interchange import's `onConflict: "update"` legs — where, as with the inverted-window refusal, it is recorded as a per-row error prefixed with the code rather than aborting the import. #### `ENTITY_ALREADY_EXISTS` A `ValidationError` whose issue carries the exported code `ENTITY_ALREADY_EXISTS` refused a create because the id is already taken. `details.entityType` says whether a node or an edge was refused and `details.kind` names its kind. ```typescript import { ENTITY_ALREADY_EXISTS_CODE, ValidationError } from "@nicia-ai/typegraph"; try { await store.nodes.Person.create({ name: "Alice" }, { id: takenId }); } catch (error) { if ( error instanceof ValidationError && error.details.issues.some((issue) => issue.code === ENTITY_ALREADY_EXISTS_CODE) ) { // Use a different id, or update the existing entity. } } ``` The code is the same whichever layer noticed, on either backend. A node create finds out from its own existence probe — but the probe and the INSERT are two statements, and PostgreSQL does not serialize two write transactions under its default READ COMMITTED isolation, so a concurrent create of the same NEW id can commit in between and the engine refuses the INSERT instead. (SQLite's `BEGIN IMMEDIATE` gives the writer slot to one transaction at a time, so its probe always sees the winner's row.) An edge create has no existence probe at all, so the engine's refusal is always what reports a taken edge id. All of these raise the same error, so a caller retrying a generated id needs one branch, not several. `details.id` names the taken id, and is present for every single-entity create. It is absent only when the refused statement inserted more than one row: the engine reports that the statement collided without saying which row did, and its transaction is already aborted, so there is nothing left to probe. No race is needed to reach that — a bulk create of edges, whose ids you supplied and which nothing probes, is refused this way on every backend. Treat `details.id` as optional if you create in bulk. This is about identity, not values. A conflict on a declared `unique` constraint raises `UniquenessError` instead, and a violated `unique: true` index declaration surfaces as the engine's own failure — neither is reshaped into this error. ### `DisjointError` Thrown when attempting to create a node that violates a disjointness constraint. ```typescript // If Person and Organization are disjoint: await store.nodes.Person.create({ name: "Alice" }, { id: "entity-1" }); try { // Same ID, different disjoint type await store.nodes.Organization.create({ name: "Acme" }, { id: "entity-1" }); } catch (error) { if (error instanceof DisjointError) { console.log(error.category); // "constraint" console.log(error.details); // { nodeId: "entity-1", attemptedKind: "Organization", conflictingKind: "Person" } console.log(error.suggestion); // "Use a different ID for the new node, or delete the existing node first..." } } ``` ### `IdentityContradictionError` Thrown when an identity mutation would make the assertion ledger contradictory — for example asserting two nodes are the same after they were asserted different, folding a same-class pair the ontology forbids, or importing an archive whose assertions conflict with the target graph. Only raised on identity-enabled graphs. ```typescript try { await tx.identity.assertSame(alice, aliceCopy); } catch (error) { if (error instanceof IdentityContradictionError) { console.log(error.code); // "IDENTITY_CONTRADICTION" console.log(error.category); // "constraint" console.log(error.details); // { // operation: "assertSame", // "assertSame" | "assertDifferent" | "fold" | "import" // a: { kind: "Person", id: "..." }, // b: { kind: "Person", id: "..." }, // reason: "different-assertion", // "different-assertion" | "same-class" | "disjoint-kinds" // conflictingAssertionId: "...", // present when an existing assertion conflicts // conflictingKinds: ["Person", "Organization"], // present when reason is "disjoint-kinds" // } console.log(error.suggestion); // "Retract the conflicting identity assertion or correct the graph ontology before retrying." } } ``` ### Identity validity errors `IdentityValidityWindowError` refuses a future start, future end, inverted window, or a second non-identical open window for one current semantic pair. Its code identifies the reason: `IDENTITY_VALIDITY_FUTURE_START`, `IDENTITY_VALIDITY_FUTURE_END`, `IDENTITY_VALIDITY_INVERTED`, or `IDENTITY_VALIDITY_OPEN_WINDOW_CONFLICT`. `IdentityEndpointValidityError` (`IDENTITY_ENDPOINT_VALIDITY`) means an explicit assertion window extends outside an endpoint node's own validity or deletion bounds. Future or inverted identity windows are user-category input errors. A second non-identical open window and an endpoint-window conflict are constraint-category errors. Both classes are package-root exports. ### `IdentityMergeConflictError` Detected at merge **plan time** when the branches being merged carry opposing or otherwise contradictory identity truth: one branch asserts a pair `same` while another asserts it `different` (directly, or transitively through a chain of `same` assertions no single branch ever wrote), a branch retracts an assertion that a different branch reasserts under a new id (a retract/reassert race — a branch that reasserts a pair it *also* retracted itself is convergent, not a conflict, and merges cleanly), or a branch asserts an identity relation over a node another branch deleted. Extends `MergeError`, so an `instanceof MergeError` catch covers it alongside the other merge failures. `merge()` and `IdentityMergeConflictError` are both exported from `@nicia-ai/typegraph/graph-merge`, not the package root. `merge()` takes an array of branches and never throws a `MergeError` — it **returns** a `Result`: ```typescript import { merge, IdentityMergeConflictError, isErr } from "@nicia-ai/typegraph/graph-merge"; const result = await merge(store, [branch]); if (isErr(result)) { if (result.error instanceof IdentityMergeConflictError) { console.log(result.error.code); // "GRAPH_MERGE_IDENTITY_CONFLICT" console.log(result.error.details); } throw result.error; } ``` ### `MergeConstraintConflictError` Returned when `merge()`, `mergeIncremental()`, or `applyMergePlan()` resolves a plan whose final graph violates a deterministic store constraint. The store remains the owner of constraint enforcement: the merge translates its typed refusal only at the commit boundary, after the transaction has rolled back. ```typescript import { isErr, merge, MergeConstraintConflictError, } from "@nicia-ai/typegraph/graph-merge"; const result = await merge(store, branches); if (isErr(result) && result.error instanceof MergeConstraintConflictError) { console.log(result.error.code); // "GRAPH_MERGE_CONSTRAINT_CONFLICT" console.log(result.error.category); // "constraint" console.log(result.error.details.constraintCode); // e.g. "CARDINALITY_ERROR" console.log(result.error.details.edgeKind); // copied from the store error console.log(result.error.cause); // the original CardinalityError, etc. } ``` Cardinality, uniqueness, endpoint, disjointness, and restricted-delete refusals share this surface when they arise from node or edge application. The planner normally co-buckets nodes with the same declared unique key, but a late store-owned uniqueness refusal uses the same completeness boundary rather than falling back to a system error. Identity truth conflicts retain `IdentityMergeConflictError`; backend, environment, and stale-plan failures retain their existing system errors. Constraint failure is atomic: neither graph writes nor merge provenance records survive. ### Merge plan and evidence errors The reviewable merge lifecycle also returns errors in its `Result` arm. It does not throw them: ```typescript import { applyMergePlan, isErr, planMerge, StaleMergePlanError, } from "@nicia-ai/typegraph/graph-merge"; const planned = await planMerge(store, branches, options); if (isErr(planned)) throw planned.error; const applied = await applyMergePlan(store, planned.data); if (isErr(applied)) { if (applied.error instanceof StaleMergePlanError) { // The reviewed artifact no longer describes the target. Plan and review again. } throw applied.error; } ``` | Error | Code | Meaning | | --- | --- | --- | | `MergePlanCapabilityError` | `GRAPH_MERGE_PLAN_CAPABILITY` | The target cannot supply a durable revision fence for a cross-time plan. Enable `revisionTracking` or `history`; the contiguous `merge()` wrappers retain their documented compatibility behavior. | | `MergePlanningStaleError` | `GRAPH_MERGE_PLANNING_STALE` | The target revision changed between the planner's opening and closing observations. This is an expected retry-and-replan outcome under concurrency: no artifact is returned, so recapture the target and create a new plan before retrying. | | `StaleMergePlanError` | `GRAPH_MERGE_PLAN_STALE` | The target moved after planning, the plan already succeeded, or another concurrent application won. No plan writes committed. | | `InvalidMergePlanError` | `GRAPH_MERGE_PLAN_INVALID` | The value failed the versioned plan schema or a semantic invariant. | | `UnsupportedMergePlanVersionError` | `GRAPH_MERGE_PLAN_VERSION_UNSUPPORTED` | `formatVersion` is not supported by this TypeGraph version. | | `MergePlanDigestMismatchError` | `GRAPH_MERGE_PLAN_DIGEST_MISMATCH` | Canonical plan content differs from the recorded digest. | | `MergePlanTargetMismatchError` | `GRAPH_MERGE_PLAN_TARGET_MISMATCH` | The plan names a different graph id from the supplied target. | | `MergePlanSchemaMismatchError` | `GRAPH_MERGE_PLAN_SCHEMA_MISMATCH` | The plan was resolved under a different active schema version or hash. | | `MergePlanOriginMismatchError` | `GRAPH_MERGE_PLAN_ORIGIN_MISMATCH` | The target has an independently-created revision clock, even if its numeric revision happens to match. | | `CandidateSourceError` | `GRAPH_MERGE_CANDIDATE_SOURCE` | A built-in candidate source failed. `details` identifies its source id, entity kind, and operation context. | | `MatchEvidenceError` | `GRAPH_MERGE_EVIDENCE` | Candidate evidence is malformed or a score is non-finite. `NaN` and infinity are refused, never serialized or silently dropped. | Plan validation and the target/schema/origin/revision fence run before canonical writes. The revision check is inside the same transaction as apply, so two concurrent attempts cannot both commit. A stale plan is not repaired or adapted: create a new plan and obtain approval for its new `digest`. Plans may contain the complete proposed application data. Their digest detects content changes and gives approval systems a stable identity, but it is not a signature and does not authenticate storage, authorize a caller, or prove who created the artifact. Protect plan data and enforce those trust decisions in the application before calling `applyMergePlan()`. ### `EndpointError` Thrown when an edge is created with invalid endpoint types. ```typescript // If worksAt only allows Person -> Company: try { await store.edges.worksAt.create(company, person, {}); // Wrong direction } catch (error) { if (error instanceof EndpointError) { console.log(error.category); // "constraint" console.log(error.suggestion); // "Check the edge definition to see which node types are allowed..." } } ``` ### `EndpointPairError` Thrown when a [source-dependent edge](/core-concepts#source-dependent-targets) receives a source/target combination that matches no declared pair. It extends `TypeGraphError` directly, so catching `EndpointError` alone does not catch it. An invalid source kind continues to produce `EndpointError`. ```typescript import { EndpointPairError } from "@nicia-ai/typegraph"; try { // Dynamic callers are checked at runtime, too. // assignedTo allows Employee -> Department and Student -> Course. await store.getEdgeCollection("assignedTo").create(employee, course, {}); } catch (error) { if (error instanceof EndpointPairError) { console.log(error.code); // "ENDPOINT_PAIR_ERROR" console.log(error.category); // "constraint" console.log(error.details); // { // edgeKind: "assignedTo", endpoint: "pair", // fromKind: "Employee", toKind: "Course", // allowedPairs: [ // { from: "Employee", to: "Department" }, // { from: "Student", to: "Course" }, // ], // } } } ``` Malformed target maps and graph registrations that widen built-in constraints fail at configuration time with `ConfigurationError`. ### `CardinalityError` Thrown when a cardinality constraint is violated. ```typescript // If worksAt has cardinality: "one" (person can only work at one company): await store.edges.worksAt.create(alice, acme, { role: "Engineer" }); try { await store.edges.worksAt.create(alice, otherCompany, { role: "Consultant" }); } catch (error) { if (error instanceof CardinalityError) { console.log(error.category); // "constraint" console.log(error.details); // { edgeKind: "worksAt", fromKind: "Person", fromId: "", cardinality: "one", existingCount: 1 } console.log(error.suggestion); // "Remove the existing edge before creating a new one, or update the existing edge..." } } ``` ### `UniquenessError` Thrown when a uniqueness constraint is violated. ```typescript // If email has a unique constraint: await store.nodes.Person.create({ name: "Alice", email: "alice@example.com" }); try { await store.nodes.Person.create({ name: "Bob", email: "alice@example.com" }); } catch (error) { if (error instanceof UniquenessError) { console.log(error.category); // "constraint" console.log(error.details); // { constraintName: "unique_email", kind: "Person", existingId: "", newId: "", fields: ["email"] } console.log(error.suggestion); // "Use a different value for the unique field, or update the existing record..." } } ``` ### `EdgeMatchIdentityConflictError` Thrown when a direct edge create collides with the edge kind's declared `matchIdentity`. Use `getOrCreateByEndpoints()` when the intended behavior is to return the existing identity owner. ## Not Found Errors ### `NodeNotFoundError` Thrown when a referenced node does not exist. ```typescript try { await store.nodes.Person.update("nonexistent-id", { name: "New Name" }); } catch (error) { if (error instanceof NodeNotFoundError) { console.log(error.category); // "user" console.log(error.details); // { kind: "Person", id: "nonexistent-id" } console.log(error.suggestion); // "Verify the node ID is correct and the node hasn't been deleted..." } } ``` ### `EdgeNotFoundError` Thrown when a referenced edge does not exist. ```typescript try { await store.edges.worksAt.update("nonexistent-edge", { role: "Manager" }); } catch (error) { if (error instanceof EdgeNotFoundError) { console.log(error.category); // "user" console.log(error.details); // { kind: "worksAt", id: "nonexistent-edge" } console.log(error.suggestion); // "Verify the edge ID is correct and the edge hasn't been deleted..." } } ``` ### `KindNotFoundError` Thrown when referencing a node or edge type that doesn't exist in the graph definition. ```typescript try { await store.query().from("NonExistentType", "n").execute(); } catch (error) { if (error instanceof KindNotFoundError) { console.log(error.category); // "user" console.log(error.details); // { kindName: "NonExistentType", entity: "node" } console.log(error.suggestion); // "Check the graph definition to see which node and edge types are available..." } } ``` ### `EndpointNotFoundError` Thrown when an edge references a node that doesn't exist. ```typescript try { await store.edges.worksAt.create( { kind: "Person", id: "nonexistent" }, company, { role: "Engineer" } ); } catch (error) { if (error instanceof EndpointNotFoundError) { console.log(error.category); // "user" console.log(error.details); // { edgeKind: "worksAt", endpoint: "from", nodeKind: "Person", nodeId: "nonexistent" } console.log(error.suggestion); // "Create the referenced node first, or verify the node ID is correct..." } } ``` ## Delete Errors ### `RestrictedDeleteError` Thrown when delete is blocked due to existing edges (when `onDelete: "restrict"`). ```typescript // If Person has edges and onDelete is "restrict": try { await store.nodes.Person.delete(alice.id); } catch (error) { if (error instanceof RestrictedDeleteError) { console.log(error.category); // "constraint" console.log(error.details); // { nodeKind: "Person", nodeId: "", edgeCount: 3, edgeKinds: ["worksAt", "authored"] } console.log(error.suggestion); // "Delete all edges connected to this node first, or change the delete behavior..." } } ``` ## Configuration Errors ### `ConfigurationError` Thrown when the store, backend, or schema definition is misconfigured. ```typescript // Using transactions on D1 (which doesn't support them): try { await store.transaction(async (tx) => { // ... }); } catch (error) { if (error instanceof ConfigurationError) { console.log(error.category); // "system" console.log(error.suggestion); // "Check the backend documentation for supported features..." } } ``` #### Definition-time unique-constraint refusals `defineGraph()` validates every node kind's `unique` constraints when the graph is defined, rather than leaving a broken `where` clause to surface as odd behavior on the first write. Three states are refused with `ConfigurationError`: - A `where` callback that **does not return a predicate** — `details` carries `kind` and `constraintName`. - A predicate naming a **field the kind's schema does not declare** — `details` adds `field` and `declaredFields`. - A `where` clause on a kind whose **schema is not an object schema** (it exposes no `.shape`, so there is no declared-field set to check the clause against) — `details` carries `kind` and `constraintName`. Refused rather than left unvalidated, because skipping the check silently would disable this guard for exactly the untyped callers it exists for. A plain `unique: [{ fields }]` on such a schema is *not* refused: it names props by key and evaluates fine against a non-object schema. All three carry only the class-level code `CONFIGURATION_ERROR`; match them by class, not by a `details.code`. The equivalent invariant on the graph-extension document path does have a stable code, `UNKNOWN_UNIQUE_WHERE_FIELD`. A constraint built **outside** `defineGraph` never passed this gate, so the non-predicate case is refused at evaluation too: `checkWherePredicate` throws the same `ConfigurationError` (with `constraintName` and `fields`) on the write path instead of treating a broken clause as one that applies to every row. All three readers of a `where` clause — definition-time validation, per-write evaluation, and persistence-time capture — now agree, because they read it through one shared function. Because the check evaluates the clause, a `where` callback now runs once at definition time in addition to its per-write evaluations — keep it pure. The check applies to node kinds whose schema exposes an object shape; edge `unique` constraints are not validated here. Statically typed callers were already unable to name an undeclared field, so this bites untyped or generated definitions. #### Definition-time `__proto__` property refusal `defineNode()` / `defineEdge()` refuse a schema that declares a property named `__proto__` with a `ConfigurationError` carrying `details.conflicts` and a `nodeType` / `edgeType` key. The name is **unstorable**, not merely reserved: Zod accepts it in a shape but drops it from every parse result — reporting success even when the field is required — so a value written to it is silently lost. It is only reachable through a computed key. `z.object({ __proto__: … })` written literally sets the shape object's own prototype instead of creating an entry, while `z.object({ ["__proto__"]: z.string() })` yields a shape whose `Object.keys` really does contain it. The graph-extension document path refuses the identical declaration with the stable issue code `RESERVED_PROPERTY_NAME`, at any nesting depth — so a nested object field named `__proto__` is refused on the same grounds as a top-level one. Before this, the two authoring paths disagreed about the same field: a typed refusal on the document path, silent data loss on the typed one. #### `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` A write guarded by a declared constraint runs its probe and its write under one per-graph mutual exclusion. That fence is transaction-scoped on both dialects (SQLite's `BEGIN IMMEDIATE`, PostgreSQL's `pg_advisory_xact_lock`), so a backend reporting `capabilities.execution.interactiveTransactions: false` — Cloudflare D1, `drizzle-orm/neon-http`, any SQLite backend built with `transactionMode: "none"` — cannot hold it, and the write is refused rather than run unfenced. Durable Objects are unaffected. `details.constraint` names which class needed the fence, because "this backend cannot fence constrained writes" is unusable advice while "your `cardinality: 'one'` edge cannot be enforced here" is actionable. The `suggestion` carries the per-class way forward. | `details.constraint` | The write it describes | | --- | --- | | `edgeCardinality` | Creating or resurrecting an edge whose `cardinality` is `one`, `unique`, or `oneActive`. | | `edgeMatchKeyConvergence` | Endpoint convergence that requires the portable transaction-scoped path: an undeclared dynamic `matchOn`, constrained cardinality, update or temporal options, derived/custom backends, or schema-aware resurrection of a tombstoned winner. A schema-declared durable `matchIdentity` removes this fence from eligible live single-item and bulk create/found paths. | | `nodeDisjointness` | Creating a node under a kind that participates in a `disjointWith` axiom. Probed only where a node comes into existence, so deletes and in-place updates are not refused. | | `nodeUniquenessScope` | Creating **or updating** a node under a `scope: "kindWithSubClasses"` unique that actually expands past the node's own kind. A `scope: "kind"` unique is backed by the uniques primary key and needs no fence. | `details.graphId` names the graph. Unconstrained writes on the same backend are untouched — see [Declared constraints require an interactive transaction](/backend-setup#declared-constraints-require-an-interactive-transaction) for what still works there. `CONSTRAINT_TRANSACTION_NOT_WRITE_FENCED` is the corresponding refusal for a caller-adopted SQLite transaction whose `DEFERRED` snapshot became stale before the constrained write could take the writer slot. Roll back that transaction and retry it with `BEGIN IMMEDIATE`; TypeGraph-owned transactions already use that mode. The refusal happens before the constraint probe, so the write is fenced or refused rather than allowed to rely on a stale decision. #### `BATCH_WRITE_UNSUPPORTED` A backend whose `capabilities.execution.unitOfWork` is `"batch"` (Cloudflare D1's `batch()`, Neon HTTP's `transaction(queries)`) fixes every statement before the first one runs and commits them together with no session in between. Every fused write on such a backend — a static batch and a certified atomic program alike — asserts the active schema version inside the very statement that writes, so a stale version writes nothing and the store reports `StaleVersionError`. See [The guard every fused write shares](/limitations#the-guard-every-fused-write-shares). A write that needs more than that one guarded statement refuses, but `BATCH_WRITE_UNSUPPORTED` is not itself a top-level error code: the enforcing gate keeps its own class and code (`CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, `UNSUPPORTED_BACKEND_CAPABILITY`, `IDENTITY_REQUIRES_ATOMIC_BACKEND`, or a plain `ConfigurationError` for `history` / `revisionTracking` / a schema commit) and nests `{ code: "BATCH_WRITE_UNSUPPORTED", reason }` under `details.batchRefusal`, naming what a closed batch cannot supply: | `details.batchRefusal.reason` | What it needs | Raised by | | --- | --- | --- | | `interactive-callback` | Hold an interactive callback transaction open across several round trips. | `store.transaction(fn)` / `store.transactionWithReceipt(fn)` | | `constraint-needs-probe` | Read a value it wrote earlier in the same write before deciding what to write next. | A declared constraint's probe-then-write (`CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, above) | | `identity` | Read and write Operational Identity's closure across several round trips inside one held transaction. | `Store` construction, or `requireAtomicIdentityBackend`, when `graph.identity` is declared | | `history` | Hold the per-graph write lock and clock open across a whole write cascade. | `history: true` or `revisionTracking: true` | | `schema-commit` | Hold one transaction across its compare-and-swap read and its activating write. | `commitSchemaVersion` / `setActiveVersion` | `SCHEMA_WRITE_FENCE_UNSUPPORTED` — the portable schema-version fence an ineligible write falls back to (see [Schema Migrations](/schema-management)) — does not carry `batchRefusal`. It is reached from many fuse failures that are not specific to a batch-tier backend (an ineligible write kind, a tombstone-resurrection write a supplied id falls through to, a derived backend, a provenance mismatch), so it states its plain limitation without guessing which of the reasons above, if any, applies. #### Write-fence declaration codes `capabilities.writeFence` resolves one of four write-fence plans a lock site consumes — see [Write fence declaration](/backend-setup#write-fence-declaration-writefence). `ConfigurationError` codes name the ways a backend's fence declaration, or its resolved plan, turns out not to cover what a write needs: | `details.code` | Raised when | | --- | --- | | `WRITE_FENCE_DECLARATION_INVALID` | The declared `writeFence` fails runtime validation: an unrecognized `mechanism` string, an unrecognized `drain` string under `mechanism: "advisory"`, or a `drain` key present on `mechanism: "engine-serialized"` / `"caller-serialized"` (`drain` applies only to `"advisory"`). `details.field` names `"mechanism"` or `"drain"`; for an unrecognized value, `details.accepted` lists the allowed strings. Raised by `resolveWriteFencePlan` before any plan is shaped — an invalid `drain` never falls through to behaving like `"quiescent"`. | | `WRITE_FENCE_SQL_UNAVAILABLE` | The resolved declaration's `mechanism` is `"advisory"` but the backend's `fenceSql` is missing the member that `mechanism`/`drain` combination needs to spell (`advisoryLockExpression`, `isolationFactExpression`, or, under `drain: "table-lock"`, `lockTables`) — or, independently of any lock plan, a session isolation-level read (recorded capture's isolation guard) finds no `fenceSql` at all. Raised at backend construction for the lock-plan case; at the point of the read for the session-fact case. | | `RECORDED_CLOCK_REQUIRES_WRITE_FENCE` | The store is constructed with `history: true` or `revisionTracking: true` — TypeGraph-owned recorded-clock allocation — against a backend whose write-fence plan resolves `unfenced`. | | `WRITE_FENCE_UNAVAILABLE` | A resolved plan cannot satisfy what a specific operation needs: either the plan is `unfenced` outright, or it is a `lock` plan whose `drain` is `"none"` meeting an operation whose `requires` is `"drain"`. `details.operation` names the operation and `details.requires` names which kind of exclusion (`"keyed"` or `"drain"`) it needed; a `drain: "none"` refusal also names the drain in the message. `"engine-serialized"` and `"caller-serialized"` satisfy either `requires` value without consulting `drain`. | | `CALLER_SERIALIZED_REFUSES_ADOPTION` | `adoptTransaction` was called on a backend whose resolved write-fence plan is `caller-serialized`. An externally owned transaction's lifetime cannot be held by the backend's in-process write-unit queue, so `store.withTransaction(externalTx)` is refused rather than let its writes silently interleave with the queue's own. `details.member` names `"adoptTransaction"`. | `RECORDED_CLOCK_REQUIRES_WRITE_FENCE` refuses at `createStore`, never mid-flush, and the message names the exact declaration line to add. `WRITE_FENCE_UNAVAILABLE` is not a `createStore`-time check: `requireWriteFence` is called from every individual lock site (the identity graph lock, the identity-enablement drain, identity DDL, trusted import, contribution DDL, recorded-clock allocation, schema-fence sites, graph-merge provenance), so it fires wherever one of those runs — inside a live transaction, mid-operation, not only at `createStore`. `WRITE_FENCE_SQL_UNAVAILABLE` and `WRITE_FENCE_DECLARATION_INVALID` both refuse earlier, at backend construction for a `createSqlBackend`-built backend (or, for the session-fact half of `WRITE_FENCE_SQL_UNAVAILABLE`, at the read that needed it), since they are about the declaration itself rather than what a specific store option or operation requires of it. `CALLER_SERIALIZED_REFUSES_ADOPTION` fires wherever `adoptTransaction` is actually called, which is never at `createStore` time. `IDENTITY_REQUIRES_WRITE_FENCE` is another write-fence-related code — see the Operational Identity guard codes table above — but is not in this table because it guards identity construction, not recorded-clock allocation. ### Caller-serialized queue codes The in-process queue a `writeFence: { mechanism: "caller-serialized" }` declaration builds (`src/backend/serialized-execution-queue.ts`) raises two more `ConfigurationError` codes, both naming `details.subject` — the SQLite dialect string for SQLite's own per-connection queue, or `"caller-serialized"` for the write-unit queue a `caller-serialized` declaration builds: | `details.code` | Raised when | | --- | --- | | `SERIALIZED_QUEUE_REENTRANT_SUBMISSION` | A queued operation was awaited from inside a transaction already running on the same queue — the transaction holds the queue's execution slot until it completes, so the nested operation could never run. Use the transaction-scoped context (`tx.nodes` / `tx.edges` / `tx.backend`) instead of the root store or backend inside a `store.transaction` callback, or move the operation outside the transaction. | | `CALLER_SERIALIZED_REQUIRES_ASYNC_CONTEXT` | The queue's reentrancy detection depends on `node:async_hooks`' `AsyncLocalStorage`, which is unavailable on this runtime (or had not finished loading). A `caller-serialized` write-fence declaration's in-process promise depends on that detection actually working, so every submission is refused rather than run without it. SQLite's own per-connection queue never raises this code: it runs without detection instead of refusing when the context is unavailable. | These codes are not part of `RECORDED_CAPTURE_GUARD_CODES` — that set is closed to the three codes documented under [Recorded-capture guard codes](#recorded-capture-guard-codes) below, and `isRecordedCaptureGuardError` does not recognize any write-fence code. ### Optimistic-retry unit codes The retry owner every `"optimistic-retry"`-tier unit of work runs through (`src/backend/capabilities/retried-unit.ts`) raises one more `ConfigurationError` code, naming `details.operation` — the same operation name `TransactionConflictError` reports for the same unit: | `details.code` | Raised when | | --- | --- | | `OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` | The unit's target is on the `"optimistic-retry"` execution tier (see [Backend Capabilities](/backend-setup#backend-capabilities)), and detecting a unit of work nested inside another one depends on `node:async_hooks`' `AsyncLocalStorage`, which is unavailable on this runtime. Running without that detection would let a nested unit's own independent retry commit against reads an outer attempt took before it ever conflicted, so the unit is refused, before its attempt ever runs, rather than run without it. A target on any other execution tier is unaffected: no nested owner exists there, so this code is never raised for it. | #### Backend capability declaration codes Custom backend declarations and capability bundles use stable `details.code` values when the declared surface disagrees with what TypeGraph can safely execute: | `details.code` | Raised when | | --- | --- | | `CAPABILITY_DECLARATION_CONTRADICTION` | `recursiveTraversal.supported` and its `reason` contradict each other: unsupported without a reason, or supported with a dangling reason. | | `RECURSIVE_TRAVERSAL_UNSUPPORTED` | A backend declares recursive traversal unsupported and a recursive query, subgraph read, or historical identity operation needs it. `details.operation` names the refusing path and `details.reason` echoes the backend declaration. | | `CONSTRAINT_CLAIM_SURFACE_MISMATCH` | The `constraintClaims` declaration and the claim members implemented by the backend disagree in either direction. | | `BUNDLE_PORT_SURFACE_MISMATCH` | A non-claim capability bundle resolves a required member as present, but the backend port used by the operation cannot reach it. Fallback-disposition members degrade through their documented fallback instead of throwing this code. | | `RECORDED_DDL_CONSTRAINT_NAME_MISMATCH` | `recordedTableDdl` names a primary-key constraint for only one of the temporary or final recorded-table name sets. | The recorded-time preview migration also throws `UnsupportedBackendCapabilityError` with `details.capability: "recordedTableDdl"` when a legacy schema needs rewriting and the custom backend does not provide its DDL callback. See [Migrating Preview Recorded Time](/schema-management#migrating-preview-recorded-time) and [Capability bundles](/backend-setup#capability-bundles) for the corresponding migration and backend-author guidance. #### Approximate retrieval with a mismatched metric `.similarTo(vector, k, { approximate: true, metric })` is refused with a `ConfigurationError` when `metric` differs from the field's declared metric. An ANN structure is built for one metric — `vec0` bakes `distance_metric` into the virtual table, libSQL's DiskANN index is built with `metric=…`, pgvector's index carries a per-metric operator class — so retrieving by the declared metric and re-scoring under the override would return the declared metric's neighbors wearing the override's scores. The two options state something that cannot both hold, so the option is refused rather than downgraded to an exact scan behind the caller's back. `details` carries `nodeKind`, `fieldPath`, `requestedMetric`, `declaredMetric`, and `indexType`; there is no stable `details.code`, so match by class and `details`. A slot declared `indexType: "none"` is **not** refused — there is no ANN structure to be bound to a metric, and the opt-in compiles to the exact scan, a degradation stated on the `approximate` option itself. A mismatched metric with no `approximate` is not refused on the query builder either; `store.search.vector` and `store.search.hybrid` refuse every mismatched override on their own broader rule. See [Approximate retrieval](/semantic-search#approximate-retrieval-for-similarto-opt-in). #### Durable edge match identity guard codes Durable edge match identity uses stable `ConfigurationError` detail codes: | `details.code` | Meaning | | --- | --- | | `EDGE_MATCH_IDENTITY_VALUE_NOT_SCALAR` | A declared identity field cannot be represented as a portable JSON scalar, or an untyped runtime value violated that declaration. | | `EDGE_MATCH_IDENTITY_KEY_TOO_LARGE` | One complete durable identity tuple exceeds the portable 2,000-byte index budget. Normal import records this against the individual edge; trusted import is atomic and refuses the whole stream. | | `EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE` | The adapter declares durable identity support, but the database is missing its columns or unique arbiter. Initialize or migrate the schema before serving writes. | | `EDGE_MATCH_IDENTITY_REQUIRES_ATOMIC_BACKEND` | Initial adoption needs an atomic empty-kind fence or materialization preflight that the custom backend does not implement. | | `DURABLE_EDGE_MATCH_IDENTITY_COMMAND_UNSUPPORTED` | A custom backend declares durable identity support but refuses the authoritative convergence command. TypeGraph fails closed because the portable read-then-write fallback has no equivalent database arbiter. | | `IMPORT_EDGE_BATCH_RETRY_REQUIRES_SAVEPOINT` | A durable import batch was refused without savepoint rollback protection, either because the backend is non-transactional or because its root/transaction statement-execution contract cannot serve savepoints. TypeGraph will not retry rows individually because that could double-attribute an already-written prefix. | The last refusal deliberately differs from optional fused-command fallback: the durable identity declaration delegates correctness to a database key, so a backend that claims the feature but refuses its command cannot safely re-enter the dynamic portable path. #### Heterogeneous edge-read guard codes `findEdgesByHeterogeneousEndpointSet` refuses mixed endpoint modes instead of guessing how incident and exact-pair rows should be interpreted: | `details.code` | Meaning | | --- | --- | | `EDGE_HETEROGENEOUS_READ_MIXED_ENDPOINT_MODES` | The request contains both incident-endpoint rows (without an opposite endpoint) and exact directed-pair rows. Supply an opposite endpoint for every row to request exact-pair matching. | | `EDGE_HETEROGENEOUS_READ_BIND_BUDGET_EXCEEDED` | The endpoint set cannot fit within the backend's bind-parameter budget. Split the request into smaller calls. | #### Operational Identity guard codes Operational Identity lifecycle failures use stable `details.code` values on `ConfigurationError`: | `details.code` | Meaning | | --- | --- | | `IDENTITY_REQUIRES_ATOMIC_BACKEND` | The selected adapter cannot provide the interactive transaction required by identity writes. | | `IDENTITY_REQUIRES_STATEMENT_EXECUTION` | The backend cannot execute the raw statements Operational Identity issues internally. | | `IDENTITY_REQUIRES_WRITE_FENCE` | Operational Identity was constructed against a backend whose `capabilities.writeFence` resolves `unfenced` — declare the capability, matching the engine's real locking support. See [Write-fence declaration codes](#write-fence-declaration-codes). | | `IDENTITY_NOT_ENABLED` | `store.identity`, `tx.identity`, `StoreView.identity`, or an identity-expanded query option was reached on a graph without `identity: { ... }` — normally caught at compile time; this is the runtime guard for a widened or `any`-typed handle. | | `IDENTITY_STORAGE_MISSING` | An identity relation disappeared after enablement, or exists without this graph's fill. Restore ledgers, or recreate and rebuild the derived closure, before serving traffic. `details.reason: "unfilled"` marks the second case: the separation relation is present but holds no row for this graph while the ledger holds a live `different` assertion across two distinct identity classes — reopen the Store (the open runs the fill) or run `rebuildIdentityClosure(store)`. A Store handle opened while the relation did not exist keeps failing until it is reopened, which is deliberate: the alternative is a confident "not separated" the moment another graph's upgrade creates the shared relation. | | `IDENTITY_UPGRADE_REQUIRES_ATOMIC_DDL` | The backend cannot publish the derived separation relation's upgrade — the `CREATE` and the fill — as one commit, on a graph that owes rows. `details.missingPorts` names what is absent: `schemaWriteTransaction` / `identityTableDdl` on the fenced path, or `executeSchemaDdl` on the schema-commit path. Refused rather than degraded, because a relation created empty and filled afterwards reads as "nothing is separated" in between. Both bundled Drizzle backends implement all three when transactions are enabled, so this is a custom-backend path. | | `IDENTITY_ENABLEMENT_PENDING` | First enablement is pending because `autoMigrate` is disabled. | | `IDENTITY_PROFILE_MIGRATION_PENDING` | A `sameIdAcrossKinds` change (a breaking `fold`↔`ignore` flip, or disabling identity) has not been applied — either it is breaking, or `autoMigrate` is disabled. | | `IDENTITY_SCHEMA_MIGRATION_PENDING` | An identity-relevant ontology change is pending because `autoMigrate` is disabled. | | `IDENTITY_SEPARATION_VIOLATION` | The derived separation relation refused a write that would place both endpoints of a current `different` assertion in one identity class. The database-level backstop beneath identity validation; reaching it means an earlier guard let a contradiction through. | | `IDENTITY_TRANSACTION_NOT_WRITE_FENCED` | SQLite refused an identity write because the enclosing transaction was begun `DEFERRED` and another connection committed before it could take the writer slot. Only reachable through `store.withTransaction(externalTx)` / `store.withRecordedTransaction(externalTx)`, where the caller owns the `BEGIN` — TypeGraph's own transactions open `BEGIN IMMEDIATE` and hold the slot from the start. SQLite cannot upgrade a stale snapshot in place, so roll back and re-run the transaction, opening it with `BEGIN IMMEDIATE`. | | `IDENTITY_SCHEMA_CONTRADICTION` | Existing nodes or assertions contradict the proposed identity profile or ontology, or the materialized closure disagrees with the assertions it was derived from. Run `rebuildIdentityClosure(store)` to recover from a closure mismatch. | | `IDENTITY_IMPORT_REQUIRES_PROFILE` | An interchange document carries an `identity` section but the target graph does not have the profile enabled. | | `IDENTITY_MERGE_REQUIRES_PROFILE` | A branch carries identity changes but the merge target graph does not have the profile enabled. | | `IDENTITY_EXPORT_REQUIRES_TEMPORAL_FIELDS` | An identity-enabled export explicitly disabled temporal fields. Remove `includeTemporal` or set it to `true`; endpoint bounds are required to validate assertion windows on import. | | `IDENTITY_IMPORT_ID_CONFLICT` | An imported assertion id already exists in the target ledger identifying different truth (relation, endpoints, or validity window). | | `RECORDED_IDENTITY_SCHEMA_MISSING` | A `history: true` open of an identity-enabled graph could not find the recorded identity relation. Bundled backends provision it, so this is rare there and more likely on a custom backend. | When an unapplied migration's **only** breaking change is the identity one, the specific pending code above wins over the generic `MigrationError` (which is attached as `cause`); a diff that also breaks nodes, edges, ontology, or indexes raises the generic `MigrationError` enumerating all of them. Identity import also raises `ValidationError` with one of these `details.issues[].code` values when an interchange document's `identity` section fails shape or integrity checks. Each issue carries the offending assertion's id structurally in `details.issues[].assertionId`, and `importGraph`/`importGraphStream` record these failures as `entityType: "identity"` entries in `result.errors` (a self-assertion — `IDENTITY_SELF_ASSERTION` — included) rather than throwing: | Issue `code` | Meaning | | --- | --- | | `IDENTITY_IMPORT_UNKNOWN_KIND` | An assertion endpoint names a node kind not in the target graph's registry. | | `IDENTITY_IMPORT_PAIR_NOT_NORMALIZED` | An assertion's `a`/`b` endpoints are not in code-point order. | | `IDENTITY_STATE_IMPORT_ENDED_ASSERTION` | A `state`-mode import (the default) contains an already-ended assertion; use `identityMode: "archival"` on export to carry ended assertions. | | `IDENTITY_IMPORT_FUTURE_VALID_FROM` | An open (current) assertion's `validFrom` is in the future, in either import mode. | | `IDENTITY_IMPORT_FUTURE_VALID_TO` | An ended assertion's `validTo` is in the future. | | `IDENTITY_IMPORT_INVALID_WINDOW` | An assertion's `validTo` precedes its `validFrom`. | | `IDENTITY_IMPORT_ENDED_BY_WITHOUT_END` | An assertion names an `endedBy` cause but carries no `validTo`; only an ended assertion has a cause. | | `IDENTITY_IMPORT_ENDED_BY_NOT_ENDPOINT` | An assertion's `endedBy` names a node that is not one of its own endpoints; a deletion cascade only ends assertions that touch the deleted node. | | `IDENTITY_SELF_ASSERTION` | An assertion's `a` and `b` name the same node. | #### Merge provenance sidecar codes `persistProvenance: true` writes to a *sidecar* graph beside the merge target, and `openProvenanceStore` refuses any sidecar graph id it cannot prove it owns. Both refusals are `ConfigurationError`s with a stable `details.code`, and both carry `details.graphId` (the sidecar id) and `details.targetGraphId`: | `details.code` | Meaning | | --- | --- | | `GRAPH_MERGE_PROVENANCE_ID_COLLISION` | The sidecar graph id is occupied by something this library did not write. `details.reason` names which state was found, and the suggestion is specific to it. | | `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED` | The backend exposes no transactional schema fence (`schemaWriteTransaction`), so the id's emptiness check and its ownership-marker write cannot commit as one unit. Not a collision — the id may well be free. An already-owned sidecar still opens on such a backend, so read-only use of an existing sidecar stays available. | The five `details.reason` values on `GRAPH_MERGE_PROVENANCE_ID_COLLISION`: | `details.reason` | The state that was found | | --- | --- | | `application-graph` | The id holds rows (in any per-graph table) or a schema that is not the sidecar's, so it belongs to an application. When a pre-marker sidecar is classified, revision-change journal entries that record its own stored `Provenance` rows are not counted; every other journal entry is. Rename the colliding graph or point the merge elsewhere. | | `empty-legacy-sidecar` | A pre-marker sidecar with no rows at all, which carries no evidence of authorship and is indistinguishable from an application graph of the same shape. | | `unupgradeable-legacy-sidecar` | A pre-marker sidecar whose rows do not verify as provenance this library wrote for *this* target, so it cannot be upgraded to an owned sidecar. | | `unowned-exact-schema-graph` | The current sidecar schema with no ownership marker. Because the marker is written *first*, this library cannot have produced this state; contents are not consulted, so an empty or provenance-shaped occupant is refused too. | | `corrupt-ownership-marker` | A `ProvenanceOwner` row that is not a valid live claim for this target — soft-deleted, schema-invalid, naming a different target, or stored under a different row id. It is never overwritten or resurrected, because it may be an application's row. | Under `persistProvenance: true` these arrive wrapped: the sidecar is opened and claimed **before** the merge commits, and either code refuses the merge as an `InvalidMergeOptionsError` (`details.option: "persistProvenance"`, `details.provenanceErrorCode` echoing the code above, the `ConfigurationError` as `cause`) with the target left unmodified. Only transient row-write failures after the commit degrade to a `warnings` entry. #### Interchange serialized-connection guard codes Two long-lived interchange streams cannot share one serialized database connection: an export snapshot holds a read transaction for the whole stream and a streaming import writes a transaction per chunk on that same connection, so the second one either nests a `BEGIN` or waits for a slot that never frees. The lease is **exclusive** — one stream of any kind per connection — so all four pairings refuse with a `ConfigurationError` rather than hanging: | `details.code` | Raised when | | --- | --- | | `INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` | An export snapshot holds the connection, detected through the shared serialized resource the two backend wrappers were marked with. | | `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT` | The same condition, reported by the object-identity detector: one SQLite backend is exporting into itself. Worth telling apart because the fix differs — pass a second backend rather than await whatever else is running. | | `INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS` | A streaming import holds the connection, in either order of discovery. | The code names *what holds the connection*; `details.requested` and `details.heldBy` (each `"export-snapshot"` or `"import-stream"`) name which pairing was actually refused, so a same-kind refusal is never reported as something it is not. `details.graphId` names the graph the refused stream was for. `"import-stream"` is the kind of every long-lived import, not only `importGraphStream`: `importGraph` holds the lease for the whole call, and `trustedImportGraph` / `trustedImportGraphStream` hold it for the whole trusted session — so those APIs throw this `ConfigurationError` as well as their own `TrustedImportError`. Connections TypeGraph cannot observe are not refused: two clients dialed at one server, or two SQLite handles on one file, are genuinely independent. See [Scaling branches and interchange](/graph-merge#scaling-branches-and-interchange) for which drivers are recognized as serialized. ##### Declaring a connection the driver hides Recognition is a duck-type over the client object, so a serialized driver TypeGraph cannot identify (`expo-sqlite`, `op-sqlite`, `sqlite-proxy`, `pg-proxy`, Bun `SQL`, a postgres-js client capped through a string it does not coerce) is left unmarked and its stream pairs are not refused. `createSqliteBackend` and `createPostgresBackend` accept a `serializedResource` declaration for that gap — `{ mode: "shared", resource: client }` — and for the reverse case, `{ mode: "independent" }`, when the detection is wrong for your topology. See [Serialized connections](/backend-setup#serialized-connections). The declaration is applied or refused, never quietly ignored: | Declaration | Outcome | | --- | --- | | `{ mode: "shared", resource }` on a connection TypeGraph did not detect, or naming the client it did detect | The named object is the serialized resource; two backends naming the same object are one connection | | `{ mode: "shared", resource }` naming a **different** object than the one detected | `ConfigurationError` (`code: "CONFIGURATION_ERROR"`) from the factory, with `details.reason: "serialized-resource-conflict"` and `details.declaredKind` / `details.detectedKind` naming what each side was | | `{ mode: "independent" }` | Honored, whatever was detected — the documented escape hatch | The conflict is refused rather than resolved because two wrappers over one connection given two different sentinels would stop being seen as a pair, which is precisely the refusal this guard exists to make. The two `*Kind` details are constructor names (`"Database"`, `"BoundPool"`), not the handles themselves: `details` is what `toLogString()` serializes, and a driver handle there would print whatever that driver stores — a `pg.Pool` keeps its `connectionString`, password included — into your logs. `{ mode: "independent" }` lifts the shared-resource arm between two distinct backend objects. It does **not** lift `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`: one SQLite backend exporting into itself holds the one snapshot transaction its own import writes through, which is a fact about a single handle rather than a claim about connection topology. Pass a second backend for that case. That surviving refusal is SQLite-only, so on PostgreSQL a backend declared independent exporting into itself is not refused either — a client that hands out independent connections is exactly what the declaration claims. #### `ExportStreamCancelledError` An export stream whose `signal` fires settles with `ExportStreamCancelledError` (`code: "INTERCHANGE_EXPORT_STREAM_ABORTED"`) rather than a silent end of stream, so a consumer never mistakes a cancelled export for a complete one. It is thrown only *after* the export has given back everything it took, so receiving it means the connection is already free. What that was depends on the backend: a transactional one rolls back the snapshot and releases the connection's stream lease; one without transactions held neither and simply abandons its remaining reads, its delivered chunks never having been a single snapshot. The message says which. `details.graphId` names the exported graph and `cause` carries the signal's own `reason` when the caller supplied one. A signal that is already aborted refuses the export before any transaction is opened. See [Cancelling an export](/interchange#cancelling-an-export). #### `ExportStreamIdleTimeoutError` An `exportGraphStream` configured with `idleTimeoutMs` settles with `ExportStreamIdleTimeoutError` (`code: "INTERCHANGE_EXPORT_STREAM_IDLE_TIMEOUT"`) when its consumer does not request another chunk within that bound. The timeout measures only the interval after a chunk is yielded; time spent waiting for the backend to produce the next chunk does not count. `details.graphId` identifies the graph and `details.idleTimeoutMs` carries the configured bound. As with explicit cancellation, a transactional export rolls its snapshot back and releases its stream lease before the error is delivered; a non-transactional export held neither and abandons its remaining reads. See [Cancelling an export](/interchange#cancelling-an-export). #### Recorded-capture guard codes `ConfigurationError` is intentionally open-shaped, but the guards that fire on a `history: true` / `revisionTracking: true` store carry a **stable, branchable `details.code`** so a portable caller does not have to substring-match the message. The three codes are exported as a set, `RECORDED_CAPTURE_GUARD_CODES`, and reachable through the `isRecordedCaptureGuardError` type guard: | `details.code` | Raised when | |----------------|-------------| | `RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION` | `store.withTransaction(externalTx)` on a history-enabled store — it has no flush point before the caller commits. Use `store.withRecordedTransaction(externalTx, fn)`. (Also a compile error on an `AdapterHistoryStore`.) | | `RECORDED_CAPTURE_RAW_SQL_DISABLED` | A raw SQL escape (`tx.sql`, `backend.executeStatement` / `executeDdl`) on a history-enabled store, where it would bypass recorded-time capture. | | `REVISION_TRACKING_RAW_SQL_DISABLED` | The same raw SQL escape on a revision-tracked store, where it would bypass the revision anchor. | Typed code cannot call `withTransaction` on an `AdapterHistoryStore`; use `withRecordedTransaction` directly. The runtime code remains useful at JavaScript and deliberately untyped boundaries. If one of those boundaries throws, `isRecordedCaptureGuardError(error, "RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION")` narrows both the error and its `details.code` without message matching. Pass a specific code to narrow to one guard, or omit it to match any. The guard narrows `error` to a `ConfigurationError` whose `details.code` is the passed `RecordedCaptureGuardCode` (or the full union when no code is given), so no untyped `details` spelunking is needed. This composes with [`tx.sqlAvailability`](/queries/temporal/#raw-sql-under-history-capture): the discriminant tells a caller *why* `tx.sql` is unusable ahead of time (`"history"` / `"revisionTracking"` vs. `"unavailable"` for a backend with no transactions), while the guard code identifies a guard that has already thrown. Between them, "history capture forbids raw SQL here" and "this backend has no transactions" (which carries **no** guard code) are cleanly distinguishable without catching-and-string-matching. #### Engine-native recorded-time codes A backend can track recorded (system) time itself by declaring `GraphBackend.recordedTime` instead of using TypeGraph's own recorded relations and clock — see [Engine-native recorded time](/queries/temporal#engine-native-recorded-time) and [Supplying `recordedTime`](/backend-authoring#supplying-recordedtime). Which ownership form a store reads under is derived from that member's presence, never declared separately, so there is no `recordedTimeOwnership` option to set. Every refusal specific to that form carries a stable `details.code`: | `details.code` | Raised when | |----------------|-------------| | `ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE` | An engine profile declares `recordedTime` without also declaring `lineage` — engine-native history keeps no recorded relations of its own for TypeGraph to derive a graph-merge change delta from. Raised at backend construction, naming both members. | | `RECORDED_TIME_UNAVAILABLE` | A caller reached `requireRecordedTime` and found `recordedTime` absent on the backend it asked — store construction under `history: true` and the shared `recordedNow()`/`revisionNow()`/receipt-stamping read, both reached only once ownership has already resolved to `"engine-native"`. | | `ENGINE_NATIVE_REVISION_TRACKING_UNSUPPORTED` | A store is constructed with `revisionTracking: true` against an engine-native backend, whether or not `history: true` is also requested — there is no TypeGraph clock for `revisionTracking` to advance; the engine's own revision is available only under `history: true`. | | `ENGINE_NATIVE_RECORDED_READ_UNSUPPORTED` | A store is constructed with an external `recordedRead` binding against an engine-native backend — there is no TypeGraph recorded relation for one to populate. | | `ENGINE_NATIVE_RECORDED_IDENTITY_UNSUPPORTED` | `store.identityAtCoordinate` at a past recorded instant, or the query compiler's historical identity traversal, is reached under engine-native recorded time — identity history reads TypeGraph's own recorded relations directly, which an engine-native backend does not populate. | | `ENGINE_NATIVE_MIGRATE_RECORDED_TIME_UNSUPPORTED` | `migrateLegacyRecordedTime` is called against an engine-native backend — the migration rewrites TypeGraph's own recorded relations, which an engine-native backend does not have. | | `RECORDED_INSTANT_OWNERSHIP_MISMATCH` | `store.asOfRecorded(instant)` receives an instant minted under the OTHER recorded-time ownership form — an `r1:` (TypeGraph-owned) instant against an engine-native store, or an `e1:` (engine-native) instant against a TypeGraph-owned store. | The profile refusal fires at backend construction; the two store-option refusals, and `RECORDED_TIME_UNAVAILABLE`'s construction arm, fire at `createStore`; the remaining codes, and `RECORDED_TIME_UNAVAILABLE`'s read arm, fire at the specific call that cannot be honored. None of these codes are members of `RECORDED_CAPTURE_GUARD_CODES` above — that set stays closed to the three TypeGraph-capture guards. ### `SchemaMismatchError` Thrown when the database schema doesn't match the expected graph definition. ```typescript try { const [store] = await createStoreWithSchema(graph, backend); } catch (error) { if (error instanceof SchemaMismatchError) { console.log(error.category); // "system" console.log(error.details); // { graphId: "my-graph", expectedHash: "", actualHash: "" } console.log(error.suggestion); // "Run migrations to update the database schema..." } } ``` ### `MigrationError` Thrown when schema migration fails due to breaking changes that require manual intervention. The `details.reason` value `"edge-match-identity-rekey"` means a populated edge kind changed or newly adopted its durable match identity. Existing rows cannot be assigned new identity keys without choosing how conflicts converge. Export the affected edges, hard-delete them, apply the schema migration, then reimport them so TypeGraph materializes and arbitrates the new durable keys. ```typescript try { const [store] = await createStoreWithSchema(graph, backend); } catch (error) { if (error instanceof MigrationError) { console.log(error.category); // "system" console.log(error.details); // { graphId: "my-graph", fromVersion: 3, toVersion: 4, reason: "Removed required field 'email' from Person" } console.log(error.suggestion); // "Review the breaking changes and perform manual migration if needed..." } } ``` ### `BaseSchemaMigrationError` Thrown by zero-DDL verified and graph-template entry points when the deployment-wide physical TypeGraph schema has not been adopted to the version required by the running library. This is separate from `MigrationError`, which describes one graph's serialized schema evolution. ```typescript try { const [store] = await createVerifiedStore(graph, backend); } catch (error) { if (error instanceof BaseSchemaMigrationError) { console.log(error.details); // { // installedVersion: undefined, // requiredVersion: 1, // reason: "missing" // } } } ``` `reason` is `"missing"`, `"stale"`, or `"newer"`. For missing or stale storage, run `createStoreWithSchema()` or `createAdapterStoreWithSchema()` once under a DDL-capable role, or apply the published external base-schema migration. A newer marker requires a TypeGraph release that supports that version. ## Query Errors ### `UnsupportedPredicateError` Thrown when using a query predicate that isn't supported by the current backend. ```typescript // Using vector similarity on a backend without vector support: try { await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(queryVector, 10)) .select((ctx) => ctx.d) .execute(); } catch (error) { if (error instanceof UnsupportedPredicateError) { console.log(error.category); // "system" console.log(error.suggestion); // "Use a backend that supports this predicate, or rewrite the query..." } } ``` ## Transaction Errors ### `TransactionClosedError` Thrown when a statement reaches a transaction-scoped backend after its transaction boundary has already returned. A transaction pins one database connection, which carries one statement at a time. When `store.transaction(...)` resolves or rejects, the driver emits `COMMIT` or `ROLLBACK` on that connection and hands it back to the pool. Any statement still in flight then has nowhere safe to go — it would execute inside somebody else's transaction — so TypeGraph refuses it. The usual source is a callback that lets work escape it. `Promise.all` rejects on its first rejection while its siblings keep running: ```typescript await store.transaction(async (tx) => { // If `a` fails, `b`'s remaining statements are orphaned. await Promise.all([tx.nodes.Doc.create(a), tx.nodes.Doc.create(b)]); }); ``` You will normally never see this error: `Promise.all` has already rejected with the original failure and discards the orphan's. It surfaces only if you await the orphaned promise yourself. To avoid orphaning writes at all, use `Promise.allSettled` and inspect the results, or await the writes in sequence. `adoptTransaction()` never closes its queue — only the caller knows when their transaction ends — so this error cannot arise there. It remains the caller's job to await every graph write before committing. ### `TransactionConflictError` Thrown when a transaction was aborted by a serialization failure or deadlock on every attempt available to it. `details.operation` names the transaction that failed, `details.attempts` the number tried, and `cause` is the last attempt's driver error — PostgreSQL's own protocol for both conditions is to re-run the whole transaction from the top, which is what this error reports as exhausted. `store.transaction()` and `store.transactionWithReceipt()` raise it with `attempts: 1` for a conflict on their single attempt; passing `retry: { attempts }` (see [Retrying on conflict](/schemas-stores/#retrying-on-conflict)) raises it only once every attempt has conflicted. Graph-merge's commit paths raise `MergeError` on the same exhaustion, with a `TransactionConflictError` as its `cause`. ```typescript try { await store.transaction(fn, { retry: { attempts: 3 } }); } catch (error) { if (error instanceof TransactionConflictError) { console.log(error.details.attempts); // 3 console.log(error.cause); // the last driver error } } ``` **The serialization covers TypeGraph's own statements, not `tx.sql`.** The raw Drizzle handle you get for writing your own relational tables in the same transaction shares the one pinned connection but bypasses the queue. Running a raw statement concurrently with a graph write — or with another raw statement — races two queries on that connection (the overlap `pg@9` removes), and the boundary cannot drain a raw statement it never saw. Await each `tx.sql` statement before the next write. ## Error Handling Patterns ### Using Error Utilities TypeGraph provides utility functions for common error handling patterns: ```typescript import { isTypeGraphError, isUserRecoverable, isConstraintError, isSystemError, getErrorSuggestion, } from "@nicia-ai/typegraph"; try { await store.nodes.Person.create(data); } catch (error) { if (!isTypeGraphError(error)) { // Not a TypeGraph error, handle differently throw error; } // Get suggestion regardless of error type const suggestion = getErrorSuggestion(error); if (isUserRecoverable(error)) { // User can fix this by providing different input return { error: error.toUserMessage(), suggestion, }; } if (isConstraintError(error)) { // Business rule violation return { error: "This operation violates a constraint", details: error.details, }; } if (isSystemError(error)) { // Infrastructure/configuration issue console.error(error.toLogString()); throw error; } } ``` ### Catch Specific Errors ```typescript import { ValidationError, NodeNotFoundError, DisjointError, } from "@nicia-ai/typegraph"; try { await store.nodes.Person.create(data); } catch (error) { if (error instanceof ValidationError) { // Handle validation failure with contextual details return { error: "Invalid data", issues: error.details.issues, entity: error.details.kind, }; } if (error instanceof DisjointError) { // Handle constraint violation return { error: "ID already used by different type" }; } throw error; // Re-throw unexpected errors } ``` ### Check Error Codes ```typescript try { await store.nodes.Person.update(id, data); } catch (error) { if (error instanceof TypeGraphError) { switch (error.code) { case "NODE_NOT_FOUND": return { error: "Person not found" }; case "VALIDATION_ERROR": return { error: "Invalid data", issues: error.details.issues }; default: throw error; } } throw error; } ``` ### Transaction Error Handling ```typescript try { await store.transaction(async (tx) => { const person = await tx.nodes.Person.create({ name: "Alice" }); const company = await tx.nodes.Company.create({ name: "Acme" }); await tx.edges.worksAt.create(person, company, { role: "Engineer" }); }); } catch (error) { // Transaction is automatically rolled back on any error if (error instanceof ValidationError) { console.log("Validation failed, transaction rolled back"); console.log("Failed on:", error.details.kind, error.details.operation); } throw error; } ``` ## Contextual Validation Utilities For library authors or advanced use cases, validation utilities are available from the schema sub-export: ```typescript import { validateNodeProps, validateEdgeProps, wrapZodError, createValidationError, } from "@nicia-ai/typegraph/schema"; // Validate node properties with full context const validated = validateNodeProps(PersonSchema, inputData, { kind: "Person", operation: "create", }); // Wrap a Zod error with TypeGraph context try { schema.parse(data); } catch (zodError) { throw wrapZodError(zodError, { entityType: "node", kind: "Person", operation: "update", id: "person-123", }); } ``` ## Error Codes Reference | Code | Error Class | Category | Description | |------|-------------|----------|-------------| | `VALIDATION_ERROR` | `ValidationError` | user | Schema validation failed | | `DISJOINT_ERROR` | `DisjointError` | constraint | Disjointness constraint violated | | `IDENTITY_CONTRADICTION` | `IdentityContradictionError` | constraint | Identity mutation would make the assertion ledger contradictory | | `IDENTITY_VALIDITY_FUTURE_START` | `IdentityValidityWindowError` | user | Identity assertion starts after the operation clock | | `IDENTITY_VALIDITY_FUTURE_END` | `IdentityValidityWindowError` | user | Identity assertion ends after the operation clock | | `IDENTITY_VALIDITY_INVERTED` | `IdentityValidityWindowError` | user | Identity assertion ends before it starts | | `IDENTITY_VALIDITY_OPEN_WINDOW_CONFLICT` | `IdentityValidityWindowError` | constraint | A different open window already represents the current semantic pair | | `IDENTITY_ENDPOINT_VALIDITY` | `IdentityEndpointValidityError` | constraint | An endpoint does not cover the explicit assertion window | | `GRAPH_MERGE_IDENTITY_CONFLICT` | `IdentityMergeConflictError` | system | Branches carry opposing identity truth | | `GRAPH_MERGE_CONSTRAINT_CONFLICT` | `MergeConstraintConflictError` | constraint | The resolved merge would violate a store constraint | | `ENDPOINT_ERROR` | `EndpointError` | constraint | Invalid edge endpoint types | | `ENDPOINT_PAIR_ERROR` | `EndpointPairError` | constraint | Undeclared source/target combination | | `CARDINALITY_ERROR` | `CardinalityError` | constraint | Cardinality constraint violated | | `UNIQUENESS_VIOLATION` | `UniquenessError` | constraint | Uniqueness constraint violated | | `EDGE_MATCH_IDENTITY_CONFLICT` | `EdgeMatchIdentityConflictError` | constraint | A direct edge write collided with its declared endpoint/property identity | | `NODE_NOT_FOUND` | `NodeNotFoundError` | user | Referenced node doesn't exist | | `EDGE_NOT_FOUND` | `EdgeNotFoundError` | user | Referenced edge doesn't exist | | `KIND_NOT_FOUND` | `KindNotFoundError` | user | Unknown node/edge type | | `ENDPOINT_NOT_FOUND` | `EndpointNotFoundError` | user | Edge endpoint node doesn't exist | | `RESTRICTED_DELETE` | `RestrictedDeleteError` | constraint | Delete blocked by existing edges | | `CONFIGURATION_ERROR` | `ConfigurationError` | system | Invalid configuration | | `SCHEMA_MISMATCH` | `SchemaMismatchError` | system | Database schema mismatch | | `MIGRATION_ERROR` | `MigrationError` | system | Migration failed | | `BASE_SCHEMA_MIGRATION_REQUIRED` | `BaseSchemaMigrationError` | system | Deployment-wide base storage requires privileged adoption | | `UNSUPPORTED_PREDICATE` | `UnsupportedPredicateError` | system | Predicate not supported | | `UNSUPPORTED_BACKEND_CAPABILITY` | `UnsupportedBackendCapabilityError` | user | The backend does not advertise a capability the call needs. `details.capability` names it — for example `vector.searchFrontierTuning` for `efSearch` on any SQLite vector or hybrid search, where the engine has no per-search ANN frontier, with `details.reason` naming the limitation | | `INTERCHANGE_EXPORT_STREAM_ABORTED` | `ExportStreamCancelledError` | user | An export stream's `signal` fired, after the export gave back everything it took. On a transactional backend that is the snapshot transaction and the connection's stream lease; on one without transactions the export held neither and simply abandoned its remaining reads. The message says which | | `INTERCHANGE_EXPORT_STREAM_IDLE_TIMEOUT` | `ExportStreamIdleTimeoutError` | user | An export stream's consumer left a delivered chunk unacknowledged past its configured `idleTimeoutMs`; the export settled its snapshot and lease before reporting the timeout | | `TRANSACTION_CONFLICT` | `TransactionConflictError` | system | A transaction was aborted by a serialization failure or deadlock on every attempt available to it. `details.attempts` is the number tried; `cause` is the last driver error | # Architecture > How TypeGraph works internally and the design decisions behind it This page explains how TypeGraph works under the hood, the design decisions that shaped it, and why certain tradeoffs were made. ## High-Level Architecture ```text ┌────────────────────────────────────────────────────────┐ │ Your Application │ │ │ │ ┌──────────────────────────────────────────────────┐ │ │ │ TypeGraph Library │ │ │ │ │ │ │ │ ┌────────────┐ ┌────────────┐ │ │ │ │ │ Schema │ │ Query │ │ │ │ │ │ DSL │ │ Builder │ │ │ │ │ └──────┬─────┘ └─────┬──────┘ │ │ │ │ │ │ │ │ │ │ └──────────────┴───────────────┘ │ │ │ │ │ │ │ │ │ ▼ │ │ │ │ ┌──────────────────┐ │ │ │ │ │ Ontology Layer │ │ │ │ │ └──────────────────┘ │ │ │ └─────────────────────────┬────────────────────────┘ │ │ │ │ │ ▼ │ │ ┌────────────────────────┐ │ │ │ TypeGraph Backend Port │ │ │ └────────────┬───────────┘ │ │ │ │ │ ┌──────▼───────┐ │ │ │ SQL Adapter │ │ │ │ (Drizzle) │ │ │ └──────┬───────┘ │ │ │ │ └───────────────────────────┼────────────────────────────┘ │ ▼ ┌─────────────────┐ │ Your Database │ └─────────────────┘ ``` TypeGraph is an **embedded library**, not a database. It runs in your application process and compiles queries to SQL. A managed local Store can own its SQLite or PGlite connection; adapter integrations can instead use a connection your application already owns. ### Capability boundaries TypeGraph's public types match its supported runtime surfaces. The default `Store` exposes graph operations; `AdapterStore` adds native transaction interoperability; transaction contexts expose a read-only backend projection. Every store exposes `store.capabilities`, the read-only runtime feature descriptor used by its backend, so portable code can branch on atomicity, vector, fulltext, or analytics support without reaching through the adapter boundary. Whenever a surface loses capabilities, TypeGraph constructs an explicit allowlist projection. Proxy overlays are reserved for decorating a surface without changing its capabilities. The internal store and transaction ports are non-enumerable symbol properties and are absent from public TypeScript contracts. They are not a JavaScript security boundary: sufficiently reflective code can discover symbol properties, as it can inspect internals in any in-process library. The guarantee applies to all documented and supported access paths. ## Core Design Principles ### 1. Embedded, Not External **Decision**: TypeGraph is a library dependency, not a separate service. **Why**: Graph databases like Neo4j require managing another piece of infrastructure. For many use cases—knowledge bases, organizational structures, content relationships—the graph is part of your application, not a standalone system. **Tradeoff**: You don't get Neo4j's broad graph-data-science suite (community detection, betweenness centrality, and similar specialized analytics), but you avoid: - Additional deployment complexity - Network latency between app and graph - Separate scaling and monitoring - Data synchronization challenges ### 2. Schema-First, Type-Driven **Decision**: Zod schemas are the single source of truth. TypeScript types are inferred, not duplicated. **Why**: In many graph systems, you define types in one place, validation in another, and database schemas in a third. This leads to drift and bugs. With TypeGraph: ```typescript const Person = defineNode("Person", { schema: z.object({ name: z.string().min(1), email: z.string().email().optional(), }), }); // TypeScript type is inferred automatically type PersonProps = z.infer; // { name: string; email?: string } ``` The schema drives: - Runtime validation on create/update - TypeScript types for compile-time safety - Database storage format - Query builder type constraints ### 3. SQL as the Execution Engine **Decision**: Compile graph queries to SQL, don't implement a custom query engine. **Why**: SQLite and PostgreSQL are battle-tested, highly optimized query engines. Rather than building another one: ```typescript // Your query store.query() .from("Person", "p") .traverse("worksAt", "e") .to("Company", "c") .select((ctx) => ({ person: ctx.p.name, company: ctx.c.name })) // Compiles to SQL with CTEs WITH person_cte AS ( SELECT * FROM typegraph_nodes WHERE kind = 'Person' AND deleted_at IS NULL ), edge_cte AS ( SELECT * FROM typegraph_edges WHERE kind = 'worksAt' AND deleted_at IS NULL ), company_cte AS ( SELECT * FROM typegraph_nodes WHERE kind = 'Company' AND deleted_at IS NULL ) SELECT p.props->>'name' as person, c.props->>'name' as company FROM person_cte p JOIN edge_cte e ON e.from_id = p.id JOIN company_cte c ON c.id = e.to_id ``` This means: - You get database-level query optimization - Indexes work as expected - Transactions are ACID - You can analyze queries with EXPLAIN ### 4. Precomputed Ontology **Decision**: Compute transitive closures at store initialization, not query time. **Why**: Semantic relationships like `subClassOf` and `implies` form hierarchies. Computing "all subclasses of Media" during every query would be expensive. Instead, when you create a store: ```typescript const store = createStore(graph, backend); // ↑ Computes: // - subClassOf closure: Media → [Media, Podcast, Article, Video] // - implies closure: marriedTo → [marriedTo, partneredWith, knows] // - disjoint sets: Person ⊥ Organization ⊥ Product ``` These closures are stored in the `TypeRegistry` and used during query compilation: ```typescript .from("Media", "m", { includeSubClasses: true }) // At compile time, expands to: WHERE kind IN ('Media', 'Podcast', 'Article', 'Video') ``` **Tradeoff**: Changing the ontology requires recreating the store. But ontologies typically change rarely compared to instance data. ### 5. Homoiconic Schema Storage **Decision**: Store the graph schema and ontology as data in the database itself. **Why**: Most ORMs and graph libraries define schemas only in application code. The database stores data but has no record of what the data means. This creates problems: - You can't understand the database without reading the application source - Schema changes are invisible—no history, no diff, no audit trail - Exports require the application to interpret the data - Multiple applications can't share schema understanding TypeGraph takes a different approach: the schema is data. When you initialize a store, the complete schema (node types, edge types, property definitions, ontology relations, precomputed closures) is serialized to JSON and stored in `typegraph_schema_versions`: ```sql SELECT schema_doc FROM typegraph_schema_versions WHERE graph_id = 'my_graph' AND is_active = TRUE; ``` From TypeScript, [`getActiveSchema`](/schema-management#what-does-this-database-already-have) runs this query and parses the result into a typed `SerializedSchema`. The stored schema includes everything needed to understand the graph: ```typescript { graphId: "my_graph", version: 3, nodes: { Person: { properties: { /* JSON Schema */ }, ... }, Company: { ... } }, edges: { worksAt: { fromKinds: ["Person"], toKinds: ["Company"], ... } }, ontology: { relations: [{ metaEdge: "subClassOf", from: "Engineer", to: "Person" }], closures: { subClassAncestors: { Engineer: ["Person"] }, // ... precomputed inference data } } } ``` This enables: | Capability | How It Works | | ---------------------------- | ------------------------------------------------------------------------------------------------- | | **Self-describing database** | Query the schema without application code—useful for debugging, admin tools, and data exploration | | **Schema versioning** | Every schema change creates a new version; previous versions are preserved for auditing | | **Change detection** | Compare stored schema to code schema to detect additions, removals, and breaking changes | | **Portable exports** | The [interchange format](/interchange) is self-contained—importers know what the data means | | **Runtime introspection** | Applications can query the schema at runtime for dynamic UI, validation, or documentation | ```typescript import { getActiveSchema, getSchemaChanges } from "@nicia-ai/typegraph/schema"; // Query the active schema at runtime const schema = await getActiveSchema(backend, "my_graph"); console.log("Node types:", Object.keys(schema.nodes)); console.log("Edge types:", Object.keys(schema.edges)); // Detect pending changes before deployment const diff = await getSchemaChanges(backend, graph); if (!diff.isBackwardsCompatible) { console.error("Breaking changes require migration"); } ``` **Tradeoff**: Schema storage adds a small amount of database overhead (one JSON document per version). The benefit is a database that explains itself. ## Data Model ### Storage Schema TypeGraph uses two core tables: ```sql -- Nodes table CREATE TABLE typegraph_nodes ( graph_id TEXT NOT NULL, kind TEXT NOT NULL, id TEXT NOT NULL, props JSON NOT NULL, -- Properties as JSON version INTEGER NOT NULL, -- Optimistic concurrency valid_from TEXT NOT NULL, -- Temporal validity valid_to TEXT, created_at TEXT NOT NULL, updated_at TEXT NOT NULL, deleted_at TEXT, -- Soft delete PRIMARY KEY (graph_id, kind, id, valid_from) ); -- Edges table CREATE TABLE typegraph_edges ( graph_id TEXT NOT NULL, kind TEXT NOT NULL, id TEXT NOT NULL, from_kind TEXT NOT NULL, from_id TEXT NOT NULL, to_kind TEXT NOT NULL, to_id TEXT NOT NULL, props JSON NOT NULL, match_identity_name TEXT, match_identity_key TEXT, version INTEGER NOT NULL, valid_from TEXT NOT NULL, valid_to TEXT, created_at TEXT NOT NULL, updated_at TEXT NOT NULL, deleted_at TEXT, CHECK ( (match_identity_name IS NULL) = (match_identity_key IS NULL) ), UNIQUE (graph_id, kind, match_identity_name, match_identity_key), PRIMARY KEY (graph_id, kind, id, valid_from) ); ``` ### Why JSON for Properties? **Decision**: Store node/edge properties as JSON, not as columns. **Why**: 1. **Schema flexibility**: Adding a property doesn't require ALTER TABLE 2. **Heterogeneous nodes**: Different node kinds have different schemas 3. **Query simplicity**: One table for all nodes, not one per kind Both SQLite (JSON1 extension) and PostgreSQL (JSONB) have efficient JSON operators: ```sql -- PostgreSQL SELECT props->>'name' FROM typegraph_nodes WHERE props->>'status' = 'active'; -- SQLite SELECT json_extract(props, '$.name') FROM typegraph_nodes WHERE json_extract(props, '$.status') = 'active'; ``` **Tradeoff**: You can't create a B-tree index on a JSON property as easily as a column. For high-cardinality filtering, consider: - PostgreSQL: Expression indexes on JSONB paths - SQLite: Expression indexes on `json_extract(...)` (or generated columns) See [Indexes](/performance/indexes) for TypeGraph utilities to define and create these indexes. ### Durable edge match identity An edge registration can promote one endpoint/property comparison from a caller convention to a stored graph-schema contract: ```text graph matchIdentity declaration │ ▼ canonical endpoint/property key │ ▼ authoritative edge command ──► database unique arbiter │ │ └──────── created ◄───────┤ found ◄───────┘ ``` The database unique constraint is the concurrency authority. A cache, an outside read, or a process-local lock cannot replace it because serverless requests may use different processes and connections. Bundled root backends can therefore lower an eligible `getOrCreateByEndpoints()` call to one statement; paths with cardinality claims, history, revision tracking, or caller-owned work use the same arbiter inside their transaction. Each decision has one owner: | Decision | Owner | | --- | --- | | Which property fields form the identity | The graph registration's named `matchIdentity` | | Whether a schema can activate or re-key it | The schema manager's all-physical-row emptiness fence | | How endpoint/property values become a portable key | The shared canonical encoder | | Which concurrent writer owns the identity | The database unique constraint | | Whether a command created or found a row | The authoritative command result | | Whether a failed import batch may retry rows | The semantic savepoint result | The semantic savepoint is broader than raw SQL transaction control. Rolling it back restores the database savepoint and TypeGraph's pending capture touches, forced revisions, and graph-lock memo together. Consumers receive the resulting decision instead of re-deriving it from an exception, row count, or follow-up read. That is what keeps direct creates, convergence, bulk writes, import, history, and operation hooks aligned. ### Temporal Model Every node and edge tracks temporal validity: ```text ┌──────────────────────────────────────────────────────────────┐ │ Node: Article#123 │ ├──────────────────────────────────────────────────────────────┤ │ Version 1: "Draft" │ valid_from: 2024-01-01 │ │ │ valid_to: 2024-01-15 │ ├─────────────────────────┼────────────────────────────────────┤ │ Version 2: "Published" │ valid_from: 2024-01-15 │ │ │ valid_to: NULL (current) │ └─────────────────────────┴────────────────────────────────────┘ ``` When you update a node: 1. The current row's `valid_to` is set to now 2. A new row is inserted with `valid_from = now`, `valid_to = NULL` This enables: - **Point-in-time queries**: "What did the graph look like on January 10th?" - **Audit trails**: "What were all the versions of this article?" - **Soft deletes**: `deleted_at` marks deletion without losing history ## Query Compilation ### The Query Pipeline ```text Query Builder → Query AST → TypeGraph SQL Fragment → Adapter → Database ``` 1. **Query Builder**: Fluent API that constructs a typed AST 2. **Query AST**: A data structure representing the query (nodes, edges, predicates, projections) 3. **SQL Generator**: Transforms the AST into TypeGraph's immutable, database-independent SQL fragment representation 4. **Adapter**: Renders the fragment for SQLite or PostgreSQL and executes it through the configured driver ### Common Table Expressions (CTEs) TypeGraph compiles traversals to CTEs, which databases optimize well: ```typescript store .query() .from("Person", "p") .traverse("authored", "e") .to("Document", "d") .whereNode("d", (d) => d.status.eq("published")); ``` Becomes: ```sql WITH step_0 AS ( -- Start: all Person nodes SELECT * FROM typegraph_nodes WHERE graph_id = $1 AND kind = 'Person' AND deleted_at IS NULL ), step_1 AS ( -- Traverse: follow 'authored' edges SELECT e.*, s.id as _from_step FROM typegraph_edges e JOIN step_0 s ON e.from_id = s.id WHERE e.kind = 'authored' AND e.deleted_at IS NULL ), step_2 AS ( -- Arrive: at Document nodes SELECT n.*, s.id as _edge_id FROM typegraph_nodes n JOIN step_1 s ON n.id = s.to_id WHERE n.kind = 'Document' AND n.deleted_at IS NULL ) SELECT step_0.props->>'name' as person, step_2.props->>'title' as document FROM step_0 JOIN step_1 ON step_1._from_step = step_0.id JOIN step_2 ON step_2._edge_id = step_1.id WHERE step_2.props->>'status' = 'published'; ``` ### Recursive CTEs for Variable-Length Paths For `recursive()` traversals with cycle prevention enabled (the default), TypeGraph generates recursive CTEs like: ```sql WITH RECURSIVE path AS ( -- Base case: starting nodes SELECT id, 1 as depth, ARRAY[id] as path FROM typegraph_nodes WHERE kind = 'Person' AND id = $1 UNION ALL -- Recursive case: follow edges SELECT n.id, p.depth + 1, p.path || n.id FROM path p JOIN typegraph_edges e ON e.from_id = p.id JOIN typegraph_nodes n ON n.id = e.to_id WHERE e.kind = 'reportsTo' AND p.depth < 10 -- Implicit cap for unbounded traversal AND NOT n.id = ANY(p.path) -- Cycle detection ) SELECT * FROM path; ``` When you opt into `cyclePolicy: "allow"` and do not project a path column, TypeGraph can use a lighter recursive shape without path-array state and cycle predicates. ## Vector Search Architecture Semantic search with embeddings works across **all** backends — pgvector on PostgreSQL, sqlite-vec on better-sqlite3, and libSQL/Turso's built-in vector engine. The behavior is selected by a pluggable `VectorStrategy`, so adding a new backend is a single strategy object with no edits to the core. ### Pluggable Strategies Each backend wires a strategy that knows how to store embeddings and compile similarity queries: | Backend | Strategy | Storage / Index | Metrics | | -------------- | --------------------------------------------------------- | ------------------------------------------------------------------- | ------------------------- | | PostgreSQL | `pgvectorStrategy` (default) | typed `vector(N)` tables, HNSW / IVFFlat | cosine, l2, inner_product | | better-sqlite3 | `sqliteVecStrategy` (when the sqlite-vec extension loads) | `vec0` virtual tables (KNN) | cosine, l2 | | libSQL / Turso | `libsqlVectorStrategy` (wired automatically) | `F32_BLOB(N)`, DiskANN ANN via `libsql_vector_idx` + `vector_top_k` | cosine, l2 | `createSqliteBackend` and `createPostgresBackend` accept a `vector?: VectorStrategy` option to override the default. The strategies, `buildVectorCapabilities`, and the complete `VectorStrategy` / `VectorSlot` authoring vocabulary are exported from `@nicia-ai/typegraph/backend`. A backend advertises its vector support as data on `backend.capabilities.vector`: ```typescript backend.capabilities.vector; // { supported, metrics, indexTypes, maxDimensions, ... } ``` ### Storage Embeddings are stored in **per-field typed tables**, one per `(graphId, nodeKind, fieldPath)`, each carrying that field's fixed dimension. Tables are provisioned by the privileged migrator (`createStoreWithSchema`, and `evolve()` for runtime-added fields), with a durable contribution marker; the runtime hot path asserts the marker and never issues DDL. They are named `tg_vec___`. Graph-scoping the table name lets multiple graphs in one database declare the same kind+field at different dimensions without collision. ```sql -- PostgreSQL with pgvector: one table per (graphId, kind, field). The kind -- and field are encoded in the table name, so rows only key by node. CREATE TABLE tg_vec_my_graph_document_embedding ( graph_id TEXT NOT NULL, node_id TEXT NOT NULL, embedding vector(1536) NOT NULL, -- pgvector type, fixed dimension per field created_at TIMESTAMPTZ NOT NULL, updated_at TIMESTAMPTZ NOT NULL, PRIMARY KEY (graph_id, node_id) ); CREATE INDEX ON tg_vec_my_graph_document_embedding USING hnsw (embedding vector_cosine_ops); -- HNSW index for fast similarity ``` `generatePostgresMigrationSQL()` runs `CREATE EXTENSION IF NOT EXISTS vector` but creates no embedding table — the per-field tables are provisioned by `createStoreWithSchema` at boot (under the privileged role). ### Query Flow The query API is storage-transparent and unchanged across backends: ```typescript .whereNode("d", (d) => d.embedding.similarTo(queryVector, 10)) ``` Compiles to a backend-specific nearest-neighbor query — for example, on PostgreSQL: ```sql SELECT * FROM typegraph_nodes n JOIN tg_vec_my_graph_document_embedding e ON e.node_id = n.id AND e.graph_id = n.graph_id ORDER BY e.embedding <=> $1 -- Cosine distance LIMIT 10; ``` The backend's vector index (pgvector HNSW/IVFFlat, sqlite-vec `vec0`, or libSQL DiskANN) handles approximate nearest neighbor search efficiently. ## Performance Characteristics ### What's Fast - **Point lookups by ID**: O(1) with primary key index - **Traversal frontiers**: Set-based SQL rounds with database-managed joins and de-duplication - **Ontology expansion**: Precomputed at initialization, O(1) at query time - **Semantic search**: ANN indexes (pgvector HNSW/IVFFlat, sqlite-vec `vec0`, libSQL DiskANN) provide sub-linear search ### What's Slower - **Deep recursive traversals**: Recursive CTEs are more expensive than simple JOINs - **Whole-graph algorithms**: WCC, label propagation, and PageRank iterate over every visible node by default, or over an explicit `nodeKinds` induced subgraph, and their selected edges - **Large property filtering without indexes**: JSON extraction is slower than column access - **Cross-kind queries**: `includeSubClasses: true` increases the WHERE IN set ### Optimization Strategies 1. **Filter early**: Apply predicates as close to the source as possible 2. **Limit results**: Always paginate large result sets 3. **Use specific kinds**: Avoid `includeSubClasses` unless needed 4. **Index JSON paths**: For frequently-filtered properties, add expression indexes 5. **Batch writes**: Use transactions to reduce disk syncs and round-trips ### Mutation execution classes TypeGraph classifies bulk writes by what must be known before SQL can be submitted: - A **closed mutation program** carries every input, fence, and refusal rule needed for the database to decide the write. Eligible node/edge creates and soft deletes use one exact-resource execution profile and can run as one native atomic exchange on bundled serverless transports. The profile is attached to the exact backend object. Derived backends do not inherit it accidentally; an already-open PostgreSQL transaction earns a separate session-bound registration. - A **resolved mutation set** requires an authoritative database preimage and application computation before its writes are known. `bulkUpsertById()` is the canonical example: stored props are merged and Zod-validated, temporal decisions are derived, and repeated IDs observe earlier batch items. An eligible distinct-ID set can cross the exact registered boundary after resolution: its guarded SQL carries the node versions or complete edge preimages that justified the after-images. Update-only sets use one guarded set statement; mixed create/update sets add a terminal database postimage assertion inside the same native exchange, so an incomplete update aborts and rolls back its creates before the transport commits. Complex sets resolve and execute inside one interactive transaction. Neither shape is mislabeled as a read-free program. On an interactive PostgreSQL root, the operation runs on the exact collection-opened, caller-supplied, or adopted transaction and returns an explicit `applied | unsupported` verdict. `unsupported` proves that no program SQL ran before the complete portable path begins. This boundary keeps transport optimization subordinate to Store semantics. A new bulk optimization must either prove its mutation is closed or name the authoritative resolution phase it preserves; it cannot read on one connection and write on another, silently discard sidecars, or duplicate an eligibility predicate beside the profile owner. Backend authors can certify the transport boundary independently of mutation eligibility with the framework-agnostic atomic transport conformance runner. The runner supplies no dialect assumptions: the author provides statements, state observers, and exact-root provenance checks, while the shared checks verify ordered result slots, bound-parameter preservation, empty programs, and all-or-nothing rollback across primary and sidecar writes. ## Why These Tradeoffs? ### Why Not a Native Graph Database? Native graph databases (Neo4j, Amazon Neptune) excel at: - Very deep traversals (10+ hops) - Broad graph-data-science suites beyond the focused built-in algorithms - Massive scale (billions of nodes) TypeGraph is designed for: - Knowledge bases with thousands to millions of nodes - Shallow to medium traversals (1-5 hops typically) - Applications that already use SQL databases - Teams that want one database to manage ### Why a TypeGraph-Owned Backend Port? The schema DSL, Store, query compiler, and SQL fragments belong to TypeGraph. They do not import a database adapter's types. This keeps the public API stable and prevents consumers from typechecking declarations for drivers and dialects they never use. Drizzle remains an implementation detail of the built-in SQLite and PostgreSQL adapters: 1. **Driver integration**: Reuses mature SQLite and PostgreSQL connections 2. **Adapter-native access**: Bring-your-own-connection entrypoints retain precise Drizzle database and transaction types 3. **Replaceable boundary**: The core depends on TypeGraph ports; adapters translate fragments and operations at the edge The package exports that boundary directly. Schema-only packages can import the graph DSL and its schema-derived types from the Drizzle-free `@nicia-ai/typegraph/core` entrypoint. Backend and search-strategy authors can import the full Drizzle-free port vocabulary, including `GraphBackend`, `AdapterBackend`, `DialectAdapter`, and `SqlFragment`, from `@nicia-ai/typegraph/backend`. ### Why Zod for Schemas? 1. **Runtime validation**: Not just types, but actual validation 2. **Inference**: `z.infer` eliminates type duplication 3. **Composition**: Build complex schemas from simple ones 4. **Ecosystem**: Widely used, lots of integrations ## Next Steps - [Performance](/performance/overview) - Benchmarks and optimization tips - [Schemas & Stores](/schemas-stores) - Complete function signatures - [Integration Patterns](/integration) - How to integrate with your stack # Authoring an engine profile > Derive a variant of a bundled SQL engine profile, and what building one from scratch still requires [Backend Setup](/backend-setup) covers using the two bundled backends. This page is for adapting one: changing a lock spelling, loosening a declared capability, or swapping the resource-audit verdict without hand-copying every other field a profile carries. ## What a profile is A `SqlEngineProfile` is the data and dialect closures one SQL engine contributes before any backend object exists: dialect tokens, the execution adapter, transaction framing, DDL provisioning, strategies, capability declarations, and an opaque `assembly` wrapping the operation-backend builder. `createSqlBackend` is the one factory that turns a profile into a `GraphBackend`, and it owns everything that is the same for every engine: - **Capability derivation** — running the shared capability tail (atomic-batch detection, vector/fulltext capability shape, contribution-rebuild support) over the profile's own `declaredCapabilities`. - **Fence resolution** — building the one write-fence target for the whole backend and resolving its plan once, so every lock site and every transaction-scoped handle agrees on the same decision. - **Member assembly** — resolving the profile's `assembly` into its operation-backend builder and late-member factory, then assembling the mirrored member groups (contribution, identity, graph-template, base-schema, index-materialization, kind-removal, schema-version). - **Marks** — auditing the backend's resource shape and applying the trust marks (root-autocommit eligibility, schema-fenced-insert eligibility, first-party standing) that gate optimizations elsewhere. `createPostgresBackend` and `createSqliteBackend` are each `createSqlBackend` applied to a profile the bundled builders produce. ## The derivation path ```typescript import { buildPostgresEngineProfile, createSqlBackend, deriveEngineProfile, } from "@nicia-ai/typegraph/adapters/drizzle/engine"; const baseProfile = buildPostgresEngineProfile(db, options); const derivedProfile = deriveEngineProfile(baseProfile, { // one or more of the derivable fields below }); const backend = createSqlBackend(derivedProfile); ``` `buildPostgresEngineProfile` and `buildSqliteEngineProfile` are the derivation base: they build a real profile against a real connection, exactly the way `createPostgresBackend` / `createSqliteBackend` do internally. `deriveEngineProfile(base, overrides)` returns a fresh profile — `{...base, ...overrides}` — with `overrides` restricted to the fields listed below. `createSqlBackend` then assembles a backend from the result through the exact same path a bundled profile takes. If your own module re-exports a derived profile as an inferred-typed `const`, give it an explicit `SqlEngineProfile` annotation — the opaque `assembly` field's internal brand is not itself exported, so `tsc` cannot name it in a declaration file it has to infer. ## What you can override Each field below is read directly off the profile object (or off the `assembly`-derived context) by exactly one place in `createSqlBackend`, with the one carve-out below — so overriding it changes the whole backend consistently. | Field | What overriding it changes | | --- | --- | | `declaredCapabilities` | The capabilities `finalizeEngineCapabilities` derives the rest of the backend's advertised capabilities from — for example, declaring `writeFence` differently changes which write-fence plan resolves. | | `fenceSql` | The lock-statement spelling the resolved fence plan carries; pass `undefined` to remove it entirely (see [Removing `fenceSql`](#removing-fencesql) below). | | `resourceAudit` | The serialized-resource verdict `createSqlBackend` records before the backend escapes. | | `autocommit` | Whether a single statement outside an explicit transaction is durable — gates the root-autocommit mark. | | `contributionRuntime` | Deps for the contribution-marker member group. | | `identityRuntime` | Deps for the identity / recorded-relation member group. | | `graphTemplateRuntime` | Deps for the graph-template member group. | | `baseSchemaRuntime` | Deps for the base-schema lifecycle member group. | | `indexMaterializationRuntime` | Deps for the index-materializations member group. | | `kindRemovalRuntime` | Deps for the kind-removals member group. | | `close` | The backend's `close` member. | `DERIVABLE_ENGINE_PROFILE_KEYS` (exported alongside `DerivableEngineProfileKey` and `DerivableEngineProfileOverrides`) is the exact set above, as an `as const` array. ### The adapter-backed carve-out `declaredCapabilities` and `resourceAudit` are otherwise freely derivable, but `deriveEngineProfile` refuses an override that would change three of their sub-fields — `declaredCapabilities.maxBindParameters`, `declaredCapabilities.execution.interactiveTransactions`, and `resourceAudit.kind` — away from the base profile's own value, naming the sub-field (`ENGINE_PROFILE_OVERRIDE_UNSUPPORTED`). This check runs against any base profile, PostgreSQL or SQLite, but it exists for the bundled PostgreSQL builder: `buildPostgresEngineProfile` reads those exact three sub-values to compute its execution adapter's own options before the profile object exists, baking a copy of each into `profile.execution`, which is not itself derivable. Deriving from a SQLite base refuses the same override even though `buildSqliteEngineProfile`'s operation backend reads `maxBindParameters` off the resolved capabilities directly and would honor a changed value — the check does not distinguish the two dialects. Every other sub-field on both objects — `writeFence`, `windowFunctions`, `clearValidTo`, `returning`, `claims`, `graphAnalytics`, `resourceAudit`'s `resource` / `identityLeaseResource`, and so on — stays freely derivable. ## What you cannot override Every other field is refused for one of these reasons: most are captured by more than the profile's head alone, so overriding only the head would leave `buildOperations`, `lateMembers`, or a member group they build reading the value the base builder closed over; `dialect` and `assembly` are refused for different reasons of their own (see the table). | Field | Why it's refused | | --- | --- | | `dialect` | The operation backend literal hardcodes it. | | `tableNames` | Captured by `buildOperations` and every transaction handle. | | `execution` | Captured by `buildOperations` and every transaction handle. | | `strategy` | Captured by `buildOperations` and every transaction handle. | | `fulltext` | Captured by `buildOperations` and every transaction handle. | | `vector` | Captured by `buildOperations` and every transaction handle. | | `provisioning` | `ensureTable`, `catalog`, and `lineage` are all captured by migrations and transaction handles. | | `assembly` | Opaque and bundled-only; a derived profile carries the base's `assembly` forward by reference, so it resolves to the identical `buildOperations` / `lateMembers` pair the base builder closed over. | An override naming any of these throws `ConfigurationError` with code `ENGINE_PROFILE_OVERRIDE_UNSUPPORTED`, naming the key, whether or not the type would have allowed it — the check runs against the overrides object's own keys at runtime, not only its declared type. These refusals are `deriveEngineProfile`'s contract, not `createSqlBackend`'s. A profile spread by hand (`{ ...base, execution: mine }`) carries the base's `assembly` by reference, so `createSqlBackend` accepts it and applies the override to some members while others keep the builder's value — exactly the split the refusal exists to prevent. Derive through `deriveEngineProfile`. ## Worked example: a custom advisory-lock spelling An engine that spells its advisory lock differently from the bundled `pg_advisory_xact_lock(hashtext($namespace), hashtext($key))` form — hashing one concatenated string instead of two separate arguments — derives a `FenceSql` and passes it as an override: ```typescript import { buildPostgresEngineProfile, createSqlBackend, deriveEngineProfile, } from "@nicia-ai/typegraph/adapters/drizzle/engine"; import { postgresFenceSql } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import type { FenceSql } from "@nicia-ai/typegraph/backend"; import { sql, type SqlFragment } from "@nicia-ai/typegraph"; function customAdvisoryLockExpression( namespace: string, key: string | number, ): SqlFragment { const keyText = typeof key === "number" ? String(key) : key; return sql`pg_advisory_xact_lock(hashtext(${namespace} || ':' || ${keyText}))`; } const customFenceSql: FenceSql = { advisoryLockExpression: customAdvisoryLockExpression, lockTables: postgresFenceSql.lockTables, isolationFactExpression: postgresFenceSql.isolationFactExpression, }; const baseProfile = buildPostgresEngineProfile(db, options); const derivedProfile = deriveEngineProfile(baseProfile, { fenceSql: customFenceSql, }); const backend = createSqlBackend(derivedProfile); ``` `advisoryLockExpression` is the custom spelling here; `lockTables` and `isolationFactExpression` are the bundled PostgreSQL builders, reused because this example leaves them unchanged — a custom `FenceSql` need not replace every member. TypeGraph derives the standalone-statement forms every lock site actually calls (`acquireKeyed`, `acquireKeyedWithIsolation`, `isolationFact`) from these two expressions, so `customFenceSql` never spells a statement and its expression separately — the two cannot disagree about what they lock or read. This is the same `customAdvisoryLockExpression` pinned by `tests/engine-profile-derivation.test.ts` against a real PostgreSQL connection, trimmed of the `customLockTables` / `customIsolationFactExpression` coverage this example doesn't need. Every write-fence lock site now spells its lock through `customFenceSql` instead of the bundled one — including the recorded graph-write fence, which fuses its lock into its own CTE (`buildLockSchemaVersionAndGraphWrite`) but reads `advisoryLockExpression` / `isolationFactExpression` off the resolved fence target rather than a hardcoded bundled spelling, so this derivation reaches it too. The ONE exception, not reachable through `fenceSql`, is the schema-commit fence (`acquireSchemaWriteFence` in `postgres.ts`): it emits a standalone, single-argument `pg_advisory_xact_lock` call through `advisoryLockSingleExpression`, baked directly into `buildPostgresEngineProfile`'s closure. It deliberately occupies a different lock space from every two-argument lock `fenceSql` spells, so it is not an oversight `fenceSql` could close even if it were derivable — reaching it needs a from-scratch profile (see [What is not derivable yet](#what-is-not-derivable-yet)). The graph-template instantiation statement is a different, already-reachable case: it is the `instantiateStatement` member of `graphTemplateRuntime`, one of the fields this same derivation can override (see the table above). ## Worked example: a portable `row`-mechanism fence An engine with no advisory-lock primitive at all — a PostgreSQL-wire engine with no working `pg_advisory_xact_lock` — declares `mechanism: "row"` instead. TypeGraph spells the keyed acquisition itself against the never-dropped fences relation, so this derivation needs no `advisoryLockExpression` at all — only the declared mechanism and its two facts, `drain` and `conflict`: ```typescript const derivedProfile = deriveEngineProfile(baseProfile, { declaredCapabilities: { ...baseProfile.declaredCapabilities, writeFence: { mechanism: "row", drain: "quiescent", conflict: "commit-time", }, }, }); const backend = createSqlBackend(derivedProfile); ``` `conflict` states which of the two ways this engine resolves two writers of one fence row: `"wait"` for a lock-based engine (the second acquirer's statement blocks, exactly like `"advisory"`); `"commit-time"` for an optimistic-concurrency engine, where both acquirers proceed and the loser's COMMIT fails. Declaring `"commit-time"` on an interactive backend (as here) derives `capabilities.execution.unitOfWork: "optimistic-retry"` — every store-owned write this backend opens now replays a real commit-time conflict as a whole unit, up to `OPTIMISTIC_RETRY_ATTEMPTS` (3) attempts, rather than surfacing the raw driver error on the first one. `drain: "quiescent"` is the simplest legal drain when nothing else needs a real table lock; pass `fenceSql.lockTables` and declare `drain: "table-lock"` instead when this engine has one. A `fenceSql.isolationFactExpression`, if this engine's wire protocol supports reading it, rides the SAME acquisition statement — pass `postgresFenceSql.isolationFactExpression` (or a custom one) as `fenceSql` to keep recorded capture and match-key convergence trusting a real fact instead of failing closed on an unknown one. ### Declaring a `serializationFailure` classifier `isSerializationFailure` (the one predicate every retry owner consults) recognizes PostgreSQL's own `40001` / `40P01` SQLSTATEs and their fixed driver-message fallback. An engine whose commit-conflict shape is something else entirely — a custom error class, a different code — declares `execution.serializationFailure` so the SAME predicate recognizes it instead of falling through to a raw, unretried failure: ```typescript const profile = buildPostgresEngineProfile(db, options); const backend = createSqlBackend({ ...profile, execution: { ...profile.execution, serializationFailure: (error) => error instanceof Error && error.message.includes("CONFLICT_ON_COMMIT"), }, }); ``` This is deliberately NOT a `deriveEngineProfile` override: `execution` is captured by `buildOperations` and every transaction handle (see [What you cannot override](#what-you-cannot-override) below), so `deriveEngineProfile` refuses it like every other field in that table. Hand-spreading `execution` this way is safe for reading `serializationFailure` itself, because `createSqlBackend` is the only reader of `profile.execution.serializationFailure` — it registers the classifier against the exact backend object it is about to return, once, at construction — while every other `execution` member (`compile`, `execute`, `runExclusive`, and so on) rides forward as the SAME function reference the base builder closed over, spread unchanged. `createSqlBackend` consults the registered classifier for `isSerializationFailure` calls that pass this backend (or one of its transactions) as `target`; every store-owned unit routed through `runRetriedUnit`, and `store.transaction`'s own retry, already does. That safety is narrow, and it does not extend to the profile object itself. `{...profile, execution: {...}}` is a plain object literal — a different object from the one `buildPostgresEngineProfile` returned — so `isFirstPartyProfile` no longer recognizes it. `createSqlBackend` gates every `markFirstPartyFactory` call on that check, so a hand-spread profile loses standing to two optimizations, silently and with no functional difference to catch in testing: the dialect-derivation write-fence fallback (moot here, since the spread profile still carries `writeFence` declared) and the lazy schema-fence lease (`withTransactionSchemaFenceLease`, `src/store/operations/write-transaction.ts`), which falls back to the conservative per-call fence instead. Accept that trade for a one-off `serializationFailure` override; a backend meant to keep first-party standing declares `serializationFailure` inside the builder function that constructs `profile` in the first place, rather than spreading the finished object afterward. ## Removing `fenceSql` `fenceSql` is the one field a derived profile can clear: pass `fenceSql: undefined` to drop the bundled spelling entirely. That alone is not enough to reach a working profile — `createSqlBackend` still resolves a write-fence plan eagerly, and a profile whose resolved `writeFence.mechanism` is still `"advisory"` with no `fenceSql` refuses with `WRITE_FENCE_SQL_UNAVAILABLE`. Pair it with a `declaredCapabilities` override that stops claiming `"advisory"` (for example, declaring `writeFence: { mechanism: "engine-serialized" }` instead — no `drain` key: that field applies only to `mechanism: "advisory"`) to actually resolve an `engine-serialized` plan that needs no lock spelling at all. ## Supplying `lineage` `EngineProvisioning.lineage` forwards onto the assembled backend's optional `lineage` member unchanged, exactly like `provisioning.catalog` forwards onto `catalog`. Neither bundled profile sets it: `buildPostgresEngineProfile` and `buildSqliteEngineProfile` both leave it `undefined`, so a store built on a bundled backend derives its `lineage` from its own recorded relations when `history: true` is on, and has none otherwise (see [Lineage and pruned diffs](/graph-merge#lineage-and-pruned-diffs)). An engine whose storage layer already tracks a whole-database revision and can answer "what changed in this graph since revision R" more cheaply than a full scan supplies `lineage` directly. Both `revision` and `changesSince` take a **session** as their first argument. Run each read on that session (`session.execute` or `session.executeRaw`); a connection captured by the strategy may see a different snapshot. A caller planning outside a transaction passes the root backend. A backend that also needs lineage inside transactions must thread it through `EngineProvisioning.lineage` so transaction handles expose the same capability. `revision()` returns an opaque token comparable only with revisions from the same lineage source. `changesSince()` must report every changed node and edge key, including inserts, updates, deletes, and resurrection. Return `{ kind: "unbounded" }` when the delta cannot be bounded. TypeGraph uses lineage to prune branch diffs when it has a TypeGraph-owned revision anchor; for stores without revision tracking, `base@V` uses a complete content fingerprint regardless of engine lineage. That fingerprint covers current identity assertions and is recomputed inside the target commit transaction. Previously minted `engine:` base tokens are retired. Test a new `lineage` against `tests/backends/integration/lineage-conformance.ts`'s `registerLineageConformanceIntegrationTests` (registered per-dialect through `createIntegrationTestSuite`, or called directly against your own backend, via `{ getStore: () => ({ backend }) }`). It registers two describes: only "lineage: recorded-relations conformance" is portable — it drives every case through `resolveLineage`, the same path a real caller takes, and is the case the bundled recorded-relations derivation passes: after N writes, `changesSince(r0)` is exactly the touched keys, `changesSince(rN)` is empty, a hard delete after a revision reports the deleted key once, and an unrecognized revision is `unbounded`. "lineage: capture-completeness evidence" is TypeGraph-specific — it exercises `recordedRelationsLineage` directly (the per-revision evidence a bare engine revision has no equivalent gap for); an engine profile's own suite should run against the conformance describe only and skip the other. ## Supplying `recordedTime` `EngineProvisioning.recordedTime` declares an engine that tracks recorded (system) time itself, rather than through TypeGraph's own capture relations and clock — a backend that declares it must also declare `lineage` (engine-native history keeps no recorded relations for TypeGraph to derive a change delta from; `createSqlBackend` refuses `recordedTime` without a co-declared `lineage` with `ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE`). `EngineRecordedTimeMembers` has two members, `source` and `revisionNow`, both `this: void`. `source(table, revision)` names the table expression `table` (`"nodes"` | `"edges"` | `"identityAssertions"`) reads its recorded rows from, AS OF `revision` — the engine's own temporal-table syntax, with the interval already folded in (a system-time `AS OF` clause, or equivalent). It replaces what a TypeGraph-relation-backed source spells as two members: the recorded relation itself (`recordedNodesTable`/`recordedEdgesTable`) and a separate `recorded_from <= r AND r < recorded_to` interval predicate. Because `source`'s own expression already scopes every row to exactly one revision, there is nothing left for a predicate to narrow — every recorded read this member backs compiles with no interval clause at all. `revision` is an opaque `{ revision, recordedAt }` pair minted by your own `revisionNow` below; never parse `revision.revision` as a number; embed it in the AS OF expression as an opaque token. `table` is never called with `"identityAssertions"` today — a recorded identity read (`Store. identityAtCoordinate` at a past instant, and the query compiler's historical identity traversal) is refused outright under engine-native ownership before any read compiles, so your implementation must still accept the shared union without that arm ever running. `revisionNow(session)` is called on two different kinds of session, and must answer differently for each: - **On a root backend** (`store.recordedNow()`, `store.revisionNow()`): the engine's current COMMITTED revision. - **On an open `transaction()` handle** (both places `TransactionReceipt.recorded` is stamped, called before that transaction's own COMMIT): the revision at which THIS transaction's writes will become visible once it commits — the engine's pending/next revision for that session, not the last one committed before it opened. TypeGraph stamps this still-uncommitted value straight into the receipt it hands back to the caller once the transaction succeeds. An engine that can only name its last-COMMITTED revision, never its own pending one from inside an open transaction, cannot implement `recordedTime`: stamping the last-committed value into a receipt would describe the state *before* the write the receipt is reporting on, and there is no correct point after COMMIT to read the right value from without reopening the race `recordedTime` exists to close. Declaring `recordedTime` changes what `history: true` means on your profile. TypeGraph's own recorded relations, clock, and write-fence-gated clock allocation are never engaged; `revisionTracking: true` is refused whether or not `history: true` is also requested (`ENGINE_NATIVE_REVISION_TRACKING_UNSUPPORTED` — there is no TypeGraph clock for it to advance, and the engine's own revision is only ever available under `history: true`); a `recordedRead` external binding is refused (`ENGINE_NATIVE_RECORDED_READ_UNSUPPORTED` — there is no TypeGraph recorded relation for one to populate); and `migrateLegacyRecordedTime` refuses outright (`ENGINE_NATIVE_MIGRATE_RECORDED_TIME_UNSUPPORTED` — it rewrites TypeGraph's own recorded relations, which your backend does not have). `RecordedInstant` anchors from a `recordedTime`-declaring store use the `e1::` form rather than TypeGraph's `r1:<16-digit revision>:` form; `asOfRecorded` refuses an anchor minted under the other ownership form with `RECORDED_INSTANT_OWNERSHIP_MISMATCH`. See [Engine-native recorded time](/queries/temporal#engine-native-recorded-time) for the full reader contract. ## Refusals you may meet | Code | When | | --- | --- | | `ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION` | The profile's resolved capabilities omit `writeFence` — `createSqlBackend` has no write-fence decision to resolve and refuses outright, naming the one capabilities line to add. | | `WRITE_FENCE_SQL_UNAVAILABLE` | The resolved capabilities declare `mechanism: "advisory"` but the profile's `fenceSql` is missing the member that mechanism/drain combination needs; or `mechanism: "row"` with `drain: "table-lock"` but no `fenceSql.lockTables`. A `"row"` target missing `tableNames.fences` is NOT refused here — it refuses the first time a keyed site actually acquires the fence row. | | `WRITE_FENCE_DECLARATION_INVALID` | The declared `writeFence` carries an unrecognized `mechanism`, `drain`, or `conflict` string; a `drain` key on a mechanism other than `"advisory"` / `"row"`; a `conflict` key on anything but `"row"`; or `conflict: "commit-time"` on a target whose own `capabilities.execution.interactiveTransactions` is `false` — that value is honored only by the `"optimistic-retry"` execution tier, which never derives without an interactive transaction to replay inside, so accepting it there would silently drop it rather than apply it. `resolveWriteFencePlan` validates the raw value (a plain-JavaScript author is not held to the discriminated-union type) before shaping a plan from it. | | `CALLER_SERIALIZED_REFUSES_ADOPTION` | `adoptTransaction` was called on a backend whose resolved write-fence plan is `caller-serialized` — an externally owned transaction's lifetime cannot be held by the backend's in-process write-unit queue. | | `CATALOG_UNAVAILABLE` | A store path that needs the backend's catalog probes (index materialization, the recorded-time schema check, the recorded-time migration's column read) finds `catalog` absent — a profile whose `provisioning.catalog` is unset builds a backend with no `catalog` member at all. | | `LINEAGE_UNAVAILABLE` | A caller reached `requireLineage` and found `lineage` absent on the backend it asked. Every OUT-OF-TRANSACTION graph-merge caller consults `lineage` through `resolveLineage`, which already falls back to the recorded-relations lineage or to a full comparison rather than hitting this refusal. `assertTargetUnchanged`'s in-transaction re-validation reads the transaction handle's `lineage` ONLY — no fallback to the root — so this fires whenever a `lineage` that anchored the plan (found on the root at plan time) is not ALSO threaded onto the transaction handle that commits it; see "Supplying `lineage`" above for how to thread it correctly. | | `ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE` | The profile declares `recordedTime` without also declaring `lineage` — engine-native history keeps no recorded relations of its own for TypeGraph to derive a graph-merge change delta from, so the engine's own `lineage` is the only source for one. Raised at `createSqlBackend` construction, naming both members. | | `RECORDED_TIME_UNAVAILABLE` | A caller reached `requireRecordedTime` and found `recordedTime` absent on the backend it asked. Store construction under `history: true` and the shared `recordedNow()`/`revisionNow()`/receipt-stamping read are the only callers today, both reached only once `recordedTimeOwnership` has already resolved to `"engine-native"`, so this is defense-in-depth rather than a reachable misconfiguration on a bundled backend. | | `ENGINE_NATIVE_REVISION_TRACKING_UNSUPPORTED` | A store was constructed with `revisionTracking: true` against a backend that declares `recordedTime`, whether or not `history: true` was also requested — engine-native has no TypeGraph clock for `revisionTracking` to advance on its own. | | `ENGINE_NATIVE_RECORDED_READ_UNSUPPORTED` | A store was constructed with an external `recordedRead` binding against a backend that declares `recordedTime` — there is no TypeGraph recorded relation for one to populate; engine-native's own recorded reads are sourced from `recordedTime.source` instead. | | `ENGINE_NATIVE_RECORDED_IDENTITY_UNSUPPORTED` | `Store.identityAtCoordinate` at a past recorded instant, or the query compiler's historical identity traversal, was reached under engine-native recorded time — identity history reads TypeGraph's own recorded relations directly, which an engine-native backend does not populate. | | `ENGINE_NATIVE_MIGRATE_RECORDED_TIME_UNSUPPORTED` | `migrateLegacyRecordedTime` was called against a backend that declares `recordedTime` — the migration rewrites TypeGraph's own recorded relations, which an engine-native backend does not have. | | `RECORDED_INSTANT_OWNERSHIP_MISMATCH` | `asOfRecorded(instant)` was called with an instant minted under the OTHER recorded-time ownership form — an `r1:` instant against an engine-native store, or an `e1:` instant against a TypeGraph-owned one. | | `ENGINE_PROFILE_OVERRIDE_UNSUPPORTED` | `deriveEngineProfile`'s `overrides` names a key outside the derivable set, or one of the three adapter-backed sub-fields with a changed value (see [the carve-out](#the-adapter-backed-carve-out)). | | `ENGINE_ASSEMBLY_UNRECOGNIZED` | The profile's `assembly` is not a value `assembleEngine` produced — a profile built by hand rather than obtained from a bundled builder (optionally adapted with `deriveEngineProfile`). | ## What non-first-party costs First-party standing is bound to the exact profile object one of the two bundled builders returned, not to a field — a derived profile is a new object neither builder ever saw, so it never carries that standing forward, even when every field is copied from a first-party profile unchanged. That costs a derived profile's backend two things: - **No dialect-derivation fallback.** `resolveWriteFencePlan`'s fallback for a profile with no `writeFence` declared is sound only for the two bundled dialects, so it never applies to a derived profile regardless — irrelevant in practice as long as `declaredCapabilities` is kept, since both bundled declarations already carry a write-fence declaration explicitly. - **No lazy per-transaction schema-fence lease.** The lease `store/operations/write-transaction.ts` takes out under `isFirstPartyFactory` is closed to a derived profile's backend; each managed write takes its own fence instead. Every gate `createSqlBackend` runs — the write-fence-declaration refusal, the `mechanism: "advisory"` without `fenceSql` refusal, the schema-fenced-insert and autocommit marks — still applies to a derived profile exactly as it does to a bundled one. A bundled profile object is frozen once its builder returns it: mutating a field on that exact object throws, rather than silently drifting the profile away from what first-party standing was granted to. `deriveEngineProfile` is unaffected — it spreads `base`'s fields into a new object literal, which does not freeze. ## What is not derivable yet Building a profile from scratch — rather than deriving a variant of a bundled one — needs an execution adapter, an operation strategy, and an operation-backend assembly, none of which is exported today. `SqlEngineProfile.assembly` is opaque, and its only constructor, `assembleEngine`, is exported from no entrypoint: it is authoring a new engine, not deriving a variant of an existing profile, and waits on a future exported assembly constructor. Until then, derivation from a bundled builder — changing a lock spelling, a capability declaration, a resource-audit verdict, or a runtime dependency bag — is the supported way to adapt a profile. # TypeGraph vs. Neo4j, LadybugDB and pgGraph: Who Wins What import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; My first pass at benchmarking TypeGraph against real graph databases ran the seven LDBC "short read" queries (IS1–IS7) against Neo4j and LadybugDB. TypeGraph on SQLite won every one of them, which should have made me suspicious rather than happy: IS1–IS7 are point lookups and one-hop walks, and an in-process engine doing direct index seeks can't lose that race to anything that pays a network round trip. So I rebuilt the harness around the queries a graph database is _supposed_ to win (shortest paths, bounded neighborhood walks, complex multi-hop reads, and whole-graph algorithms) and ran 17 queries against five engines at two scales. The short version: - **TypeGraph on SQLite still wins every point read, usually by 10–100x.** - **The engines that keep an in-memory graph index win whole-graph algorithms by three to four orders of magnitude**, and the gap gets wider as the data grows. I'm publishing both results rather than only the flattering one. ## The setup All five engines run through one shared harness ([`packages/benchmarks/src/real/`](https://github.com/nicia-ai/typegraph/tree/bench/pggraph-comparison-v2/packages/benchmarks/src/real)): | Engine | Version | | ------------------------------ | ---------------------------------------------------------------------------- | | SQLite | 3.53.2 (via `better-sqlite3` 12.11.1) | | PostgreSQL (TypeGraph backend) | 18.1 (`pgvector/pgvector:pg18` image) | | Neo4j | `neo4j:2026.05.0` server image, `neo4j-driver` 6.2.0, GDS plugin | | LadybugDB | `@ladybugdb/core` 0.18.0 | | pgGraph | Evokoa pgGraph 0.1.8 (`ghcr.io/evokoa/pggraph:0.1.8`, bundles PostgreSQL 17) | pgGraph is the new entrant, and I think it's clever. Rather than being a separate database, it's a Postgres extension that builds a derived CSR (compressed sparse row) index over ordinary normalized tables and exposes traversal and pathfinding as SQL functions. Its point reads are tuned Postgres, and the CSR index only comes into play once a query traverses. The 17 queries are the original IS1–IS7; IC13 (shortest path) and IC14 (weighted shortest path); BFS3, a bounded neighborhood walk; three complex reads (IC2, IC8, IC9); and four graph-algorithm queries: GA_DEGREE, GA_WCC (weakly connected components), and GA_BFS / GA_SSSP (whole-component reachability and shortest-path depth from a seed). I held the comparison to two rules. Every result is checked with a value-level digest, per row, across every engine that runs the query, and every run below passed, since a fast wrong answer shouldn't count. And nothing is skipped silently: an engine without a comparable primitive for a query reports a typed `gap`, so a `gap` cell means the engine can't run that query in a comparable form, not that I didn't get around to it. Scales are **SF1** (9,892 persons, 361K directed `knows` edges) and **SF10** (65,645 persons, 3.88M `knows`, 21.9M comments). Each is one run on an EC2 host with shared vCPUs, so treat sub-millisecond cells and anything flagged noisy as order-of-magnitude. ## SF1 p50 latency in milliseconds unless noted. Fastest engine per row in **bold**. | Query | typegraph-sqlite | typegraph-postgres | neo4j | ladybugdb | pggraph | | -------------------- | ---------------: | -----------------: | -------: | --------: | -------: | | IS1 | **0.03** | 0.76 | 4.84 | 1.11 | 0.93 | | IS2 | **1.85** | 21.5 | 39.9 | 76.0 | 19.9 | | IS3 | **0.24** | 1.82 | 4.07 | 4.5 | 1.47 | | IS4 | **0.02** | 0.98 | 3.01 | 0.39 | 0.91 | | IS5 | **0.03** | 1.08 | 3.1 | 1.61 | 0.94 | | IS6 | **0.07** | 1.97 | 3.23 | 3.93 | 1.89 | | IS7 | **0.07** | 2.15 | 5.74 | 7.83 | 1.76 | | IC13 (shortest path) | 3.88 | 22.3 | 3.24 | 9.91 | **2.18** | | IC14 (weighted SP) | 5586 | **4860** | gap | gap | gap | | BFS3 | **220** | 1139 | 468 | 1539 | 357 | | IC2 | 51.8 | 523 | **36.2** | 89.1 | 347 | | IC8 | **3.06** | 17.3 | 3.73 | 24.2 | 8 | | IC9 | 2236 | 15069 | 3388 | **775** | 3498 | | GA_DEGREE | **0.03** | 1.09 | 2.97 | 1.8 | 0.72 | | GA_WCC | 7269 | 24237 | 18.7 | gap | **9.53** | | GA_BFS | 221 | 1828 | **24.5** | gap | 288 | | GA_SSSP | 220 | 1742 | **23.6** | gap | 284 | ### Point reads TypeGraph on SQLite takes every IS row, often by one to two orders of magnitude, because it's the only engine here that pays no network round trip and no per-call query planning overhead. GA_DEGREE and IC8 go the same way, because underneath the "algorithm" and "complex read" labels they're point lookups too. This is the case for an embedded graph. Most of what an application asks its graph all day looks like IS1–IS7 (fetch this person, their recent posts, who they know), and on IS1 Neo4j takes 4.84ms where SQLite takes 0.03ms. ### Graph algorithms GA_WCC, GA_BFS, GA_SSSP, and IC13 are what a graph engine's specialized index exists for, and here the CSR engines are in a different league: - **pgGraph wins GA_WCC (9.53ms) and IC13 (2.18ms)**, running union-find and shortest path directly over its CSR index. - **Neo4j with the GDS plugin wins GA_BFS (24.5ms) and GA_SSSP (23.6ms).** Without GDS, Neo4j answers these with Cypher path enumeration, which works but is slow. With GDS it projects the graph into memory once and runs the same kind of set-based traversal pgGraph does. - **TypeGraph is three to four orders of magnitude slower on GA_WCC** (7,269ms on SQLite, 24,237ms on Postgres, against pgGraph's 9.53ms). That last gap isn't a bug I can fix. TypeGraph's algorithms run as rounds of SQL, one window or aggregate query per round, however well indexed, while pgGraph and GDS hold the graph in an in-memory structure built for this access pattern, and query tuning won't make SQL iteration behave like a CSR traversal. ## SF10 The SF10 run happened **before** I added the GDS plugin to the Neo4j setup, so Neo4j's graph-algorithm rows show `gap` here even though Neo4j can run them. For Neo4j on those queries, the SF1 table above is the current picture. `s` = seconds; otherwise milliseconds. | Query | typegraph-sqlite | typegraph-postgres | neo4j | ladybugdb | pggraph | | -------------------- | ---------------: | -----------------: | ------: | --------: | -------: | | IS1 | **0.03** | 1.06 | 5.22 | 1.12 | 0.93 | | IS2 | **2.35** | 77.2 | 133 | 104 | 22.6 | | IS3 | **0.42** | 6.23 | 53.8 | 15.7 | 1.87 | | IS4 | **0.03** | 0.91 | 3.33 | 0.90 | 0.96 | | IS5 | **0.04** | 8.67 | 5.52 | 2.68 | 0.96 | | IS6 | **0.07** | 3.69 | 8.09 | 5.50 | 2.02 | | IS7 | **0.07** | 7.29 | 8.23 | 8.38 | 1.80 | | IC13 (shortest path) | 35.0 | 339 | 82.2 | 161 | **2.25** | | IC14 (weighted SP) | 74.6s | **57.3s** | gap | gap | gap | | BFS3 | **1.7s** | 7.2s | 3.9s | 9.3s | 2.2s | | IC2 | 141 | 724 | **138** | 296 | 843 | | IC8 | **5.24** | 220 | 92.1 | 47.6 | 15.3 | | IC9 | 11.9s | 76.7s | 15.5s | **3.1s** | 25.4s | | GA_DEGREE | **0.06** | 1.05 | 3.02 | 2.07 | 0.82 | | GA_WCC | 119.9s | 506.0s | gap | gap | **69.0** | | GA_BFS | 2.8s | 22.3s | gap | gap | **2.0s** | | GA_SSSP | 2.8s | 22.3s | gap | gap | **1.9s** | Point reads barely move at 10x the data, as you'd expect from an indexed seek. The more interesting number is how GA_WCC grows: | Engine | SF1 | SF10 | Growth | | ------------------ | ------: | -----: | -----: | | pggraph | 9.53ms | 69.0ms | ~7x | | typegraph-sqlite | 7269ms | 119.9s | ~16x | | typegraph-postgres | 24237ms | 506.0s | ~21x | pgGraph grows a bit less than linearly with the data, while TypeGraph's round-by-round algorithm grows faster than linearly because more edges also means more rounds, so the gap gets bigger as the data grows. If you need whole-graph analytics over millions of edges on a schedule, use a specialized engine for that job. ### SQLite beats Postgres, even inside TypeGraph Both TypeGraph backends run the same logical algorithms, and SQLite is 4–8x faster on the heavy ones at SF10 (GA_WCC 119.9s vs 506.0s, GA_BFS 2.8s vs 22.3s, IC9 11.9s vs 76.7s). My first guess was network round trips, but that doesn't hold up: GA_BFS already issues one `INSERT ... RETURNING` per round, Postgres is on loopback, and the ratio stays about 8x at both scales, whereas a fixed per-call cost would shrink as a share of a longer run. That points at a per-row cost in how Postgres executes these queries, which I haven't tracked down yet. The one exception is IC14, the weighted shortest path, where Postgres wins. Its set-based frontier expansion handles a large, unbounded Dijkstra better than SQLite's row-at-a-time version. ## IC14: the one only TypeGraph ran IC14 asks for the lowest-cost path between two people rather than the fewest hops. In this lineup, nothing else could run it in a comparable form: Neo4j's GDS has no stored `knows` weight to project, pgGraph's shortest path counts hops only, and LadybugDB's weighted shortest path isn't wired into the harness yet. TypeGraph answers it on both backends with `store.algorithms.weightedShortestPath`, at real LDBC scale (5.6s / 4.9s at SF1, 75s / 57s at SF10), with byte-identical results across the two. Those times aren't fast, and the benchmark makes them look worse than real use would, because it picks random pairs near the graph's diameter, which is the worst case for single-source Dijkstra. Real weighted-path questions tend to be between related, nearby entities, where the same algorithm stops early. ## Loading Loading is where TypeGraph still trails. It improved a lot between my first run and this one, because in the meantime 0.37 shipped a trusted initial import, and the benchmark loader now uses it. `importGraph` validates every row it writes: schema shape, edge endpoints, cardinality, conflicts. That's the right default, since most imports come from somewhere you don't fully trust. But when you're filling a brand-new database from an export you produced and already validated, every one of those checks is wasted work. `trustedImportGraph` and `trustedImportGraphStream` skip them. They bypass the normal write pipeline, drop secondary indexes, insert straight into the tables, then rebuild the indexes and refresh statistics, all in one transaction. Loading the same 200,000 nodes and 200,000 edges into a fresh SQLite database both ways: ```text run 1 — importGraph: 6783ms trustedImportGraphStream: 2455ms run 2 — importGraph: 7253ms trustedImportGraphStream: 2206ms run 3 — importGraph: 5048ms trustedImportGraphStream: 2116ms ``` Trusted import was 2.5–3x faster on every run. The streaming form takes a header, then node chunks, then edge chunks, so the loader here reads the LDBC CSVs in two bounded passes instead of holding a multi-million-row graph in memory. Skipping validation needs a narrow contract. The node and edge tables must be completely empty, and TypeGraph only checks stream order and kind names. Property shapes, endpoints, and duplicate-free IDs are on you. Features whose extra writes it would otherwise skip (history, uniqueness constraints, `searchable()` and `embedding()` fields) are rejected outright. Point it at a database with rows in it and it throws before touching anything: ```text TrustedImportError: Trusted import requires globally empty TypeGraph node and edge tables. details: { tables: ["typegraph_nodes", "typegraph_edges"], reason: "database_not_empty" } ``` For anything that isn't a one-time load of a fresh, dedicated database, use `importGraph` (untrusted data, conflicts, history, search fields) or collection `bulkInsert` (trusted data going into a database that isn't empty). Here's what it did to the benchmark: | Engine | SF1 load | Earlier run | SF10 load | | ------------------ | -------: | ----------: | --------: | | ladybugdb | 46s | 41.6s | 373s | | neo4j | 66s | 71.4s | 409s | | pggraph | 120s | — | 1,119s | | typegraph-sqlite | 164s | 587.1s | 2,199s | | typegraph-postgres | 344s | 669.3s | 3,278s | TypeGraph on SQLite went from 587s to 164s (~3.6x) and on Postgres from 669s to 344s (~1.9x). SQLite is now within about 1.4x of pgGraph, while Postgres is still about 3x slower than pgGraph at both scales, because pgGraph loads through batched `INSERT`s tuned for its own schema and TypeGraph's Postgres path still uses prepared statements instead of `COPY`. Switching to `COPY` is next, and unlike most of what this benchmark turned up, it would help every Postgres user. ## pgGraph as an accelerator pgGraph is the most interesting result here, less for the rows it wins than for how it works. It indexes tables that already look a lot like the ones TypeGraph's Postgres backend writes (normalized nodes and edges in Postgres, queryable over the same connection), which is why its point reads look like tuned Postgres while its traversal numbers look like a graph engine's. **None of this is built yet.** The pgGraph driver in this benchmark loads its own copy of the data into its own schema. It doesn't sit on top of a TypeGraph-managed database. But the numbers suggest the shape of a pairing: use TypeGraph for the point reads and everyday writes it already wins, and when profiling finds a real whole-graph workload (components, centrality, unweighted shortest path at scale), build a pgGraph index over the same tables and send just that query through it. I'm excited about that direction, because it keeps everything in one database without giving up anything on the queries applications run most. ## What the slow rows actually mean It's tempting to read every slow cell as a TypeGraph bug. Mostly they aren't: - **IC9 is a modeling artifact.** LDBC models a message's creator as an edge, so ranking a feed means fetching a sort key for every candidate first: 1.2M of them at SF1 to return the top 20. A real feed would put the owner on the item and index `(owner, created_at)`, which `defineNodeIndex` already supports, so the fix belongs in the schema rather than the engine. - **The whole-graph algorithms are architectural.** Tuning has helped (a delta-frontier rewrite of connected components nearly halved the Postgres GA_WCC time during this work), but it won't close a 1,000x gap. Pairing with something like pgGraph for those queries is the realistic answer. - **IC14 is benchmark-amplified**, as above. And the caveats: one run per scale on a shared-vCPU host, so several cells are noisy and should be read as order-of-magnitude. What I'm confident in regardless: every query passes value-level parity across every engine that runs it, and the direction of every finding is far larger than run-to-run noise. ## Try it - [The harness](https://github.com/nicia-ai/typegraph/tree/bench/pggraph-comparison-v2/packages/benchmarks/src/real): all five engine drivers and the EC2 runner - Full results and investigation notes: [`sf1-results.md`](https://github.com/nicia-ai/typegraph/blob/bench/pggraph-comparison-v2/packages/benchmarks/reports/sf1-results.md), [`sf10-results.md`](https://github.com/nicia-ai/typegraph/blob/bench/pggraph-comparison-v2/packages/benchmarks/reports/sf10-results.md) - [Trusted initial import](/interchange#trusted-initial-import): the full contract - [TypeGraph 0.35: Faster Almost Everywhere](/blog/typegraph-0-35-performance): the fixes an earlier run of this benchmark turned up # Bring Your Own Database import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; The first versions of TypeGraph depended on Drizzle and assumed the database underneath was either a `pg` pool or `better-sqlite3`. That held up until people started running it on Cloudflare D1, inside Durable Objects, over Neon's HTTP driver, in PGlite, and on engines that speak the Postgres wire protocol but handle locking very differently from Postgres. Each of those broke an assumption somewhere, usually as a SQL error from deep inside a query the engine couldn't run, and over three releases (0.38, 0.51, and 0.57) I've been removing those assumptions. This post covers all three: how Drizzle became optional, how a backend can now declare what it can't do before a query fails, and how an engine that locks differently can describe that to TypeGraph. ## Drizzle is an adapter now Before 0.38, every `Store` carried Drizzle's types whether your code touched them or not. A strict TypeScript project that only ever called `store.nodes.Person.create(...)` still had to resolve Drizzle's dialect declarations to typecheck, which is a lot of ORM to pull into your type checking just to create a node. Now the portable `Store` has no Drizzle in it. The managed factories own the connection and hand you a complete store: ```typescript import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local"; const store = await createLocalSqliteStore(graph, { path: "./graph.db" }); const alice = await store.nodes.Person.create({ name: "Alice" }); ``` ```text created: LBPSjEoGPqI0P3C6kaJ5M Alice capabilities.execution.interactiveTransactions: true ``` `createLocalPgliteStore` does the same for Postgres-in-WASM. The full graph API is there, `store.transaction(...)` included. When you do want to write your own tables on the same connection, you opt in with `createAdapterStore`, and `tx.sql` hands you the native transaction: ```typescript await store.transaction(async (tx) => { await tx.nodes.Document.update(documentId, props); if (tx.sqlAvailability !== "available") { throw new Error(`Native transaction unavailable: ${tx.sqlAvailability}`); } await tx.sql.insert(documentVersions).values(versionRow); }); ``` `sqlAvailability` is a discriminant because raw SQL is sometimes off on purpose: on a history-enabled store, writing around TypeGraph would skip recorded-time capture, so `tx.sql` isn't available there. ## Drizzle is now an optional dependency In 0.51 `drizzle-orm` became an optional peer dependency, and ten entrypoints don't need it installed at all: the root, `backend`, `core`, `schema`, `indexes`, `graph-extension`, `interchange`, `profiler`, `graph-merge`, and `provenance`. That list is enforced by tests: a fixture imports all ten with `drizzle-orm` missing from `node_modules`, and source and build-output checks fail if an import ever drags it back in. The last three routes to Drizzle (recorded-time migration DDL, a claim comparison, and some removal-statement builders) moved to portable code with golden tests pinning byte-identical SQL on both dialects. If you use a managed store or a `/adapters/drizzle/...` entrypoint you still need Drizzle, and if your package manager skips optional peers you'll need to install it yourself. The managed factories tell you so with a typed error that includes the `npm install` command, instead of a module-resolution stack trace. ## Backends say what they can't do Removing the dependency was the easier part. The harder problem is that TypeGraph emits SQL some engines can't run, and it used to find that out in production, halfway through a query. Recursive CTEs are the clearest case, since variable-length traversals, subgraph extraction, and a few identity reads depend on them. An engine without them can now say so: ```typescript const capabilities: Partial = { recursiveTraversal: { supported: false, reason: "engine has no WITH RECURSIVE / equivalent", }, }; ``` With that declared, the operations that need recursion throw a `ConfigurationError` naming the operation, instead of sending the engine SQL it can't parse. `weightedShortestPath` falls back to walking the path one hop at a time, which returns the same answer in more round trips. Leaving the capability out means it's supported, so every existing custom backend keeps working as it did. ## How your engine keeps writers apart Some writes (Operational Identity, and the recorded-time clock behind `history` and `revisionTracking`) need exactly one writer per graph at a time. TypeGraph used to pick the lock by checking which dialect it was talking to, which went badly if your engine said "postgres" but had no advisory locks. Now the backend declares how it keeps writers apart, which for the two bundled engines is one line each: ```typescript // PostgreSQL writeFence: { mechanism: "advisory", drain: "table-lock" } // SQLite writeFence: { mechanism: "engine-serialized" } ``` A custom backend that hosts identity or recorded history without declaring one is rejected at construction, and the error message prints the line to add for your dialect. 0.57 added two more mechanisms. `caller-serialized` is your promise that nothing else writes to the database; TypeGraph enforces the in-process half and trusts you with the rest. `row` is for engines with no advisory locks at all: TypeGraph takes the lock by upserting a row in its own `typegraph_fences` table. If your engine settles write conflicts at commit time rather than blocking, declare `conflict: "commit-time"` and TypeGraph retries its own transactions as a whole unit when they lose, up to three attempts. ## Engine profiles The piece I'm happiest about is that 0.57 made the two bundled backends _data_. `createPostgresBackend` and `createSqliteBackend` are now the same function, `createSqlBackend`, applied to an engine profile: the dialect, how it executes, how it provisions, and what it declares. That means you can take a bundled profile, change what's different about your engine, and get a real backend: ```ts import { buildSqliteEngineProfile, createSqlBackend, deriveEngineProfile, } from "@nicia-ai/typegraph/adapters/drizzle/engine"; import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; const { db } = createLocalSqliteBackend(); const base = buildSqliteEngineProfile(db); const derived = deriveEngineProfile(base, { declaredCapabilities: { ...base.declaredCapabilities, writeFence: { mechanism: "row", drain: "quiescent", conflict: "commit-time", }, }, }); const backend = createSqlBackend(derived); console.log(backend.capabilities.execution?.unitOfWork); // "optimistic-retry" ``` (It's SQLite here only because that's what runs on a laptop.) You never set `unitOfWork` yourself; TypeGraph works it out from what the engine declared rather than guessing from the engine's name. Derivation is narrow on purpose. You can override declared capabilities, the lock SQL, and a handful of runtime hooks, but not `dialect` or `execution`, because the bundled builders capture those in more than one place, and I'd rather throw at construction than hand you a backend that's half one engine and half another. ## One error for "you lost the race" A Postgres serialization failure or deadlock used to surface as whatever the driver felt like throwing. Now every conflict comes back as `TransactionConflictError`, with the driver error as `cause`, and `store.transaction()` can retry for you: ```ts // SQLite never raises 40001; this stands in for what PostgreSQL would. function serializationFailure(): Error { return Object.assign(new Error("could not serialize access"), { code: "40001", }); } let attempts = 0; await store.transaction( async (tx) => { attempts += 1; await tx.nodes.Account.create({ owner: "ada", balance: 100 }); if (attempts < 3) throw serializationFailure(); }, { retry: { attempts: 3 } }, ); console.log(attempts, (await store.nodes.Account.find()).length); // 3 1 — two rolled-back attempts left nothing behind ``` The callback reruns from the top, so it has to be safe to run more than once: read inside it, don't reuse values from outside it, and don't cause side effects that escape the transaction. ## Branches your host can copy `branch()` normally copies a graph by streaming it into a fresh database. 0.57 also adds `forkedWorkingCopyStrategy`, which hands the copy to whatever your host is good at, such as a file copy, `CREATE DATABASE ... TEMPLATE`, or a provider's branch API. Because the fork is a physical copy of the database, it keeps things a streamed copy can't, including recorded history, so a forked branch can answer `asOfRecorded` queries from before the fork was taken. ## What isn't there yet You can't build an engine from scratch yet. `deriveEngineProfile` adapts one of the two bundled profiles, and building a profile from nothing needs pieces that aren't exported. No third engine has been run through this in production yet either: the commit-time retry path is tested by injecting conflicts into real SQLite and PGlite transactions rather than against an engine that produces them on its own. If you have one, I'd like to hear how it goes. ## Upgrading From 0.37, the Drizzle-specific entrypoints moved under `/adapters/drizzle` and the old paths are gone rather than aliased: | 0.37 | 0.38 and later | | ------------------ | ----------------------------------- | | `/sqlite` | `/adapters/drizzle/sqlite` | | `/sqlite/local` | `/adapters/drizzle/sqlite/local` | | `/sqlite/libsql` | `/adapters/drizzle/sqlite/libsql` | | `/postgres` | `/adapters/drizzle/postgres` | | `/postgres/pglite` | `/adapters/drizzle/postgres/pglite` | `/sqlite/local` and `/postgres/pglite` now mean the managed store factories. If you read `tx.sql`, switch to `createAdapterStore` / `createAdapterStoreWithSchema`; if you don't, nothing changes. Custom backend authors: `capabilities.pessimisticLocks` is now `capabilities.writeFence`, and conflicts should be matched as `TransactionConflictError` rather than by SQLSTATE. The [0.57.0 changelog](/changelog#0570) has the full list. ## Try it - [Backend Setup](/backend-setup#drizzle-free-entrypoints): entrypoints, capabilities, and the parity matrix - [Write fence declaration](/backend-setup#write-fence-declaration-writefence): every mechanism and what each error means - [Authoring an engine profile](/backend-authoring): the derivable fields and worked examples - [GitHub](https://github.com/nicia-ai/typegraph) # Merges That Wait import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; [Graph Merge](/blog/graph-merge) plans before it applies: you fork a working copy with `branch()`, stage changes on it, and `planMerge()` shows you the exact write set before anything touches the target. That works well as long as staging, planning and applying all happen in one run of one process. The workflows that most want a review step don't look like that. Say a nightly job proposes loyalty-point adjustments that a person approves the next morning, a deploy lands in between, and the thing feeding the branch is a queue that occasionally delivers the same message twice. An in-memory branch handle and a plan that goes stale as soon as anything is written can't cope with any of that. It took three releases to fix, and I think the result is one of the more unusual things TypeGraph can do. 0.56 made the review itself durable by storing it as graph data, 0.67 made the working copy durable so a branch can be closed in one process and reopened in another, and 0.68 made it safe to feed that branch from an at-least-once queue. ## The review lives in the graph A merge plan is tied to the target's revision, so if the target changes after planning, applying the plan fails. That rule is what keeps a stale plan from clobbering newer data, but it gets in the way if you want the review (the proposal, the decision, who approved it) stored as graph data next to the record it concerns, because _writing the review down_ is itself a write to the target, and by the time anyone approves the plan it's stale. 0.56 resolves this by separating what was reviewed from the plan you eventually apply. `planCandidateWriteSetReview()` captures the candidate changes, the plan, the policy, and a baseline of the target as one immutable, content-digested artifact: ```typescript const review = unwrap( await planCandidateWriteSetReview({ target: store, makeBackend, policy: { id: "manual-acceptance-v1", context: { requiredApprovals: 1, authorizedReviewers: ["reviewer:maya"] }, }, writeSet: { formatVersion: 1, sourceId: "catalog-review", target: await captureCandidateWriteSetTarget(store), nodes: [ { kind: "Item", id: proposal.id, properties: { label: proposal.label, status: "accepted" }, }, ], edges: [], }, }), ); ``` You store it as ordinary graph data, keyed by its own digest, along with the reviewer's decision. `Artifact`, `Decision` and `evidence` here are ordinary kinds the example defines itself, so this needs no special schema support: ```typescript const artifact = await store.nodes.Artifact.create( { content: JSON.stringify(review) }, { id: review.digest.value }, ); await store.edges.evidence.create(proposal, artifact, { note: "review" }); const decision = await store.nodes.Decision.create({ approved: true, reviewDigest: review.digest.value, reviewer: "reviewer:maya", }); await store.edges.evidence.create(decision, artifact, { note: "approval" }); ``` Applying the original plan at this point fails: ```typescript const stale = await applyMergePlan(store, review.plan); // stale.error is a StaleMergePlanError ``` Recording the review and the approval moved the target, so the plan is stale. This is the part I like: durable review doesn't get an exemption from the staleness check, because if recording an approval could quietly un-stale a plan, the check would mean nothing. Instead of reusing the plan, you call `revalidateCandidateWriteSetReview()`, which reads the stored artifact, re-plans the retained candidate against the current target, and tells you what it found: | `checked.status` | What it means | | ---------------- | ------------------------------------------------------------------------------------------------------------------ | | `compatible` | Nothing that matters moved. You get a fresh `plan` and the original `reviewDigest`. | | `changed` | Policy, options, baseline entities, or plan fields differ from what was reviewed. Get a new review and approval. | | `incompatible` | Graph id, schema identity, or revision origin don't match. This approval can't be used against this target at all. | In the example only the review and approval records were added, so the result is `compatible`. I was careful to keep `compatible` meaning only that nothing relevant moved; it says nothing about whether the caller is allowed to act, so the example checks both before spending the plan: ```typescript if (checked.status !== "compatible") { throw new Error("A new review and approval are required"); } if (checked.reviewDigest.value !== approval.reviewDigest) { throw new Error("Approval does not identify the validated review"); } const report = unwrap(await applyMergePlan(store, checked.plan)); ``` Authenticating the stored decision and enforcing `authorizedReviewers` are still your job. TypeGraph tells you whether the plan is safe to apply and leaves the question of who may apply it to you. ## A working copy that outlives its process The branch itself was still tied to one process, because a `branch()` result holds its store and close handle in memory and disposing it deletes the fork. 0.67 adds `branchDurable()`, which forks a working copy that persists and hands back a small JSON descriptor you can put on a queue. Here's the loyalty ledger's nightly job: ```typescript const created = unwrap(await branchDurable(base, host)); const staged = created.branch.store; await staged.nodes.Account.update(ada.id, { points: 160 }); await staged.nodes.Account.create({ name: "Grace", points: 40 }); // Releases this process's connection and writer lease. The working copy stays. await created.branch.close(); await queue.put(JSON.stringify(created.descriptor)); ``` The descriptor is the only thing that leaves the process: ```json { "kind": "sqlite-file-host", "version": 1, "graphId": "loyalty", "definitionHash": "ec9dd68d2fbf14e3", "branchId": "gluQHM58QmB1LFjZIJEdA", "base": "ec9dd68d2fbf14e3#s1\u0000revision:kHeqC6EE55ENVL3a_np2R:r1:0000000000000001:2026-09-20T19:16:21.492Z", "store": { "id": "35da40a3-1820-4b9f-b9b4-2117e653ded8" }, "schemaAnchor": { "version": 1, "hash": "ec9dd68d2fbf14e3" } } ``` TypeGraph doesn't ship a durable host. It defines the contract (a `DurableWorkingCopyStrategy` with `create`, `seal`, `reopen`, `abort` and `destroy`), and your host decides where a working copy lives, whether that's a directory, a database, or a provider's branch API. The descriptor's `store` field is the host's opaque locator, and everything else in it belongs to TypeGraph. The examples here ran against a small file-backed SQLite host, across separate `node` processes. The next day a different process reads the descriptor and reopens the branch, and what comes back is an ordinary `GraphBranch`, so everything after that is the merge API you already know: ```typescript const descriptor = JSON.parse(await queue.get()); const branch = unwrap(await reopenDurableBranch(graph, descriptor, host)); const plan = unwrap(await planMerge(base, [branch])); const report = unwrap( await applyDurableMergePlan({ target: base, branch, descriptor, strategy: host, plan, }), ); await branch.close(); unwrap(await destroyDurableBranch(descriptor, host)); ``` ```text plan: 2 node upserts, 0 conflicts merged: {"nodes":2,"edges":0,"identity":{"asserted":0,"retracted":0}} base now: Grace=40, Ada=160 reopen after destroy refused: true ``` `applyDurableMergePlan()` applies the approved plan through the target Store transaction. The former optional host-native merge hook was retired because a database merge that commits internally can cross the transaction boundary that checked the target revision. A host may still use native database branches for its durable working copies. A descriptor is a document your application stored and handed back later, so TypeGraph treats it as untrusted input. At fork time the host seals the true origin, and every reopen and destroy is checked against it, so relabeling the branch id makes the reopen fail: ```text Durable branch descriptor does not match the working copy the host attested for its store locator: the descriptor's TypeGraph fences disagree with the origin recorded at fork. This is a tampered, relabeled, or wrong-branch descriptor. ``` The same check stops you from destroying branch B with branch A's descriptor, or reopening with a graph that reuses the id `"loyalty"` but defines `Account` differently. ## The message that arrives twice Suppose the branch is fed from a queue where "award Ada 25 points" can arrive twice. The award has to apply exactly once, and whoever is downstream (a notification, an audit log) has to hear about every applied award eventually, even if the worker dies right after committing. What that calls for is a transactional outbox scoped to the working copy, and 0.68 builds one into the durable-branch contract. You hand `operateDurableBranch()` an idempotency key, a `mutation` describing the change, and `metadata` to keep as evidence: ```typescript const outcome = unwrap( await operateDurableBranch(descriptor, host, { idempotencyKey: "award-7731", metadata: { source: "orders-queue", messageId: 7731 }, mutation: { op: "award", account: adaId, points: 25 }, }), ); ``` TypeGraph never interprets `mutation`; it digests it together with `metadata`, hands the host the request, and validates what comes back. The host applies the change and writes an evidence row in one database transaction, which in this host is a single SQLite transaction on one connection. To check the rollback, I made it throw after the graph write and before the evidence insert: ```text DurableOperationError | GRAPH_MERGE_OPERATION | Durable operation failed: injected failure after the graph write evidence: undefined accounts: Grace=40, Ada=160 ``` There's no evidence row, and Ada is still at the staged 160. In the real run, the worker commits the award (Ada goes from 160 to 185) and then dies before telling anyone. A fresh process picks up the descriptor, and the queue redelivers the same message: ```text recover pid 76873 | Ada on branch: 185 redelivered: replayed | Ada on branch: 185 ``` `replayed` returns the evidence from the first run and applies nothing, so Ada stays at 185 instead of 210. If you reuse the key with a different payload, even just different metadata, the host rejects it without writing: ```text changed payload: DurableOperationConflictError | GRAPH_MERGE_OPERATION_CONFLICT Ada on branch: 185 ``` The evidence rows act as the outbox. Each one starts with `delivered: false`, and you can't destroy the branch while any are undelivered: ```text destroy: DurableEvidenceUndeliveredError | GRAPH_MERGE_OPERATION_UNDELIVERED | refusing to destroy "60367e1c-…": undelivered operation evidence remains ``` To deliver them, you scan for undelivered rows and mark each one after publishing it: ```typescript const page = unwrap( await scanDurableOperations(descriptor, host, { limit: 100 }), ); for (const evidence of page.operations) { if (evidence.delivered) continue; await publishDownstream(evidence); // your outbox consumer unwrap( await markDurableOperationDelivered( descriptor, host, evidence.idempotencyKey, ), ); } unwrap(await destroyDurableBranch(descriptor, host)); ``` A crash at any point in that sequence means, at worst, that a downstream consumer hears about an award twice; the award itself is never applied twice or silently dropped. ## Limits - **There's no first-party host and no queue.** Whether your host's transaction is atomic is up to your database. TypeGraph validates what the host reports back but can't check your storage. - **Exclusion is the host's job.** The example host's writer lease is a lock file. When I left one behind, as a killed process would, reopening failed until it was cleared. A real host wants a lease that expires. - **Plan against a quiet branch.** Feeding operations into a branch while a reviewer plans against it means planning against a moving target. - **Delivery is at-least-once.** Make downstream writes idempotent on the key. - **You probably don't need this for a one-shot job.** If you stage, plan and apply in one run, `branch()` is unchanged and simpler. ## Try it - [Durable candidate review](/graph-merge#durable-candidate-review-in-the-target-graph) - [Durable host-native branches](/graph-merge#durable-host-native-branches) and [atomic operations and evidence](/graph-merge#atomic-operations-and-immutable-evidence) - [Example 27](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/27-durable-merge-review.ts): the durable review, end to end - [GitHub](https://github.com/nicia-ai/typegraph) # The Cheapest Citation Lineage Isn't the Shortest One import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; When I introduced TypeGraph I said there was no PageRank and no community detection, and that if you needed those you wanted a real graph database. That's still partly true (more on that at the end), but as of 0.38 the list is a lot shorter. Until now, `store.algorithms` answered questions about two nodes at a time, like how to get from A to B or what's within three hops of A. Some questions need the whole graph at once: whether a dataset is one connected body or several islands, which nodes matter most structurally rather than by raw count, or whether communities emerge from the topology on their own. Answering those means running an algorithm round after round over the entire graph, from one consistent snapshot, until it converges. That needs more machinery than a single traversal, including a pinned transaction, a temporary working table, and a clear rule for what happens when the rounds don't settle. 0.37 added `weaklyConnectedComponents` and `weightedShortestPath`. 0.38 added global and personalized `pageRank` and deterministic `labelPropagation`, the same algorithm the LDBC Graphalytics benchmark uses for community detection. All five run as SQL against the store, so you don't export the graph to a separate analytics engine and keep a second copy in sync. They're also a lot of fun to play with, so the rest of this post runs them on a small citation graph. ## The corpus This is the same citation graph as the [research-copilot example](/examples/research-copilot): 18 landmark ML papers, 55 authors, 14 topics, and 37 real citation edges, from Rumelhart, Hinton & Williams' 1986 backprop paper through LLaMA in 2023. That example runs point queries; here I run the whole-graph algorithms over the same data. ```typescript const backend = createExampleBackend(); const [store] = await createStoreWithSchema(graph, backend); // Ingested 18 papers, 55 authors, 14 topics, 37 citation edges. ``` ## One body of work, or islands? Treat `cites` as undirected and partition by connectivity: ```typescript const components = await store.algorithms.weaklyConnectedComponents({ edges: ["cites"], nodeKinds: ["Paper"], }); ``` ```text 18 papers partition into 1 component(s): • component of 18 paper(s), rooted at "Adam: A Method for Stochastic Optimization" ``` Every paper reaches every other through some chain of citations. On a real, messy dataset this is the sanity check to run first. `nodeKinds: ["Paper"]` keeps authors and topics out of it, so more than one component would mean a separate sub-literature rather than a lightly cited author hanging off the edge. ## The cheapest lineage isn't the shortest one This is my favorite result in the post, and getting to it took one false start. Citations always point from a newer paper to an older one, so the obvious edge weight is `yearGap`, the number of years a citation reaches back. The trouble is that every hop steps strictly backward in time, so the gaps along _any_ route from A to B telescope to exactly `A.year - B.year`. Every path ties, and weighting by `yearGap` is just `shortestPath` with extra arithmetic. Squaring the gap breaks the tie. `yearGapCost = yearGap²` is convex, so one 27-year leap costs `27² = 729` while the same span covered in several small steps costs much less. The question becomes "what's the smoothest chain of ideas between these two papers," and that's where `weightedShortestPath` starts disagreeing with `shortestPath`: ```typescript const hopPath = await store.algorithms.shortestPath(from.id, to.id, { edges: ["cites"], maxHops: 10, }); const byConvexCost = await store.algorithms.weightedShortestPath( from.id, to.id, { edges: ["cites"], weightProperty: "yearGapCost" }, ); ``` ```text transformer → backprop: shortestPath (fewest hop): 2 hops transformer(2017) → adam(2014) → backprop(1986) weighted by yearGapCost: 4 hops totalWeight=353 transformer(2017) → dropout(2014) → alexnet(2012) → lenet(1998) → backprop(1986) ◀── more hops, lower convex cost clip → backprop: shortestPath (fewest hop): 3 hops clip(2021) → simclr(2020) → dropout(2014) → backprop(1986) weighted by yearGapCost: 7 hops totalWeight=359 clip(2021) → gpt2(2019) → bert(2018) → transformer(2017) → dropout(2014) → alexnet(2012) → lenet(1998) → backprop(1986) ◀── more hops, lower convex cost ``` The fewest-hops route from the Transformer paper to backprop makes a 28-year jump through Adam. The convex-cost route takes four smaller steps (Dropout, AlexNet, LeNet) and comes in at less than half the cost. From CLIP it walks seven hops, almost straight down the history of deep learning. Both are legitimate answers to "what's the best path" between the same two papers, and I like that the convex-cost one reads like a syllabus. Weights are checked before any traversal round runs, so a negative or non-numeric `yearGapCost` anywhere in the selected edges throws `InvalidEdgeWeightError` up front instead of producing a wrong answer. ## PageRank vs. counting citations A citation count tells you how many papers cite this one. PageRank tells you how much a paper matters given _who_ cites it: a citation from an important paper is worth more, and that carries through the graph. ```typescript const pageRankScores = await store.algorithms.pageRank({ edges: ["cites"], nodeKinds: ["Paper"], direction: "out", // random surfer follows citations forward, toward the classics }); ``` ```text PR-rank score cites raw-rank Δ title ──────────────────────────────────────────────────────────────── 1 0.24209 6 1 · Learning representations by back-propagating errors 2 0.09060 3 7 +5 ImageNet Classification with Deep Convolutional N... 3 0.07559 3 6 +3 Efficient Estimation of Word Representations in V... 4 0.07424 4 2 -2 Attention Is All You Need 5 0.05868 4 3 -2 BERT: Pre-training of Deep Bidirectional Transfor... 6 0.05827 1 13 +7 Gradient-Based Learning Applied to Document Recog... 7 0.05818 3 5 -2 Dropout: A Simple Way to Prevent Neural Networks ... 8 0.05592 3 4 -4 Deep Residual Learning for Image Recognition ``` Backprop wins both rankings, which is no surprise since it's the root of the whole corpus. The row I find interesting is 6th place. LeNet has exactly **one** citation in this corpus, which puts it 13th by count, but PageRank moves it up seven places because that one citation comes from AlexNet, which is heavily cited itself. A plain count would treat it like any other citation. ## What matters to CLIP, specifically Personalized PageRank runs the same iteration, but instead of jumping to a random node it keeps jumping back to a seed you choose. The question changes from "important globally" to "important from where CLIP is standing": ```typescript const personalized = await store.algorithms.personalizedPageRank({ edges: ["cites"], nodeKinds: ["Paper"], direction: "out", seeds: [{ id: clip.id, kind: "Paper" }], }); ``` ```text PPR-rank score global-rank Δ title ────────────────────────────────────────────────────────────────── 1 0.25884 17 +16 Learning Transferable Visual Models From Natural ... 2 0.12804 1 -1 Learning representations by back-propagating errors 3 0.09396 5 +2 BERT: Pre-training of Deep Bidirectional Transfor... 4 0.07889 4 · Attention Is All You Need 5 0.06382 3 -2 Efficient Estimation of Word Representations in V... 6 0.05500 9 +3 Language Models are Unsupervised Multitask Learners 7 0.05500 15 +8 A Simple Framework for Contrastive Learning of Vi... 8 0.05500 16 +8 An Image is Worth 16x16 Words: Transformers for I... ``` CLIP itself jumps from 17th to 1st, since every jump lands back on it, and SimCLR and ViT, both cited directly by CLIP and both well outside the global top 10, climb eight places each. Changing only the seed gives you a ranking for a different question, and it's the one I'd reach for in a "related work" or recommendation feature. ## Do research communities fall out? Label propagation finds communities by having every node adopt the most common label among its neighbors, round after round, over the undirected version of `cites`. Run it strictly first: ```typescript const converged = await store.algorithms.labelPropagation({ edges: ["cites"], nodeKinds: ["Paper"], onMaxIterations: "throw", // default }); ``` ```text onMaxIterations: "throw" raised GraphAlgorithmConvergenceError — the undirected citation graph oscillates (tree / even-cycle structure that mirrors labels back and forth). ``` That error is expected. In synchronous label propagation a node doesn't vote for itself, so a tree-shaped neighborhood (common once you flatten a citation DAG into an undirected graph) can flip two labelings back and forth forever, and more iterations won't fix it. The default throws rather than handing you whatever labels the last round happened to land on. If you want that fixed-round answer, which is what the LDBC Graphalytics benchmark specifies, ask for it: ```typescript const fixedRound = await store.algorithms.labelPropagation({ edges: ["cites"], nodeKinds: ["Paper"], onMaxIterations: "return", }); ``` ```text 3 communities: ── community of 7 ── Adam: A Method for Stochastic Optimization [Optimization] Learning representations by back-propagating errors [Optimization, DeepLearning] Dropout: A Simple Way to Prevent Neural Networks ... [DeepLearning, Optimization] Gradient-Based Learning Applied to Document Recog... [CNN, ComputerVision] Sequence to Sequence Learning with Neural Networks [RNN, NLP, DeepLearning] Very Deep Convolutional Networks for Large-Scale ... [CNN, ComputerVision] Efficient Estimation of Word Representations in V... [Embeddings, NLP] ── community of 7 ── BERT: Pre-training of Deep Bidirectional Transfor... [Transformer, NLP, SelfSupervised] Learning Transferable Visual Models From Natural ... [Contrastive, MultiModal, ComputerVision] Chain-of-Thought Prompting Elicits Reasoning in L... [LanguageModel, Reasoning, NLP] Language Models are Unsupervised Multitask Learners [Transformer, NLP, LanguageModel] LLaMA: Open and Efficient Foundation Language Models [Transformer, LanguageModel, NLP] Attention Is All You Need [Transformer, Attention, NLP] An Image is Worth 16x16 Words: Transformers for I... [Transformer, ComputerVision, DeepLearning] ── community of 4 ── ImageNet Classification with Deep Convolutional N... [CNN, ComputerVision, DeepLearning] Momentum Contrast for Unsupervised Visual Represe... [Contrastive, SelfSupervised, ComputerVision] Deep Residual Learning for Image Recognition [CNN, ComputerVision, DeepLearning] A Simple Framework for Contrastive Learning of Vi... [Contrastive, SelfSupervised, ComputerVision] ``` The algorithm only saw undirected `cites` edges (the topic tags are printed for you and were never fed to it), yet it separated the optimization and classic-vision foundations, the transformer and language-model era, and the contrastive self-supervised vision cluster purely from who cites whom. ## What's still missing Shortest path (weighted and unweighted), reachability, neighborhoods, degree, connected components, label propagation, and global and personalized PageRank cover a lot of ground, but strongly connected components, topological sort, betweenness/closeness/eigenvector centrality, and Louvain/Leiden community detection aren't in `store.algorithms`. For those, pull the edge list out with `.query().traverse()` or `store.subgraph()` and hand it to an in-memory library like [graphology](https://graphology.github.io/). Scale matters too. These run as rounds of SQL, which keeps everything in one database and is fine for graphs the size most applications have, but on millions of edges the engines that hold the graph in a specialized in-memory index are orders of magnitude faster at whole-graph work. The [benchmark post](/blog/benchmarking-typegraph-neo4j-ladybugdb) has the numbers, including the unflattering ones. ## Try it - [Graph Algorithms](/graph-algorithms): every algorithm, shared options, temporal behavior, and PageRank tolerance notes across backends - [Research Copilot](/examples/research-copilot): the same corpus, run through the point-query algorithms - [Example 32](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/32-graph-analytics.ts): the runnable source behind this post - [GitHub](https://github.com/nicia-ai/typegraph) # Merging Two Feeds That Disagree About the Same Patient import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; Point two ingestion agents at overlapping data (an EHR export and a claims feed, say) and tell them to "just write everything to the graph", and you end up with two patient nodes for one person, each holding half the care history and neither aware of the other. Most pipelines then add a nightly dedupe job and hope nothing reads the graph in between. I think append is the wrong default for graphs, which is why 0.31 ships `@nicia-ai/typegraph/graph-merge`. It's the feature I've been most eager to get into people's hands. You branch a store, let each writer work on its own copy, and then fold the branches back in: entities are resolved, edges are repointed onto the surviving nodes, disagreements are reported instead of silently overwritten, and the merge records who contributed what. ## Branch, write, merge `branch()` records the base store's current state and hands back an isolated working copy. Writers use the ordinary store API against it. `merge()` diffs every branch against the base and runs one pipeline to fold them back in: ```text stage (diff every branch) → generate candidates (exact unique · blocking key · similarity) → cluster (group nodes that are the same entity) → canonicalize (pick a survivor, union properties, resolve conflicts) → repoint + dedupe edges onto survivors → reconcile delete/modify and types → commit transactionally + build the report ``` The whole thing is deterministic: clusters resolve by stable keys and conflicts are decided by an explicit `branchOrder` rather than by which branch happened to arrive first, so merging the same branches in any order commits the same graph. ## Two ways to be the same patient [Example 18](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/18-fhir-graph-merge.ts) runs this on a small FHIR-flavored care graph. An EHR branch and a claims branch each record the same two patients, and each gets both identities wrong in a different way: - **Anna Rivera** (EHR) and **Ana Rivera** (claims) share `MRN-001`, and because `mrn` is declared `unique`, they're matched regardless of how the name is spelled, without any similarity threshold. - **Mohammed Ali** (EHR, `MRN-204`) and **Mohamed Ali** (claims, `MRN-205`) have _different_ MRNs, so nothing forces them together. They share a birth date, which puts them in the same blocking bucket, and they collapse because their fulltext name similarity clears the configured `0.78` threshold. Landing in the same bucket only means two records get compared. The test suite also covers the other side of the threshold with **Zoe Adams** and **Quinn Webb**, who share a birth date and land in the same bucket but score near zero on name similarity, so they stay two separate patients. ```typescript const mergeOptions: MergeOptions = { resolve: { Patient: { block: (node) => node.birthDate, similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.78, }, }, onPropertyConflict: "flag", branchOrder: [EHR_BRANCH, CLAIMS_BRANCH], provenance: true, }; const result = await merge(base, [ehr, claims], mergeOptions); ``` Both pairs collapse to one canonical patient each: ```text merged nodes: 9 merged edges: 10 entity resolutions: 2 ``` That's six nodes staged on the EHR branch plus five on claims, minus the two patients that folded into their counterparts, and all ten edges survive without duplicates. ## Disagreements get flagged, not hidden Merge doesn't quietly pick a value when branches disagree. The properties that don't agree are flagged, and the flags also show which identity mechanism was in play: ```text conflicts: - Patient.fhirId on patient-ana: claims-agent="Patient/claims-ana", ehr-agent="Patient/ehr-anna" - Patient.name on patient-ana: claims-agent="Ana Rivera", ehr-agent="Anna Rivera" - Patient.fhirId on patient-mohamed: claims-agent="Patient/claims-mohamed", ehr-agent="Patient/ehr-mohammed" - Patient.mrn on patient-mohamed: claims-agent="MRN-205", ehr-agent="MRN-204" - Patient.name on patient-mohamed: claims-agent="Mohamed Ali", ehr-agent="Mohammed Ali" ``` There's no `mrn` conflict for `patient-ana`, because Anna and Ana share `MRN-001` exactly. The Mohammed/Mohamed pair does flag `mrn`, since their merge was based on name similarity and never required the MRNs to agree. ## Every edge lands on the survivor This is the part I like best: everything attached to either duplicate ends up on the one canonical patient. Here it is read back through each survivor's `forPatient` edges rather than a table scan: ```text Ana Rivera (MRN-001, 1974-03-09) - Encounter: Hypertension follow-up (2026-04-11T09:30:00-07:00) - Encounter: Kidney function review (2026-04-14T10:00:00-07:00) - MedicationRequest: Lisinopril 10 MG Oral Tablet - Take one tablet by mouth daily - Observation: Blood pressure panel = 152/96 mmHg (high) - Observation: Estimated glomerular filtration rate = 54 mL/min/1.73m2 (low) Mohamed Ali (MRN-205, 1990-08-21) - Encounter: Cardiology consult (2026-05-02T13:00:00-07:00) - Observation: LDL cholesterol = 168 mg/dL (high) ``` Ana Rivera's five resources came from both branches (the EHR encounter and medication, the claims encounter and lab result), and they now hang off a single patient node that neither branch created on its own, so anyone reading this record sees the full history. If an edge had been repointed wrongly, it would be missing from the list. ## Who contributed what With `provenance: true`, the merge reports which branch contributed each committed node and edge: ```text provenance: - ehr-agent: 6 node(s), 6 edge(s) - claims-agent: 5 node(s), 4 edge(s) ``` That report lives only as long as the call. Pass `persistProvenance: true` and each contribution is also written as a durable `{branch, sourceId} → canonical` row in a separate provenance graph on the same backend, so "what did this provider ever contribute?" is a query you can run next month, without adding anything to your domain schema. ## Merging into a graph that kept moving `merge()` is a snapshot operation: every branch must fork from the target's _current_ state, or it fails with `BaseVersionMismatchError` rather than risk clobbering newer data. That works when you fork, write, and merge in one round, but real ingestion keeps going, and [Example 19](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/19-incremental-merge.ts) covers folding new batches into a target that has already moved on. In that example a company knowledge base already has `Acme Corp` (`acme.com`), and a provider batch reports the same company under another spelling along with one company that's new: ```text Target before: [ 'Acme Corp (acme.com)' ] Target after: [ 'Acme Corp (acme.com)', 'Globex (globex.io)' ] No duplicate was created: the provider's "ACME Corporation" merged onto the committed "Acme Corp" via the shared domain. ``` `mergeIncremental()` finds the already-committed row by its unique `domain` and merges onto it instead of creating a duplicate. It's the same mechanism as the shared-MRN case, matched against live data instead of another branch: ```typescript const result = await mergeIncremental({ forkPoint, // the frozen ancestor the branch forked from target, // the live committed graph, which may have advanced branches: [provider], options: { resolve: { Company: { similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.9, }, }, onPropertyConflict: "flag", onBasePropertyConflict: "flag", // required: never let a stale branch value win persistProvenance: true, }, }); ``` `onBasePropertyConflict: "flag"` is required so that a stale branch can't overwrite something newer than the point it forked from. If the live target changed the same row after the fork, the target's value wins and the disagreement is reported. ## Try it - [Graph Merge](/graph-merge): entity resolution, blocking, similarity strategies, conflicts, scaling guards, and determinism - [FHIR Graph Merge](/examples/fhir-graph-merge) and [Incremental Graph Merge](/examples/incremental-merge): the docs walkthroughs - [Example 18](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/18-fhir-graph-merge.ts) and [Example 19](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/19-incremental-merge.ts): the runnable source behind this post - [GitHub](https://github.com/nicia-ai/typegraph) # Embeddings Can't Find a SKU import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; Type `PROD-1005-E` into a search box and you expect one result: the product with that SKU. Vector search is bad at this in a way tuning won't fix: a rare alphanumeric code barely moves an embedding, so cosine similarity ranks on everything _else_ in the query, and the product whose exact code you typed ends up under a pile of vaguely similar ones. BM25 handles it easily, because it counts terms and a token that appears in exactly one document wins however odd it looks. Rather than trying to make embeddings better at exact matches, I wanted both signals fused in one place, so 0.21 added native fulltext search and hybrid retrieval that combines it with vector search, and 0.24 finished the job on SQLite. ## Fulltext is a field modifier Mark a field `searchable()` and TypeGraph keeps a BM25 index in sync on every write. On Postgres that's `tsvector` + GIN; on SQLite it's FTS5. You don't need an Elasticsearch cluster for this, or a sync job to feed one. ```typescript const Product = defineNode("Product", { schema: z.object({ name: searchable({ language: "english" }), description: searchable({ language: "english" }), sku: searchable({ language: "english" }), category: z.enum(["outerwear", "footwear", "accessories", "climbing"]), embedding: embedding(16).optional(), }), }); ``` Query it with `store.search.fulltext()`, or use `$fulltext.matches()` as a predicate inside an ordinary query, where it combines with metadata filters and traversals in the same SQL statement: ```typescript const activeOuterwear = await store .query() .from("Product", "p") .whereNode("p", (p) => p.$fulltext .matches("lightweight", 10) .and(p.status.eq("active")) .and(p.category.eq("outerwear")), ) .select((ctx) => ({ sku: ctx.p.sku, name: ctx.p.name })) .execute(); ``` ## Hybrid, fused with RRF `store.search.hybrid()` runs the vector search and the fulltext search and merges them with Reciprocal Rank Fusion. RRF only looks at rank positions, so it doesn't care that a cosine score and a BM25 score live on completely different scales, which makes it the least fiddly fusion method I know of. ```typescript const hybridHits = await store.search.hybrid("Product", { limit: 5, vector: { fieldPath: "embedding", queryEmbedding, metric: "cosine", k: 20 }, fulltext: { query: "waterproof shell", k: 20, includeSnippets: true }, fusion: { method: "rrf", k: 60, weights: { vector: 1, fulltext: 1.25 } }, }); ``` Fulltext worked on both backends from 0.21, but hybrid needs the backend to run a vector search, which SQLite couldn't do yet, so on SQLite the hybrid call threw `ConfigurationError`. 0.24 gives SQLite a real vector search on top of `sqlite-vec`, built to the same shape as the Postgres version, and the hybrid call now takes the same options and returns the same results on both. SQLite also turned out to be the faster of the two. On the project's search benchmark (500 documents, 384 dimensions), SQLite hybrid runs in **0.8ms** against Postgres's 2.5ms, partly because SQLite runs in-process and doesn't pay for a round trip. ## Running it [Example 15](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/15-fulltext-hybrid-search.ts) seeds nine outdoor-gear products (parkas, shells, a climbing harness, ski goggles), each with a searchable name, description, and SKU, plus a small embedding. A BM25 query for `"waterproof jacket"` finds the Expedition Parka, and the snippet shows why: ```text Heavily insulated jacket for alpine expeditions and extreme cold. Waterproof outer shell with down fill. ``` Then the SKU: ```text Query: "PROD-1005-E" (looking up an exact SKU) 1. [PROD-1005-E] Climbing Harness Pro score=1.6328 ``` It comes back as the only result, where a pure vector search on that string ranks unrelated products higher because the SKU barely registers in the embedding. Hybrid is where it gets interesting. Here's `"waterproof shell"` with `k=20` on each side, RRF `k=60`, and fulltext weighted at `1.25`: ```text 1. [PROD-1001-A] Expedition Parka score=0.0357 (v#3, f#3) 2. [PROD-1002-B] Arctic Shell score=0.0356 (v#6, f#1) 3. [PROD-9901-Z] Legacy Rain Shell score=0.0355 (v#5, f#2) 4. [PROD-1007-G] Compression Socks score=0.0164 (v#1, f—) 5. [PROD-1008-H] Hiking Daypack 25L score=0.0161 (v#2, f—) ``` The `(v#, f#)` tags are each hit's rank on each side. Arctic Shell was fulltext's top pick but only sixth by vector, so pure vector search would have buried it, and RRF lifts it to second because doing well on _either_ side counts. Compression Socks went the other way: they were the vector side's top hit, but with no fulltext match at all they end up at the bottom of the list. ## It composes with everything else The same fusion is on the query builder as `.fuseWith()`, so a hybrid search can carry ordinary predicates. Filtering to `status = "active"` drops the discontinued Legacy Rain Shell in the same query, without a post-filter: ```typescript const builderHybrid = await store .query() .from("Product", "p") .whereNode("p", (p) => p.$fulltext .matches("waterproof shell", 20) .and(p.embedding.similarTo(queryEmbedding, 20)) .and(p.status.eq("active")), ) .fuseWith({ k: 60, weights: { vector: 1, fulltext: 1.25 } }) .select((ctx) => ({ sku: ctx.p.sku, name: ctx.p.name })) .limit(5) .execute(); // → Arctic Shell, Expedition Parka, Hiking Daypack 25L, Ski Goggles UV400, // Compression Socks ``` There are four query modes for what a search box actually receives: `websearch` for Google-style syntax (`"ski goggles" -compression`), `phrase` for exact adjacency, `plain` for all-terms-must-match, and `raw` when you want the engine's native `tsquery` or FTS5 `MATCH` syntax. And if the index falls behind (you added a `searchable()` field, or a bulk write went around the store), `store.search.rebuildFulltext()` backfills it page by page: ```text Before rebuild: "waterproof" → 0 hits (fulltext rows cleared). Rebuilt: kinds=Product processed=9 upserted=9 cleared=0 skipped=0 After rebuild: "waterproof" → 3 hits restored. ``` ## Try it - [Fulltext Search](/fulltext-search): RRF tuning, adding fulltext to existing data, and alternate Postgres strategies (pg_trgm, ParadeDB, pgroonga) - [Semantic Search](/semantic-search): the vector half - [Example 15](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/15-fulltext-hybrid-search.ts): the runnable source behind this post - [GitHub](https://github.com/nicia-ai/typegraph) # An Infinite Supply of Graph Databases import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; TypeGraph turns a SQLite (or PGlite, or Postgres) connection into a typed property graph. If you point that connection at a Cloudflare Durable Object's own storage, you **no longer have to provision the database at all**. Every user, agent conversation, or workspace can have its own graph database without anyone ever creating it, and I think that's a lot of fun. I built a companion repo to show it off, [cf-do-typegraph](https://github.com/nicia-ai/cf-do-typegraph). It wires TypeGraph into a single Durable Object class and runs three real apps on top, entirely on your laptop under `wrangler dev`. It's MIT-licensed. ## Naming is the provisioning step Here's the whole multi-tenant boundary: ```ts // the whole multi-tenant boundary, trimmed from src/worker.ts const stub = env.GRAPH_DO.get(env.GRAPH_DO.idFromName(`${example}:${tenant}`)); return stub.fetch(request); // that tenant's graph, and nothing else, lives here ``` `idFromName` is effectively the provisioning API. The Durable Object it names, and the private SQLite database inside it, comes into existence on the first request and hibernates as soon as nothing is using it, so there's no `CREATE DATABASE`, no row to add to a `tenants` table, and no migration to run against a database that didn't exist five minutes ago. Giving every user, agent, or document its own graph stops being a capacity-planning question, because each one is just a name. Minting a lot of them is just a loop: ```bash curl -X POST "localhost:8787/api/spawn/notes?count=25&prefix=demo" ``` ```json { "example": "notes", "seeded": ["demo-1", "demo-2", "...", "demo-25"] } ``` ```bash curl "localhost:8787/api/notes/demo-7/search?q=hibernation" # its own private graph ``` That endpoint forwards 25 `/seed` requests to 25 different names and lets each Durable Object materialize when its request arrives. There's nothing more to it than that. ## Isolation and shared transactions Making each tenant a Durable Object has two consequences that matter more than the provisioning trick. **Isolation is physical.** Each tenant's SQLite file is its own `ctx.storage` rather than a slice of a shared table. The usual SaaS approach is a `tenant_id` column plus a promise that every query remembers its `WHERE tenant_id = ?`, but here there's no shared table to put that column in, so a bad join or a missing filter can only ever see one tenant's data. **The graph and the app's data live in the same file and the same transaction.** The Durable Object boots its store the same way on every request: ```ts // src/do/graph-do.ts createSqliteBackend(drizzle(this.ctx.storage)); ``` `ctx.storage` is the same SQLite the application would use for its regular tables, and TypeGraph doesn't need a separate connection or service, so writing an app row and the graph edge that describes it can be one `store.transaction(...)`. There's no sync job between your data and your graph, and nothing for them to drift apart on. ## Three apps, one Durable Object class The repo runs the same `GraphDO` class three ways, distinguished only by what you name the tenant. I picked each example because it needs a query that's painful to hand-write in SQL. ### `authz`: a graph per workspace Zanzibar-style permission checks are reachability questions, and TypeGraph's ontology does the reasoning you'd otherwise hand-roll in a `WITH RECURSIVE`: ```ts // src/examples/authz/graph.ts ontology: [ implies(owner, editor), // owner ⇒ editor ⇒ viewer implies(editor, viewer), inverseOf(memberOf, hasMember), // group membership, read backwards, zero stored reverse edges ], ``` A permission check is one `shortestPath` call over the edge kinds the requested relation implies: ```bash curl "localhost:8787/api/authz/acme/check?user=alice&relation=editor&doc=roadmap" ``` ```json { "allowed": true, "relation": "editor", "subject": "user:alice", "via": { "grantSubject": "group:staff", "grantRelation": "editor", "grantTarget": "folder:engineering" }, "pathNodeIds": [ "user:alice", "group:eng", "group:staff", "folder:engineering", "folder:specs", "doc:roadmap" ], "pathEdgeKinds": ["memberOf", "memberOf", "editor", "parentOf", "parentOf"] } ``` Alice never got a direct grant on that doc, and the answer comes with the path that proves she has access anyway: nested group membership (`alice` → `eng` → `staff`), one group-level `editor` grant, and two levels of folder inheritance, all inferred at query time from five stored edges without a permissions table. ### `notes`: a graph per user This one is a personal wiki with `[[wiki-links]]`, and the fun part is that **FTS5 runs inside the Durable Object's own SQLite**, so there's no search service to keep warm: ```bash curl "localhost:8787/api/notes/me/search?q=reachability" ``` ```json { "hits": [ { "id": "note:reachability", "title": "Reachability", "score": "...", "snippet": "Reachability asks whether you can get from one node to another..." } ] } ``` Backlinks don't need any stored rows. The graph only writes a forward `linksTo` edge when a note contains `[[Some Target]]`, and one ontology declaration lets a `linkedFrom` traversal read those same edges backwards: ```ts // src/examples/notes/graph.ts ontology: [inverseOf(linksTo, linkedFrom)], ``` There's no second edge to keep in sync, so backlinks can't drift from the links that produced them. ### `agent-memory`: a graph per conversation Each session is its own graph, seeded by replaying a hand-written conversation one message at a time: ```text "I'm Ada, and I'm PMing the Halo launch." "Grace is my lead engineer — we've shipped three products together." ... "New development: we're partnering with Northwind, and Grace is their main contact." ``` That eighth message is my favorite part of the repo. `Organization` isn't a node kind in the compile-time schema, so it arrives at runtime, gets validated the same way an LLM's proposed schema change would be, and is applied live: ```ts // src/examples/agent-memory/replay.ts const validation = validateGraphExtension(evolution.extension, { strict: true, }); current = await current.evolve(validation.data); ``` The graph grows a new node kind and a new `worksAt` edge kind in the middle of a conversation, while the Durable Object is running. The change is persisted, so it's still there the next time the object wakes from hibernation. And because the store boots with `{ history: true }`, the same graph can answer "what did the agent believe after message 4?", before Northwind existed, by reading at a recorded point in time instead of the live state. ## The same pattern without Cloudflare The underlying idea, **a small, typed graph database per tenant, addressed by name**, works well whenever tenants are numerous, mostly idle, and shouldn't share a query surface. Durable Objects make it especially easy (SQLite at the edge, isolated per object, and hibernating for free), but TypeGraph doesn't know or care that one is underneath, and the same pattern works with: - **A directory of SQLite files.** One file per tenant, opened through `createLocalSqliteStore`. Provisioning is choosing a file path. - **A pool of PGlite files.** Postgres-in-WASM via `createLocalPgliteStore`, one file per tenant, no server process. - **A database per Neon branch.** Real network Postgres, with the standard `createPostgresBackend` pointed at a different connection string per tenant, provisioned through Neon's API instead of `idFromName`. I built the demo on Durable Objects because they need the least infrastructure to try: there's no server to run, and you don't need an account to develop against them. ## Tested against the real thing Every example runs its tests against real Durable Objects via `@cloudflare/vitest-pool-workers`, not mocks, including a cross-tenant isolation test that seeds one tenant and checks that a same-shaped, differently named tenant reads back empty. `pnpm dev` runs the whole thing (landing page, live force-directed graph view, all three APIs) locally for free. Workers AI, the one billable piece, is only used for optional semantic search and entity extraction, and ships commented out. ## Try it - [cf-do-typegraph](https://github.com/nicia-ai/cf-do-typegraph): `pnpm install && pnpm dev`, then open `localhost:8787` - [Bring Your Own Database](/blog/bring-your-own-database): the backend abstraction that lets the same store run on a different SQLite underneath - [Agent Memory That Knows Why It Believes Things](/blog/truth-maintenance-for-agent-memory): the bitemporal history behind the `agent-memory` replay - [An Agent That Grows Its Own Schema](/blog/runtime-schema-evolution): `store.evolve()`, the mechanism behind the live `Organization` extension - [GitHub](https://github.com/nicia-ai/typegraph) # Introducing TypeGraph import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; Every project I've worked on that needed real structure (a knowledge base, an org chart, memory for an agent) ended up with the same stack: an ORM for the actual data, a vector store added when someone wanted semantic search, and eventually, once the relationships got interesting, a graph database off to the side with a sync job keeping it roughly in line with the other two. Each of those systems has its own idea of what a `Person` is and its own consistency model, so the schema drifts between them. Meanwhile TypeScript already has Zod, a schema language good enough to describe all of it, and none of those systems treats it as the source of truth. TypeGraph is the other option: keep the graph inside the application, as a library, and store it in the database you already run. ## It runs in your process TypeGraph writes through your existing SQLite or Postgres connection and commits in the same transaction as the rest of your data. You don't deploy a graph server or keep one running, and your app doesn't make a network hop to reach its own graph. ## One schema You describe your data once, in Zod: ```typescript const Person = defineNode("Person", { schema: z.object({ name: z.string(), role: z.string() }), }); const worksOn = defineEdge("worksOn", { schema: z.object({ since: z.string() }), }); const graph = defineGraph({ id: "org", nodes: { Person: { type: Person } }, edges: { worksOn: { type: worksOn, from: [Person], to: [Person] } }, }); ``` From that one definition TypeGraph derives runtime validation, the TypeScript types, the storage layout, and what the query builder will let you write. If you've ever kept an ORM schema, a folder of hand-written interfaces, and a Cypher cheat sheet in sync by hand, you know why I wanted this. ## Relationships that mean something A foreign key tells you two rows are related, but it can't tell you that a `Podcast` is a kind of `Media`, that `marriedTo` implies `knows`, or that a `Person` and an `Organization` can never be the same thing. In most codebases those rules live in a comment, or in the head of whoever wrote the migration. In TypeGraph, edges are first-class and typed, and they carry their own properties. The ontology layer (`subClassOf`, `implies`, `disjointWith`, `equivalentTo`) does real work: a query for `Media` can include podcasts, a `knows` traversal can pick up spouses, and a write that would make a person and an organization share an identity fails. Traversals compile to SQL, so a three-hop walk is one query rather than a loop of lookups. ## Vectors are just another field Embeddings are a field type. When you declare one, TypeGraph stores and indexes it on whichever backend you're running: ```typescript const Document = defineNode("Document", { schema: z.object({ title: z.string(), embedding: embedding(1536) }), }); const similar = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(queryVector, 10)) .select((ctx) => ctx.d) .execute(); ``` `.similarTo()` is a predicate like any other, so it composes with filters and traversals in the same query. You don't ask a vector database for a list of IDs and then run a second query to find out what those IDs are. ## What it isn't TypeGraph isn't trying to be Neo4j. There's no PageRank, no community detection, no distributed storage, and at this stage no traversal algorithms beyond the queries you write yourself. It's built for thousands to millions of nodes living next to the rest of your data. If your graph _is_ the product and it has billions of edges, you want a dedicated graph database. ## Where it stands It's early, a handful of releases in. The DSL, the ontology layer, vector search, and both backends are solid. What comes next depends on what people build with it, so if you try it and hit a wall, open an issue. ## Try it - [What is TypeGraph?](/overview) and the [Quick Start](/getting-started) - [Ontology & Reasoning](/ontology): `subClassOf`, `implies`, `disjointWith`, `equivalentTo` - [Semantic Search](/semantic-search) - [GitHub](https://github.com/nicia-ai/typegraph) # Turning Two Agents' Event Streams Into One Canonical Graph import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; Imagine a sales bot and a support bot that both talk to the same person without knowing it. The sales bot knows her as "Jane Doe", VP Engineering, at `jane@acme.com`, while the support bot has "J. Doe", VP Eng, at the same email, along with two companies the sales bot has never heard of, one of which the support bot retracts a day later. Both bots emit durable, at-least-once event streams of what they've seen. Event logs are great at delivery, ordering, and replay, but they don't resolve entities, and you can't ask a Kafka topic what an agent believed last Tuesday. That work happens in a layer between the log and whatever reads it, which is what [Materializing Event Logs](/materializing-event-logs) describes. [`@nicia-ai/agent-stream-graph`](https://github.com/nicia-ai/agent-stream-graph) is the reference implementation, built on 0.35's `store.transactionWithReceipt()` and a couple of 0.36 additions. It pulls together three things I'd built separately (bitemporal history, idempotent writes, and graph merge), and seeing them click into one pipeline was a lot of fun. ## Projecting events idempotently Streams redeliver changes after a crash, a reconnect, or a replay, so the second delivery of a change has to land on the same row as the first. In practice that means `upsertById` for nodes and `getOrCreateByEndpoints` for edges, and a bare `create` only when the source event carries its own unique id that you pass through as the TypeGraph id. ```typescript const project: Projector = async (belief, change) => { switch (change.shape) { case "person": { if (change.operation === "delete") { await belief.nodes.Person.delete(asNodeId(change.key)); return; } await belief.nodes.Person.upsertById(change.key, { name: change.value.name, email: change.value.email, title: change.value.title, }); return; } case "employment": { await belief.edges.worksAt.getOrCreateByEndpoints( { kind: "Person", id: change.value.person }, { kind: "Company", id: change.value.company }, {}, ); return; } } }; ``` ## Resuming after a crash `consume()` checkpoints the last processed offset and a recorded-time anchor after every change, and uses `store.transactionWithReceipt()` to know whether the projector actually wrote anything. The demo runs the sales bot's stream, stops it after two of its four changes, then simulates the nastiest crash window there is: a change gets projected, and the process dies before the checkpoint lands. ```text (b) Durable consumer — resume from checkpoint, replay safely after crash consumer ran, then crashed after 2 messages durable cursor: last offset = 002 sales-bot belief so far: 1 person, 1 company crash window: projected 003, then died before checkpoint uncheckpointed anchor existed: 2026-07-14T16:32:30.397Z durable cursor is still: 002 belief already has worksAt edges: 1 restarted — replayed 003, then processed 004 (2 messages) sales-bot belief now: 1 person, 1 company — Jane Doe (VP Eng & Product) worksAt edges after replay: 1 (no duplicate edge) re-run (at-least-once): 0 messages processed; belief unchanged: 1 person, 1 company ``` The cursor still says `002` even though `003` was already projected, which is the crash window the demo sets up deliberately. On restart `003` replays, `getOrCreateByEndpoints` finds the edge it already wrote, and processing carries on. Running the whole stream a third time processes nothing because the cursor is already at the end. ## Rebuilding from offset zero A full rebuild is a different case: replay every event from the beginning into a fresh belief graph, for recovery or a schema migration. Before 0.36, that rewrote every row even when the replayed value matched what was already there, so each rebuild cost a wasted write per event and grew recorded history by the length of the log. `createStore(graph, backend, { coalesceUnchangedUpserts: true })` makes a value-identical redelivery a true no-op that writes nothing, adds no history row, and doesn't advance the revision. A stream that actually changes a value and later changes it back still writes both times, since those changes really happened. 0.36 also adds [`tx.measure()`](/schemas-stores/#scoped-receipts-txmeasure), which scopes a receipt to the writes made by one callback, so a materializer's own cursor bookkeeping in the same transaction can't be mistaken for projector output. ## What did each bot believe, and when? Each bot's belief graph has history enabled, so `book.anchorFor(stream, offset)` gives you a recorded-time coordinate for any processed offset. That means you can read exactly what _this_ bot believed at _that_ point in its own stream: ```text (c) What did each agent believe, at which offset? sales-bot's own belief graph: @offset 001: people: Jane Doe (VP Engineering) | companies: — @offset 004: people: Jane Doe (VP Eng & Product) | companies: Acme Corp (Series A) (title corrected) support-bot's own belief graph (same person, different surface form): @offset 004: people: J. Doe (VP Eng) | companies: Umbrella (unverified), Acme (Series B), Globex (Series C) @offset 005: people: J. Doe (VP Eng) | companies: Acme (Series B), Globex (Series C) (Umbrella retracted) → Same email, but neither agent alone knows 'Jane Doe' and 'J. Doe' are one person. The retracted company also remains visible in past belief. ``` The support bot's belief at offset 4 still shows "Umbrella (unverified)", even though it was deleted at offset 5, because a read at a past offset reconstructs the belief as it was then. ## One canonical graph Each bot's belief exports through streaming interchange into a [graph merge](/blog/graph-merge) branch, and `mergeIncremental()` folds it into a canonical graph. It's the same entity resolution as in the graph merge post, run once per stream as it arrives: ```typescript await importGraphStream( agentBranch.store, exportGraphStream(belief, { includeTemporal: true }), { onConflict: "update" }, ); const result = await mergeIncremental({ forkPoint, target: canonical, branches: [agentBranch], options: { resolve: { Person: { similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.9, }, Company: { similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.9, }, }, onPropertyConflict: "flag", onBasePropertyConflict: "flag", branchOrder: [branchId], persistProvenance: true, }, }); ``` The sales bot merges first, as a clean append. The support bot merges second, and that's where the same-person, different-spelling problem gets resolved: ```text [wave 1] merged sales-bot — conflicts: 0 [wave 2] merged support-bot — conflicts: 4 conflict: Company.name on c1: support-bot="Acme"; kept "Acme Corp" conflict: Company.stage on c1: support-bot="Series B"; kept "Series A" conflict: Person.name on p1: support-bot="J. Doe"; kept "Jane Doe" conflict: Person.title on p1: support-bot="VP Eng"; kept "VP Eng & Product" canonical now: 1 person, 2 companies — Jane Doe (VP Eng & Product) provenance — sales-bot contributed to 3 canonical entities provenance — support-bot contributed to 3 canonical entities ``` "Jane Doe" and "J. Doe" collapse into one canonical person, and every disagreement is flagged instead of silently overwritten. Globex, which only the support bot ever saw, joins as a new company. Umbrella, which the support bot retracted before the merge, never shows up at all. And the canonical graph has history too, so it time-travels across merge waves: ```text canonical, time-travelled: asOfRecorded(after wave 1): 1 person, 1 company asOfRecorded(after wave 2): 1 person, 2 companies ``` So from two unreliable streams you end up with one canonical graph, and you can still ask any of the three graphs what it believed at any point along the way. ## A bigger example The demo above is small on purpose: one person, two bots, five offsets. The repo's main demo (`pnpm demo`) runs the same mechanics against a set of wiki entries that mention overlapping concepts under different names, and converges them into one canonical concept graph with source attribution for every mention. ## Try it - [Materializing Event Logs](/materializing-event-logs): idempotent projectors, cursor bookkeeping, transaction receipts, and mapping event time onto the two temporal axes - [Graph Merge](/graph-merge): the entity resolution this builds on - [`agent-stream-graph`](https://github.com/nicia-ai/agent-stream-graph): `pnpm demo` and `pnpm demo:mechanics` - [GitHub](https://github.com/nicia-ai/typegraph) # Replacing the N+1 Loop With One Statement import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; Most apps have a page like this one: a workspace lists its published documents, and each row shows the title, how many versions the document has, the note on the newest version, and the three most recent comments. Every piece of that is a one-hop read from the document, so the first version anyone writes is a loop: ```typescript const documents = await store .query() .from("Document", "document") .whereNode("document", (document) => document.status.eq("published")) .orderBy("document", "title") .select((ctx) => ctx.document) .execute(); const rows = await Promise.all( documents.map(async (document) => { const versionEdges = await store.edges.hasVersion.findFrom(document); const versions = await store.nodes.Version.getByIds( versionEdges.map((edge) => edge.toId), ); const commentEdges = await store.edges.hasComment.findFrom(document); const comments = await store.nodes.Comment.getByIds( commentEdges.map((edge) => edge.toId), ); // Sort in JavaScript: newest version, version count, three latest comments. return summarize(document, versions, comments); }), ); ``` Counted at the driver, that loop issues 33 statements for eight documents, one for the list and four per document. On in-memory SQLite you'd never notice, but with the database across a network every one of those is a round trip, the count grows with the length of the list, and you're loading every comment to keep three and every version to keep one. This is the N+1, the oldest performance bug in the book, and over releases 0.58 to 0.65 I went after it properly. The same page now costs **one SQL statement**, and when TypeGraph can't do a batch in one statement it throws before running anything instead of quietly issuing more. (Every number here is a statement count, measured by wrapping better-sqlite3's `Statement` and `exec` in the script that ran the code. None of them are wall-clock times. What a round trip costs on your network is yours to multiply in.) ## Step one: reads that say what they want `store.neighbors()` joins each edge to the node on its far end in a single statement, and applies ordering and the limit before hydrating anything, so "the newest version" fetches one row. `store.countNeighbors()` is the matching count and hydrates nothing at all. ```typescript const [latest] = await store.neighbors(document, { edges: ["hasVersion"], orderBy: { by: "node", field: "sequence", direction: "desc" }, limit: 1, }); const versions = await store.countNeighbors(document, { edges: ["hasVersion"], }); ``` Two calls per document is 16 statements for the page, which is better than 33 but still grows with the length of the list. ## Step two: `batchOnce()` `store.batchOnce()` is the part I'm happiest with. It takes any number of independent reads and runs them as exactly one SQL statement: each read becomes a CTE, and a single JSON envelope carries every result set back, in input order. The callback can return a runtime-sized array, which means the loop above translates almost literally: ```typescript const results = await store.batchOnce((read) => documents.flatMap((document) => [ read.neighbors(document, { edges: ["hasVersion"], orderBy: { by: "node", field: "sequence", direction: "desc" }, limit: 1, }), read.countNeighbors(document, { edges: ["hasVersion"] }), ]), ); // results[2 * i] is document i's newest version, results[2 * i + 1] its count. ``` That takes the page from sixteen statements to one. Inside a transaction, `tx.batchOnce()` binds to the open connection, so the batch sees writes made earlier in the same callback and is still one statement. ### It won't quietly fall back A lot of "batch" APIs are really a `for` loop, and you only find out they issued N queries when a latency graph shows it a month later. `batchOnce()` has no fallback path, so if it can't do the job in one statement it throws before executing anything: ```typescript await store.batchOnce((read) => Array.from({ length: 501 }, () => read.countNeighbors(document, { edges: ["hasVersion"] }), ), ); // ConfigurationError: store.batchOnce() accepts at most 500 reads in one statement. ``` Five hundred reads run as one statement, and 501 throws without issuing any. It won't chunk to fit the backend's bind-parameter limit either, and reads that can't be embedded (edge-collection `batchFind*` calls, for example) are rejected instead of being run on the side. Since the name promises one statement, I'd rather it fail loudly than break that promise. The older `store.batch()`, by contrast, runs its queries one after another. It always did, and the docs now say so plainly because people were putting it in hot paths expecting a single round trip. ## Step three: shape it in SQL A batch of `neighbors()` calls still has one read per parent, so eight documents means sixteen reads inside the statement, and the 500-read ceiling works out to 250 documents. When the question is "for every document in this result", a relation answers it with a fixed number of reads however long the list is. 0.61 added `project()`, derived relations, and aggregates; 0.63 added `topPerPartition()`: ```typescript const published = () => store .query() .from("Document", "document") .whereNode("document", (document) => document.status.eq("published")); const versions = published() .traverse("hasVersion", "link") .to("Version", "version") .project((fields) => ({ documentId: fields.document.id, versionId: fields.version.id, sequence: fields.version.sequence, note: fields.version.note, })) .asRelation(); const versionCounts = versions .groupBy((columns) => [columns.documentId]) .aggregate((columns) => ({ documentId: columns.documentId, versions: expr.count(columns.versionId), })); const latestVersions = versions.topPerPartition({ partitionBy: (columns) => [columns.documentId], orderBy: (columns) => [ { expression: columns.sequence, direction: "desc" }, { expression: columns.versionId }, ], limit: 1, }); ``` `topPerPartition()` uses `ROW_NUMBER()`, so ties don't widen the limit. Both orderings end in an id because the API can't know your ordering is unique, and a repeatable winner needs a tiebreaker. The comments are the same shape with a limit of three, plus an ordered collection that folds the winners into one array per document: ```typescript const recentComments = published() .traverse("hasComment", "link") .to("Comment", "comment") .project((fields) => ({ documentId: fields.document.id, commentId: fields.comment.id, body: fields.comment.body, postedAt: fields.comment.postedAt, })) .asRelation() .topPerPartition({ partitionBy: (columns) => [columns.documentId], orderBy: (columns) => [ { expression: columns.postedAt, direction: "desc", nulls: "last" }, { expression: columns.commentId }, ], limit: 3, }) .groupBy((columns) => [columns.documentId]) .aggregate((columns) => ({ documentId: columns.documentId, recent: expr.collect(columns.body, { orderBy: [ { expression: columns.postedAt, direction: "desc", nulls: "last" }, { expression: columns.commentId }, ], }), })); ``` Relations are batch members like any other read, so the whole page, titles included, is one statement: ```typescript const titles = published() .orderBy("document", "title") .select((ctx) => ({ id: ctx.document.id, title: ctx.document.title })); const [documents, counts, latest, comments] = await store.batchOnce( () => [titles, versionCounts, latestVersions, recentComments] as const, ); ``` Joined on `documentId`, a row comes out as: ```text { id: "1BMs3-6PEJYXtjB9BcrsU", versionCount: 3, latest: "v3 of doc 2", recent: ["comment 5 on doc 2", "comment 4 on doc 2", "comment 3 on doc 2"] } ``` The script asserts that these eight rows are identical to what the loop produced. Here's how the versions of the page compare: | Page for eight documents | Statements | | -------------------------------------------- | ---------: | | Loop of `findFrom()` and `getByIds()` | 33 | | Loop of `neighbors()` and `countNeighbors()` | 16 | | `batchOnce()` of the `neighbors()` reads | 1 | | `batchOnce()` of four relations | 1 | `neighbors()` fits when you already hold a handful of sources, and relations fit when you want "every parent in this result". `topPerPartition()` needs window functions and ordered `expr.collect()` needs ordered aggregates; a backend without them throws a typed error before executing. ## The rest of the set-shaped toolkit A few smaller pieces from the same stretch, all aimed at the same habit of doing one thing per row: - **`bulkFindFrom()` / `bulkFindTo()`** are `findFrom()` / `findTo()` for a whole page of endpoints. Index `i` of the result holds the edges of input `i`, an endpoint with no edges gets an empty array, and `limitPerInput` caps each endpoint's fan-out. Fifty people's jobs cost one statement per endpoint kind (split only when the bind-parameter budget requires it) instead of fifty. ```typescript const people = await store.nodes.Person.find({ limit: 50 }); const jobsPerPerson = await store.edges.worksAt.bulkFindFrom(people); ``` - **`store.bulkFindEdgesTo()`** does the inbound direction across several node and edge kinds at once. Twelve documents and two edge kinds was one statement; calling `findTo()` per kind per document was 24. - **`updateWhere()`** is a set-based, transactional update that returns how many rows changed. The selector is mandatory (`where`, `exists`, a candidate query, or an explicit `all: true`), so you can't wipe out a whole kind by forgetting a filter: ```typescript const result = await store.nodes.Person.updateWhere({ patch: { active: false }, where: (person) => person.lastSeen.lt(cutoff), exists: [ { edgeKind: "worksAt", direction: "out", relatedKind: "Company", whereRelated: (company) => company.field("status").string().eq("closed"), }, ], }); // { affectedCount: number } ``` - **`in()` / `notIn()` take list parameters**, so `field.in(param("ids"))` binds a runtime-sized list in a prepared query instead of forcing you to rebuild the query per call. - **Subgraphs got per-edge-kind windows** (keep only the newest N edges of a noisy kind), and `subgraph()` works inside `batchOnce()`. Five roots awaited in a loop cost 10 statements on SQLite; batched, one. - **Mixed-kind cursor pages**: a query can start from several kinds (`.from(["Person", "Team"], "entity")`), and a `.page()` can sit in a batch next to unrelated reads, so a directory of people and teams can be read as one ordered stream at one statement per page. - **`executeChecked(version)`** folds the "has another isolate changed the schema?" probe into the read itself, for serverless deployments that cache the schema per isolate. Probing and then reading took 2 statements, and the checked read takes 1. A moved schema throws `SchemaChangedError`, even when the query would have matched no rows. ## What one statement doesn't buy you - **It saves round trips, but the database does the same work.** Every member of a batch still runs its own plan. For one large closure on Postgres, the direct `subgraph()` can beat the batched form, and `topPerPartition()` bounds the rows returned, not necessarily the rows scanned. - **Results are materialized.** `batchOnce()` returns JSON envelopes, doesn't stream, and has no byte cap. Bound your reads with limits and projections. - **The counts are SQLite driver counts.** They show the shape of the improvement, and your actual latency depends on your network. ## Try it - [`store.batchOnce()`](/schemas-stores#storebatchoncebuildreads-options): the full contract, including what can and can't be embedded - [Top-N per parent](/queries/relations#top-n-per-parent) and [Ordered collections](/queries/relations#ordered-collections) - [`updateWhere()`](/schemas-stores#updatewhereparams) and [`bulkFindEdgesTo()`](/schemas-stores#storebulkfindedgestoparams-options) - [GitHub](https://github.com/nicia-ai/typegraph) # Tracking Which Records Are the Same Person import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; Say a bank's KYC onboarding writes a `Person` node for each customer, keyed by the bank's customer id, and its fraud pipeline writes a `CaseSubject` for everyone named in an investigation. When the person under investigation is a walk-in customer, the case system reuses the bank's id, so one human being ends up as two nodes of different kinds, and nothing in the graph says they're the same person. Most graph libraries handle "these are the same thing" at the type level, and TypeGraph used to as well. The `sameAs` and `differentFrom` ontology factories related _kinds_ rather than rows, and `differentFrom` never checked anything about actual data, which is of no use when a compliance team needs to say that this particular customer is that particular case subject. Both factories are deprecated now. Their replacement, `store.identity`, is one of my favorite things in the library. It's a ledger of claims about individual nodes, either that two of them are the same entity or that they're provably not. You can query it at any point in time and retract a claim without deleting it, the claims carry through traversals and merges, and the database itself won't commit a set of claims that contradicts itself. ## Folding on a shared id Identity is opt-in per graph. Turning it on also decides what happens when nodes of different kinds share an id: ```typescript const graph = defineGraph({ id: "compliance", nodes: { Person: { type: Person }, CaseSubject: { type: CaseSubject }, Organization: { type: Organization }, }, edges: {}, ontology: [disjointWith(Person, Organization)], identity: { sameIdAcrossKinds: "fold" }, }); ``` With `"fold"`, live nodes of different kinds that share an id are the same entity automatically. (`"ignore"` gives you the ledger without that implicit join, for graphs where a shared id is a coincidence.) The two systems above never have to coordinate: ```typescript const person = await store.nodes.Person.create( { name: "Alex Rivera" }, { id: "cust-8842" }, ); const subject = await store.nodes.CaseSubject.create( { note: "Named in fraud case FC-2201" }, { id: "cust-8842" }, ); await store.identity.membersOf(person); // [{ kind: "CaseSubject", id: "cust-8842" }, // { kind: "Person", id: "cust-8842" }] ``` A fold starts when the node actually exists, not at its `validFrom`. A node created today with a backdated validity window shows up in historical reads, but it doesn't retroactively fold anything in the past. ## Saying it explicitly Most matches don't come with a shared id. Here an investigator links a walk-in customer to an actor in an unrelated case, under two ids that have nothing in common: ```typescript const walkIn = await store.nodes.Person.create( { name: "J. Alvarez" }, { id: "person-119" }, ); const caseActor = await store.nodes.CaseSubject.create( { note: "Signed the wire authorization in case FC-2214" }, { id: "case-77" }, ); const linked = await store.identity.assertSame(walkIn, caseActor); // { action: "created", assertion: { id, relation: "same", a, b, validFrom } } await store.identity.assertSame(walkIn, caseActor); // { action: "existing", assertion: } — idempotent ``` Asserting the same thing twice gives you the existing assertion back, so a job that re-runs after a crash doesn't need to know whether it got that far last time. There are bulk forms (`bulkAssertSame`, `bulkAssertDifferent`) that return one result per input pair, in order. Claims are also checked against the ontology before anything is stored. Since `Person` is `disjointWith` `Organization`, folding the walk-in onto a company fails: ```typescript await store.identity.assertSame(walkIn, acmeFreight); // throws IdentityContradictionError // details: { operation: "assertSame", reason: "disjoint-kinds", a, b } ``` Claims can also be bounded in time. A merged account that was later split, or a case subject that was only correct for the window an investigation covered, gets a half-open validity window: ```typescript await store.identity.assertSame(alice, legacyAlice, { validFrom: "2020-01-01T00:00:00.000Z", validTo: "2022-01-01T00:00:00.000Z", }); ``` Both endpoints have to exist for the whole window, and contradictions are checked across every overlapping stretch of time, including through chains of `same` claims. ## Changing your mind Suppose forensic review later shows that `person-119` and `case-77` are two different people who happened to share a wire-authorization signature. You want to record that, but the ledger currently says they're the same entity, and it won't hold both claims at once: ```typescript await store.identity.assertDifferent(walkIn, caseActor); // throws IdentityContradictionError // details: { operation: "assertDifferent", reason: "same-class", a, b } ``` You retract the old claim first, explicitly: ```typescript const ended = await store.identity.retractAssertion(linked.assertion.id); // ended.validTo: "2026-08-21T18:03:11.442Z" — the fold ends, timestamped await store.identity.assertDifferent(walkIn, caseActor); // { action: "created", ... } — now recorded as provably different await store.identity.areSame(walkIn, caseActor); // false await store.identity.areDifferent(walkIn, caseActor); // true ``` Nothing was deleted: the retracted assertion and its `validTo` are still readable, and a query at a time before the retraction still sees one entity, so if an auditor asks why these two records were treated as one customer in August, you can answer them. Soft-deleting a node ends its current assertions too, and records the node as the reason (`endedBy`), so a later reader can tell _why_ a claim ended. ## A backstop in the database Everything so far is application code deciding whether a write is allowed, and application code can be wrong. `assertSame` checks against state it just read, so a bug somewhere else that writes identity rows directly could still commit the contradiction the API rejects. So identity also keeps a table of which identity classes are held apart by a `different` claim, with a database `CHECK` constraint on it. When a transaction merges two classes, it rewrites those rows in the same batch. If two classes that are supposed to be separate get merged anyway, the rewrite violates the constraint and the database aborts the transaction: ```typescript // IdentitySeparationViolationError // details: { graphId, enforcedBy: "database", classKey, // assertionId, a, b } ``` The API catches contradictions first and gives you a useful error, and the constraint is there for the case where the API has a bug. ## Traversals that understand it Queries can expand through identity classes, per hop, and it's off by default: ```typescript const results = await store .query() .from("Person", "person") .traverse("authored", "edge", { includeIdentityMembers: true }) .to("Document", "document") .select((ctx) => ({ edge: ctx.edge, document: ctx.document })) .execute(); ``` That hop follows `authored` edges from the person _and_ from every other node in their identity class, and returns the real edge and target rows with duplicates removed. At the current time, each hop looks classes up through an index on the maintained closure, so the cost tracks the rows you start from and the size of their classes, not how many classes the graph holds. Measured on SQLite, with each `Person` folded to a `Company` and an `Alias` sharing its id, before and after the closure became index-seekable: | source rows | fan-out | matching edges | before | after | | ----------- | ------- | -------------- | --------- | ----- | | 250 | 1 | 250 | 67 ms | 6 ms | | 1000 | 1 | 1000 | 1077 ms | 9 ms | | 2000 | 1 | 2000 | 4616 ms | 19 ms | | 1000 | 8 | 8000 | 8611 ms | 13 ms | | 500 | 200 | 100,000 | 51,602 ms | 77 ms | That closure only describes the present, which is worth planning around. A hop at a _historical_ time has to rebuild classes from the assertion ledger for the whole graph, once per statement, even if you only asked about one node, so queries about the present are much cheaper than queries about the past. Identity claims also carry through interchange and graph merge. Exports carry assertions (current ones by default, ended ones too if you ask for an archival export), and graph merge treats identity as part of what it diffs. Two branches that make opposing claims about the same pair, or that contradict each other through a chain of `same` claims neither wrote alone, fail at plan time with `IdentityMergeConflictError`, and the merge checks again inside its own transaction before committing. ## Staging the duplicate you came to resolve Ingestion is where identity gets messiest. Say a provider feed arrives with a patient carrying MRN `MRN-4471` and the graph already has a `Patient` with that MRN. That's expected, since deciding whether the two are the same person is why the ingestion pass exists, but `Patient` declares `unique("mrn")`, so the second row can't be written. If you stage it on an ordinary `branch()`, staging fails at the duplicate before entity resolution has seen anything. `ingestionBranch()` makes a working copy that defers node uniqueness and nothing else, so schema validation, endpoint checks, disjointness and edge cardinality all still apply while staging. The usual answer to this problem is a "skip validation" flag on the import, which I didn't want, because it lets every other kind of bad row in alongside the duplicates you meant to allow. ```typescript const incoming = unwrap( await ingestionBranch(base, makeBackend, { id: asBranchId("provider-a") }), ); const imported = await importGraph(incoming, providerDocument, { onConflict: "error", onUnknownProperty: "error", }); if (!imported.success) throw new Error("Provider import was rejected"); const alias = await incoming.nodes.Patient.getById( asNodeId("incoming-patient"), ); if (alias === undefined) throw new Error("Imported patient was not found"); // `canonicalPatient` was read from the base before forking. await incoming.identity.assertSame(canonicalPatient, alias); ``` The duplicate MRN and the claim that explains it now sit on the same branch and reach merge planning together. The handle only exposes the assertion methods, without identity reads, retractions, transactions, or access to the underlying store, so it can't be used to get around the deferred constraint. At merge time the original uniqueness rule applies again. `applyMergePlan()` checks uniqueness against the entire resolved write set inside the target transaction, so a valid key handoff or swap goes through as one set (a row-by-row check would see two owners in the middle and reject it). If resolution still leaves two live owners of one MRN, the merge fails with `MergeConstraintConflictError` and writes nothing, so the deferral only gives you room to review the duplicate before merging. ## What it costs A graph without identity enabled pays nothing: no identity SQL, locks, or closure work. With it on: - **One writer per graph for identity-affecting writes on Postgres.** They serialize on a per-graph advisory lock. Other graphs and all reads are unaffected. - **First-time enablement is heavy.** It briefly takes a `SHARE` lock on the shared nodes table, which blocks writes for every graph in that database, and loads the whole graph to build the initial closure. Do it in a quiet window. - **Switching `fold` and `ignore` is a breaking schema change,** because it changes every `areSame` answer on existing data. It needs an explicit migration. - **It needs real transactions.** The bundled SQLite and Postgres drivers have them. Cloudflare D1 and `drizzle-orm/neon-http` don't, and reject an identity-enabled graph at construction. `store.identity` records and propagates the identity claims _you_ make, so deciding that two records are the same person is still up to you or your entity-resolution pass. What it adds is a durable record of those decisions, including when each was made and when it stopped being true. ## Try it - [Operational Identity](/identity): the full guide, including migrating off `sameAs`/`differentFrom` - [Constraint-aware ingestion branches](/graph-merge#constraint-aware-ingestion-branches) - [Graph Merge](/graph-merge): branch, resolve, and merge identity back - [GitHub](https://github.com/nicia-ai/typegraph) # Five Bugs That Didn't Crash import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; A bug that crashes is the easy kind, because it tells you where it is. The ones I worry about in a database library are the ones that return a reasonable-looking answer (zero rows, a successful write, a clean startup) while the data is wrong, so you find out weeks later or not at all. Over the last stretch of releases (0.44 through 0.50, plus one older fix) I've fixed five of those in TypeGraph. None of them warrants a post of its own, but together they show how I want the library to behave, and one of them you could hit just by doing what TypeGraph's own error message told you to do. ## Rows that exist at no point in time A valid-time window is half-open, so `asOf(t)` returns a row when `valid_from <= t < valid_to`. If you backfilled a record you knew had already ended, you'd pass a `validTo` in the past and no `validFrom`, and the write stamped its own instant as `validFrom`, which put the start after the end. A window that runs backwards contains no `t` at all, so the row was stored, counted, and exported, but no temporal read could ever return it. Since 0.48, a write that creates a row (or resets its window) with a `validTo` at or before its own instant and no `validFrom` stores no lower bound instead, so the row reads as "ended at T, start unknown", which is what you meant. A future `validTo` behaves as before. Custom backends can call the same helper the built-in ones use, `resolveStampedValidityLowerBound`, so the rule lives in one place. While I was in there, 0.47 added the thing people were faking with delete-and-recreate: `clearValidTo: true` reopens a window you closed too early, on the same row, with `oneActive` edges rechecked because reopening can create a second active edge. ```typescript await store.nodes.Employment.updateById(employmentId, { patch: { department: "Research" }, clearValidTo: true, }); ``` ### Upgrading doesn't repair old rows, on purpose Older versions could store inverted windows, and upgrading leaves them alone. I think that's the right call: an upgrade that made invisible rows start appearing in historical queries would change your reports and replays without telling you which rows had moved, which is its own quiet bug. So the repair is an explicit operator action, with a dry run: ```typescript import { repairInvertedValidityWindows } from "@nicia-ai/typegraph"; const report = await repairInvertedValidityWindows({ backend: anyBackend, relations: "live-and-recorded", mode: "report", }); // report.counts.recordedNodes === undefined means NOT SCANNED, never "clean" ``` `mode: "apply"` rewrites what `report` counted. Stop writers first, pass the raw backend on a history-enabled store, and repair `"live-and-recorded"` unless you have a reason not to, because fixing only the live rows leaves the recorded history carrying the same backwards window and `asOfRecorded` will keep serving it. The [runbook](/schema-management#repairing-inverted-validity-windows) has the details. ## Constraints that `importGraph` walked straight past TypeGraph lets you declare hierarchy-wide uniqueness (`scope: "kindWithSubClasses"`), `disjointWith(Person, Organization)`, and edge cardinality (`one`, `unique`, `oneActive`). Plain per-kind uniqueness was always backed by a real database key, but these three were enforced by a writer that took the per-graph write lock, checked, and then wrote. That works only as long as every writer takes the lock, and `importGraph` didn't. It also skipped the disjointness and cardinality checks entirely, so an import could commit a `Person` and an `Organization` with the same id, or three edges on a `cardinality: "one"` relationship, and report success. Import is also the path most likely to be carrying data you didn't write yourself. As of 0.50, each of those constraints is also backed by a reservation row whose primary key admits exactly one owner. Taking the constraint means winning that insert, so a writer holding no lock still loses the race it should lose. Import now enforces all three and reports rejected rows in its `errors` like any other violation. Every store and import write also goes through a single write pipeline now. Import had drifted because it had its own hand-built write path, and nobody noticed which rules it was missing. There are two caveats. This only covers writes that go through TypeGraph, so raw SQL inserting into the node or edge tables skips the reservation. And databases created before 0.50 need the new edge-claims table, which the normal bootstrap or the generated migration SQL provides. As with validity windows, the fix prevents new violations and leaves old ones where they are. To find those: ```typescript for (const violation of await store.verifyConstraintFences()) { console.warn(violation.family, violation.target.axis, violation.target.key); } ``` The audit reads the nodes and edges themselves rather than the new reservation tables, because a database written before those tables existed has no reservations in it, and an audit that only looked there would report zero violations on exactly the data you're worried about. It writes and repairs nothing, since picking which of two conflicting rows survives destroys data either way, and that decision should be yours. ## A search index that said it was fine Fulltext and vector search need their own tables. TypeGraph creates them on first use and writes a marker row saying so, and from then on it trusted the marker without checking that the tables were still there. When one went missing out of band (a partial restore, a migration that recreated a schema, or an edge runtime that lost a file), the store opened cleanly and then failed on the first search with a raw driver error about a missing relation. This mostly showed up on Cloudflare Durable Objects, where [one SQLite database per tenant](/blog/infinite-graph-databases) means thousands of small databases that nobody is watching individually. That error is now a `ContributionUnavailableError` with `state: "physical-storage-missing"` and rebuild guidance. The check only runs on the error path, so healthy stores don't pay for it. There's also a three-step ladder for when search is broken, from cheapest to most drastic: | Step | Call | Writes | | ------- | --------------------------------------- | ------------------------------------- | | Probe | `store.probeContributions()` | Nothing. Safe on a replica | | Repair | `store.repairContributions()` | Marker rows, `IF NOT EXISTS` | | Rebuild | `store.rebuildContribution("fulltext")` | Deletes and refills this graph's rows | Start at the top and stop when the probe says `ready`. Rebuild is the only fix when the table exists in a shape the current code no longer produces, and it needs a maintenance window. Vector storage can't be rebuilt, because the embeddings exist only in the table a rebuild would drop, so TypeGraph won't drop it. The [troubleshooting guide](/troubleshooting#contribution-health-probe-repair-rebuild) walks through each state. ## A migration that deleted your runtime kinds This one stings. [Runtime schema evolution](/blog/runtime-schema-evolution) lets an agent add node and edge kinds with `store.evolve()`, and those kinds live in the stored schema rather than in your TypeScript graph definition. `migrateSchema()` committed whatever graph you passed it, so if you migrated with your compile-time graph, every kind added at runtime disappeared from the schema while its rows stayed in the tables where nothing could reach them. The usual way you'd end up calling `migrateSchema()` was the library's own error message telling you to review the changes and run it, so following my advice hid your data. 0.44 folds the stored runtime additions in before committing, the same way store creation already did. It also throws if a migration would drop a kind that still holds rows, unless you pass `{ discardDroppedKindRows: true }`. Two races in the same area got fixed alongside it: - A schema commit could land while another writer was mid-write against the version being replaced. Managed writes now recheck their schema version while holding a lock the commit also needs, so a stale write fails instead of landing against a schema that no longer accepts it. - Removing a kind and adding it back before its cleanup ran made the old rows reappear next to the new ones, and the cleanup then skipped them because the kind was live again. `evolve()` now won't re-add a kind while its cleanup is pending, and cleanup rechecks the schema under the same lock. ## An `implies()` that folded nonsense rows This is the older one, from 0.35. `implies(edgeA, edgeB)` means "a traversal of `edgeB` can include `edgeA` edges", which only makes sense if `edgeA`'s endpoints could stand in for `edgeB`'s. Nothing checked that, so this was accepted: ```typescript edges: { authored: { type: authored, from: [Author], to: [Paper] }, covers: { type: covers, from: [Paper], to: [Topic] }, }, ontology: [implies(authored, covers)], ``` and a `covers` traversal with `expand: "implying"` would fold in rows that start at an `Author`. Now it fails when the graph is built or loaded: ```text ConfigurationError: implies("authored", "covers") is endpoint-incompatible: from kind(s) [Author] declared on "authored" cannot be assigned to any of "covers"'s from kind(s) [Paper]. ``` The check also runs on schemas loaded from the database, so a bad relation saved by an older version fails on first load after upgrading. The same release made projected ids keep their `NodeId` brand through `.select()` and gave every fixed-shape error class a typed `details`, so fewer mistakes make it to runtime in the first place. ## What they have in common All five fixes follow the rules I now hold the whole library to. If TypeGraph can't do what you asked, it throws and tells you why rather than returning an empty result that looks like a clean one, and it doesn't rewrite your data behind your back, even to fix it. You get a report first and decide what to do. ## Try it - [Repairing inverted validity windows](/schema-management#repairing-inverted-validity-windows) - [Claim relations, and what they do not promise](/backend-setup#claim-relations-and-what-they-do-not-promise) - [Contribution health: probe, repair, rebuild](/troubleshooting#contribution-health-probe-repair-rebuild) - [Schema Management](/schema-management): migrations and kind removal # An Agent That Grows Its Own Schema import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; If you follow a clinical trial through the literature, you see it registered first, then published as a paper, and occasionally, years later, retracted. A retraction notice is a kind of record that wasn't in anyone's schema when the trial was registered, and usually wasn't there when the paper was indexed either. TypeGraph's schema normally lives in code: you declare `defineNode` and `defineEdge` in TypeScript, you get full type inference, and adding a new kind of thing means a code change and a deploy. I think that's the right default and I'm keeping it, but it doesn't fit an ingestion pipeline pulling from an API you don't control, a multi-tenant app where each tenant brings its own shape, or an agent that discovers a new kind of record halfway through a corpus. 0.25 adds **graph extensions** for those cases. A schema change can be proposed at runtime, validated as strictly as the compile-time path, and committed atomically without a redeploy. ## Proposing a change An extension is a plain JSON document (new node and edge kinds, their property types, unique constraints, indexes) built with `defineGraphExtension` and committed with `store.evolve()`: ```typescript const proposal = defineGraphExtension({ nodes: { Paper: { properties: { title: { type: "string", minLength: 1 }, doi: { type: "string", minLength: 1 }, year: { type: "number", int: true, min: 1900, max: 2100 }, }, unique: [{ name: "paper_doi_unique", fields: ["doi"] }], }, }, }); const evolved = await store.evolve(proposal); ``` The type vocabulary is small on purpose: strings, numbers, booleans, enums, and one level of array/object nesting. An LLM-written schema stays readable, and it can't smuggle in a Zod refinement or a function. A malformed proposal throws `GraphExtensionValidationError` with per-field issues before anything touches the database. The commit is checked against the active schema version, so two writers racing to extend the same graph get a `StaleVersionError` to retry on, not a silent overwrite. Reads get the same care. TypeScript can't see a kind that didn't exist at compile time, so runtime kinds go through string-keyed versions of the query builder, which check kind names against the live schema: ```typescript const rows = await store .query() .fromDynamic("Paper", "p") .traverseDynamic("authoredBy", "a") .toDynamic("Author", "u") .select((ctx) => ({ paper: ctx.p, author: ctx.u })) .execute(); ``` A typo in a kind name throws `KindNotFoundError`, where a schemaless store would quietly return zero rows and leave you to find out later. ## Letting an agent drive A toy example is fine for the API, but I wanted to know whether the loop holds up against messy, real, multi-stage data. So I built [`typegraph-clinical-demo`](https://github.com/pdlug/typegraph-clinical-demo): the same machinery end to end, over real trial registrations from ClinicalTrials.gov, their publications from PubMed, and retractions and corrections from CrossRef. It's all public bibliographic metadata, with no patient data. An LLM agent watches the corpus arrive in three stages and proposes an extension after each one. The proposals are **blind**: the agent sees sample records and the names of kinds already in the graph, and I didn't give it any hints about types, searchable fields, or which identifier should be unique. Property types, optionality, searchability, and constraints are all inferred from the samples. - **Stage 1, registration.** The agent sees about 1,100 ClinicalTrials.gov records and proposes `ClinicalTrial`. - **Stage 2, publication.** PubMed records referencing those trials arrive, and the agent proposes `Publication` with a `referencesTrial` edge back to stage 1. - **Stage 3, retraction.** Retraction notices and corrections arrive. Nobody designs for these up front, because most trials never get one. The agent proposes `PublicationEvent` with a `correctsPublication` edge back to stage 2. Each stage is a real `store.evolve()` against a real store, followed by a bulk ingest under the new schema. ## When the proposal is valid but wrong Validation catches malformed proposals, but it can't catch one that's internally consistent and still wrong for the data, like a field the agent marked required because the three samples it saw all happened to have it. So after each accepted proposal the demo runs a **smoke test**: ingest the full sample into a scratch store seeded with every earlier extension. Failures go back to the agent in the same structured `{path, code, message}` form validation uses. Stage 3 is where it fires, on a real run: ```text STAGE 3: post-publication discourse agent attempt=1 validator ACCEPTED +nodes=[PublicationEvent] +edges=[correctsPublication] smoke test FAILED 2 issue(s) — proposal validates but doesn't fit the data: [INGEST_INVALID_TYPE] artifactDoi: expected string, received undefined [INGEST_INVALID_TYPE] date: expected string, received undefined agent attempt=2 validator ACCEPTED +nodes=[PublicationEvent] +edges=[correctsPublication] smoke test PASSED schema fits the sample PublicationEvent nodes: 922/922 ingested (0 skipped) ``` The first proposal saw three samples, all `correction` events with `artifactDoi` and `date` filled in, and made both required. The full-sample ingest hit `comment` events without them and failed. On attempt 2 the agent made both optional, the smoke test passed, and the graph committed. The whole stage, repair included, takes about 16 seconds and one extra model call. I love watching this loop work. Nobody explained the mistake to the model in prose; it got the same machine-readable errors a developer would have seen and fixed its own schema. Because the smoke test runs against a scratch store, a failed attempt leaves no stray rows or half-claimed unique keys for the next attempt to trip over. Only a proposal that survives gets committed to the real store. ## The payoff After stage 3 the graph has three kinds and two edges that didn't exist when the demo started, and one query walks all of them: ```typescript const rows = await store .query() .fromDynamic("PublicationEvent", "event") .traverseDynamic("correctsPublication", "correction") .toDynamic("Publication", "pub") .traverseDynamic("referencesTrial", "reference") .toDynamic("ClinicalTrial", "trial") .select((ctx) => ({ event: ctx.event, publication: ctx.pub, trial: ctx.trial, })) .execute(); ``` On the real corpus it surfaces four documented retraction chains, including SCIPIO (cardiac stem cells, Lancet 2011, retracted 2019) and a 2018 nilotinib trial outcome, plus three more the corpus turned up on its own: the Anil Potti genomic-predictor case and the Mehra et al. hydroxychloroquine retraction that halted multiple registered trials in 2020 among them. Getting there took no migrations, just three `store.evolve()` calls and a query written against kind names that were plain strings until the agent defined them. ## You don't need a frontier model for this The demo includes an eval harness (`pnpm eval`) that runs the same blind three-stage pipeline across a lineup of models and scores whether each first proposal survives validation and the smoke test. The cheapest model that cleared all three stages on the first try was a 26B-parameter open-weight MoE with 4B active parameters. It came in 4x cheaper than Gemini 3.1 Flash Lite (which also went three for three) and 5x cheaper than GPT 5.4 nano (which needed one repair). The full three-stage run, repair included, takes under 30 seconds and costs about **$0.003** in API calls. The loop needs a structured error channel and a model that can read a JSON Schema error and try again, and when the schema layer does its job, a small model handles that fine. ## Try it - [Graph Extensions](/graph-extensions): the full reference - [Agent-Driven Schema](/examples/agent-driven-schema): a minimal, in-repo version of the same loop - [`pdlug/typegraph-clinical-demo`](https://github.com/pdlug/typegraph-clinical-demo): clone it, `pnpm demo`, and watch the schema grow - [GitHub](https://github.com/nicia-ai/typegraph) # Schema Changes That Roll Back With Everything Else import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; [Runtime Schema Evolution](/blog/runtime-schema-evolution) ended with an agent adding a kind for retraction notices to a live graph. This post uses a cut-down version of that: a `Retraction` kind added to a graph of publications. `store.evolve()` commits a change like that in a transaction of its own and hands back a new Store. That's fine when the schema change is the whole job, but not when your service keeps its own bookkeeping next to the graph, say a `schema_audit` table recording which schema version each import ran under. If you do it in the obvious order, `evolve()` commits first, then a second transaction writes the retraction row and the audit row. When something downstream throws in that second transaction, this is what's left: ```text evolve(), then a failed import: active schema version 2 schema_audit rows 0 Retraction node rows 0 ``` The schema is at version 2 but nothing else moved, so there's a `Retraction` kind that no ledger entry mentions and no row uses. Nothing is corrupt, but your bookkeeping and the database now disagree about which schema exists, because the schema change was the one step that couldn't join the transaction. As of 0.62 it can join, and the rest of this post shows how. ## Plan outside, apply inside The setup is a graph with one kind, the extension the agent proposed, and a Drizzle table for the audit ledger in the same database. ```typescript const Publication = defineNode("Publication", { schema: z.object({ doi: z.string(), title: z.string() }), }); const graph = defineGraph({ id: "trials", nodes: { Publication: { type: Publication } }, edges: {}, }); const retractions = defineGraphExtension({ nodes: { Retraction: { properties: { doi: { type: "string", minLength: 1 }, reason: { type: "string" }, }, }, }, }); const schemaAudit = sqliteTable("schema_audit", { id: integer("id").primaryKey({ autoIncrement: true }), schemaVersion: integer("schema_version").notNull(), recordedAt: text("recorded_at"), }); ``` The change is split into planning and applying. `planEvolution()` runs before any transaction opens, validates the extension against the active schema without writing anything, and returns an immutable plan that names the schema it starts from and the one it produces: ```typescript const plan = await store.planEvolution(retractions); ``` ```json { "status": "change", "graphId": "trials", "baseline": { "version": 1, "hash": "89654f9e5c1debfb" }, "result": { "version": 2, "hash": "aec6b8d50bf21bf8" }, "requirements": [ { "kind": "new-kind", "entity": "node", "kindName": "Retraction" } ] } ``` Doing the planning outside means the slow, fallible part (validation, diffing) never holds a lock. For ordinary additions like a new kind or an optional scalar field, applying the plan needs no entity scans and no DDL. Applying takes _your_ transaction. better-sqlite3 is synchronous, so Drizzle's `db.transaction()` can't take an async callback, and a small helper drives `BEGIN`/`COMMIT`/`ROLLBACK` on the one connection. With node-postgres or libSQL you'd pass the `nativeTx` from `db.transaction(async (nativeTx) => …)` instead; [the cross-store transactions recipe](/recipes#cross-store-transactions-drizzle--typegraph) covers both. ```typescript async function inTransaction( db: BetterSQLite3Database, run: () => Promise, ): Promise { db.run(sql`BEGIN`); try { const result = await run(); db.run(sql`COMMIT`); return result; } catch (error) { db.run(sql`ROLLBACK`); throw error; } } ``` ```typescript const outcome = await inTransaction(db, async () => { const applied = await store.withEvolvedTransaction(db, plan, async (tx) => { const notices = tx.getNodeCollection("Retraction"); if (notices === undefined) throw new Error("Retraction kind missing"); await notices.create({ doi: "10.1/a", reason: "fabricated" }); }); db.insert(schemaAudit) .values({ schemaVersion: applied.receipt.schema.version, recordedAt: String(applied.receipt.recorded), }) .run(); return applied; }); const current = await store.refreshSchema({ minVersion: outcome.receipt.schema.version, }); ``` Inside the callback, `tx` already sees the new schema: `Retraction` is writable even though the committed schema doesn't have it yet. The receipt carries the exact schema version and hash the transaction produced (and, on a history-enabled store, the recorded-time anchor), so the audit row stores what the import actually ran under instead of what the code assumed. Treat the receipt as provisional until your outer `COMMIT` succeeds; after that, `refreshSchema()` hands the new schema to the Store you keep around. To check the rollback, I threw an error after the ledger insert and ran the same plan twice, once failing and once clean: ```text attempt 1 (fails after the callback): receipt: schema v2, recorded r1:0000000000000002:2026-09-20T19:12:56.201Z after rollback: active schema version 1 schema_audit rows 0 Retraction node rows 0 attempt 2 (same plan): receipt: schema v2, recorded r1:0000000000000002:2026-09-20T19:12:56.202Z after commit: active schema version 2 schema_audit rows 1 Retraction node rows 1 ``` The failed attempt got as far as a receipt and a ledger row visible inside the transaction, and none of it survived; when I checked the recorded-history table directly it had no `Retraction` rows either. The retry reused the same plan, got the same recorded revision, and committed, so the schema, the rows, the history, and your own table commit or roll back together. If another writer evolves the graph between planning and applying, the apply throws `StaleVersionError` and the transaction is yours to roll back and replan. ## No-ops, and schema-only checkpoints A pipeline that runs the same wiring on every deploy will mostly produce no-op plans. Plan an extension the Store already has and you get `status: "noop"`, which doesn't need the exclusive schema lock, so an ordinary `withRecordedTransaction()` is enough: ```typescript const again = await current.planEvolution(retractions); if (again.status === "noop") { await inTransaction(db, () => current.withRecordedTransaction(db, async (tx) => { const notices = tx.getNodeCollection("Retraction"); if (notices === undefined) throw new Error("Retraction kind missing"); await notices.create({ doi: "10.1/b", reason: "duplicate publication" }); }), ); } ``` The opposite case is a schema change that is itself the event worth recording. On a history-enabled store, `tx.requestRecordedRevision()` puts the change on the recorded timeline even when no entities change: the receipt reports zero writes, schema version 2, and a recorded anchor. ## Merges and reviewed writes that bring their own kinds 0.61 added `applyMergePlanInTransaction()` for applying an approved merge plan next to your own SQL. But a merge plan is built against a specific schema, so a merge that introduces a kind the target doesn't have yet couldn't share a commit with the evolution that adds it. 0.62 closes that gap: `branchForEvolution()` forks a branch that already has the planned kinds, and `planMergeForEvolution()` plans against the resulting schema. ```typescript const evolutionPlan = await target.planEvolution(retractions); const futureBranch = unwrap( await branchForEvolution(target, evolutionPlan, makeIsolatedBackend), ); try { await futureBranch.store .getNodeCollectionOrThrow("Retraction") .create({ doi: "10.1/a", reason: "fabricated" }); const mergePlan = unwrap( await planMergeForEvolution(target, evolutionPlan, [futureBranch]), ); await inTransaction(db, () => target.withEvolvedTransaction(db, evolutionPlan, (tx) => applyMergePlanInTransaction(target, tx, mergePlan), ), ); } finally { await futureBranch.close(); } ``` A plan built the ordinary way, against the old schema, is rejected inside the evolved transaction with `MergePlanSchemaMismatchError` before anything is written. 0.65 does the same for candidate write sets, the branch-free way to propose changes as a JSON document a reviewer can read. `planCandidateWriteSetForEvolution()` plans the agent's proposed `Retraction` records against the schema the pending evolution will produce, and nothing is written until the reviewed plan is applied inside `withEvolvedTransaction()`. Because review takes time and the target can move in the meantime, there are two different stale outcomes. If a write lands _while_ the plan is being built, planning returns a `MergePlanningStaleError`, and you recapture the target and replan. If the target changes after you already hold a finished plan, applying throws `StaleMergePlanError` inside the outer transaction, and the schema change rolls back with it. I injected a write after planning to check this, and the active schema was still at version 1 afterwards. ## Limits - **Plans stay in one process.** A plan is an in-memory token that can't be serialized or reconstructed, so plan and apply in the same process. - **The default adapter only does DML.** A plan that needs new storage (a vector slot, identity work) is rejected before the callback runs, unless you use a privileged adapter configured with `schemaProvisioning: "transactional"`. Indexes stay outside too: call `materializeIndexes()` after commit. - **Only your database's writes are atomic.** The ledger row is atomic with the schema change because they share a connection. A message you publish to a queue from inside the callback is not. - **Driver support varies.** SQLite adoption needs a native connection with an observable transaction state (better-sqlite3 has one); HTTP-only drivers can't adopt schema transactions at all. Custom `StoreEvolution` implementations need `planEvolution()` and `refreshSchema()`, and custom adapters must now declare `schemaProvisioning` explicitly. ## Try it - [Plan outside and apply inside a caller-owned transaction](/graph-extensions#plan-outside-and-apply-inside-a-caller-owned-transaction) - [Merge after schema evolution in one caller transaction](/graph-merge#merge-after-schema-evolution-in-one-caller-transaction) and [candidate write sets for a planned schema](/graph-merge#candidate-write-sets-for-a-planned-schema) - [Adopted schema transactions](/backend-setup#adopted-schema-transactions): per-driver requirements - [GitHub](https://github.com/nicia-ai/typegraph) # One Round Trip Per Write import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; I turned on statement logging, turned off prepared statements, and created one node. This is what went over the wire: ``` begin SELECT ... schema-version fence probe SELECT ... duplicate/endpoint check INSERT ... (the write) commit ``` That's five round trips to write one row. On a pooled connection sitting next to the database you'd never notice, but a lot of people run TypeGraph from edge functions against managed Postgres, where a round trip is more like 45ms and each of those statements pays it in full, so the row takes roughly a quarter of a second to exist. When I captured a first-run provisioning flow for one graph, it issued 97 statements, and about 40 of them, spread across just seven writes, were schema-version probes, duplicate checks, and `begin`/`commit` rather than actual writes. That was [issue #533](https://github.com/nicia-ai/typegraph/issues/533), and closing it took two releases. The short version: an eligible write and every check it depends on now go to the database together, as one request. A note on the numbers: everything below is a count of requests crossing the transport, which is what I measured. Latency figures are that count times an assumed 45ms round trip, not wall-clock benchmarks. ## Checks and write, in one statement The old path was a sequence of reads, each deciding whether the write could proceed: whether the schema version still matched, whether the row was a duplicate, whether the edge's endpoints existed, and whether the cardinality constraint held. Each of those was another round trip, and the whole sequence was wrapped in a transaction so nothing could change between the checks and the insert. In 0.52, for the common shape on Postgres (schema-managed, generated ID, no history or identity tracking), all of those checks and the insert compile into a single statement, and the database evaluates the conditions and performs the write atomically. That takes **five or six sequential requests down to one**, which at 45ms removes roughly 180–225ms of waiting from every such write. On Neon's WebSocket driver, a typical get-or-create that misses drops from five requests to one. When a write needs something a single statement can't carry (history capture, a caller-owned transaction, call-level `matchOn`), it takes the ordinary transactional path, which is still there underneath for everything the fast path doesn't cover. ## Batches on drivers without transactions The bigger win is on Neon HTTP, Cloudflare D1, and libSQL, which don't give you an interactive transaction at all, only one-shot atomic batches. Before 0.52, a bulk write on one of them ran statement by statement, so it was slow and a partial failure left partial data behind. Now TypeGraph compiles eligible bulk writes into a precompiled program that those drivers execute as a single atomic batch. `nodes.bulkInsert()` becomes one request, edge batches validate their endpoints inside the write instead of reading candidate endpoints first, and batches with a cardinality constraint (`one`, `unique`, `oneActive`) carry the constraint checks in the same batch. Counting actual `execute()` / `batch()` calls on libSQL: | Bulk edge write | Before | After | | ----------------------------- | -----: | ----: | | Unconstrained | 1 | 1 | | With a durable match identity | 6 | 1 | | With a cardinality constraint | 8 | 1 | Since it's all one batch, a conflict anywhere rolls the whole thing back. I wanted proof that the conflict check actually did something, so I wrote its regression test by deleting the conflict clause, watching duplicates land, and then putting the clause back and watching the same input get rejected. ## 0.53: the writes that have to look first 0.52 fused the writes that create things, which left the ones that need to know what's already there: updates, upserts, and deletes that have to release the uniqueness claims their row held. 0.53 moved those onto the same programs. - **Single-row `update()` and `delete()`** are one read plus one guarded write, rather than a transaction around both: 60–67% fewer requests. - **`bulkUpsertById()`** is two requests: one batched read, one atomic write. - **`bulkReplaceById()`** is new. Each item is a complete document, so there's nothing to read first. Creates, replacements, resurrections of deleted rows, uniqueness claims, and fulltext and vector index updates all go in one request. - **Large D1 upserts stay atomic.** D1 allows 100 bind parameters per statement, so the program splits into many statements inside one atomic batch. That raised the ceiling from 17 nodes and 6 edges to 512 nodes and 187 edges per call. Anything bigger falls back to the regular path instead of building an unbounded request. - **Postgres transaction sessions run the same programs** through a savepoint, so a rejected program doesn't poison the transaction around it. There's one trade-off to know about: the fused `update()` is optimistic. It doesn't hold a lock between its read and its write, so if the row moved underneath it, it re-reads and tries again, up to four times, and then throws `DatabaseOperationError`. Under sustained contention on the same row, that can fail a write that a transaction-capable backend used to serialize for you. I think that's the right trade for most workloads, but not for something like a hot counter. ## Letting the database decide: match identities Fusing an edge get-or-create needed something the database could arbitrate on its own, without a lock or an in-process cache: a real unique constraint. So 0.52 also added **durable edge match identities**. An edge can declare a named set of fields that identify it: ```typescript const worksAt = defineEdge("worksAt", { schema: z.object({ role: z.string() }), }); edges: { worksAt: { type: worksAt, from: [Person], to: [Company], cardinality: "many", matchIdentity: { name: "employment", fields: ["role"] }, }, }, ``` TypeGraph stores that key on every edge row and backs it with a unique index, on both SQLite and Postgres. Once that exists, `getOrCreateByEndpoints` no longer has to read, decide, write, and hope nothing raced, because it's a single conditional insert the database resolves atomically, which is what makes the one-request path safe on a stateless edge worker with nothing in memory to lean on. Changing a match identity is a breaking schema change and is rejected while the edge kind has rows. Call-level `matchOn` still works for matching you don't want in the schema, through the regular transactional path. ## Upgrading If you use a bundled backend (`createLocalSqliteBackend`, `createPostgresBackend`, or one of the serverless factories), you don't need to change any code. There are two things to know. An existing or externally provisioned database has to be opened once through `createStoreWithSchema()` (or get TypeGraph's generated base-schema migration) before the zero-DDL verified-store paths will run; until then they fail early with `BaseSchemaMigrationError`. And `store.transaction()` now throws on a backend without interactive transactions instead of quietly running your callback without one. If you maintain a custom backend, `GraphBackend.commands` is now required and a few capability flags changed shape. The [authoritative command sessions](/backend-setup#authoritative-command-sessions) section has the migration. ## Try it - [Batch write patterns](/performance/overview#batch-write-patterns) and [remote edge convergence](/performance/overview#remote-edge-convergence): which writes are eligible, and the measured counts - [Upgrading deployment-wide base storage](/backend-setup#upgrading-deployment-wide-base-storage) - [Changelog](/changelog#0530) for 0.53.0, and [0.52.0](/changelog#0520) - [GitHub](https://github.com/nicia-ai/typegraph) # Letting an Edge's Targets Depend on Its Source import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; This schema looks reasonable, but it allows more than you probably meant: ```typescript const assignedTo = defineEdge("assignedTo", { from: [Employee, Student], to: [Department, Course], }); ``` You want employees assigned to departments and students assigned to courses, but array-valued `from` and `to` declare the Cartesian product, so every source may point at every target and both `store.edges.assignedTo.create(student, department)` and `create(employee, course)` succeed. Until 0.55 the only way to rule those two combinations out was to split the edge into two kinds and query them separately. ## Say which source gets which targets `to` can now be a map from source kind to its own allowed targets: ```typescript const assignedTo = defineEdge("assignedTo", { from: [Employee, Student], to: { Employee: [Department], Student: [Course], }, }); const graph = defineGraph({ id: "assignments", nodes: { Employee: { type: Employee }, Student: { type: Student }, Department: { type: Department }, Course: { type: Course }, }, edges: { assignedTo }, }); ``` `create(employee, department)` and `create(student, course)` still work, and the other two combinations fail before anything is written: ```text EndpointPairError: assignedTo: undeclared endpoint pair { edgeKind: "assignedTo", endpoint: "pair", fromKind: "Employee", toKind: "Course", allowedPairs: [ { from: "Employee", to: "Department" }, { from: "Student", to: "Course" }, ] } ``` In typed code you won't get that far, because the collection's `create()` signature narrows per source and handing it an employee and a course is a compile error. The runtime check covers the paths a type checker can't see, such as dynamic collections, bulk writes, and imports. An endpoint kind that isn't in `from` at all still throws the existing `EndpointError`; `EndpointPairError` is specifically for two kinds that are each valid but were never declared together. The map itself is checked when you call `defineEdge()`: every kind in `from` needs an entry, extra keys aren't allowed, and no entry may be empty. A mistake throws `ConfigurationError` at that point rather than turning up later as a confusing write failure. ## Narrowing, never widening An edge's map is its outer bound, and a graph or a runtime [graph extension](/blog/runtime-schema-evolution) can register a narrower version: ```typescript edges: { assignedTo: { type: assignedTo, from: [Employee], to: { Employee: [Department] }, // this graph has no students }, } ``` A registration can't add a pair the edge never declared. That includes the easy mistake of registering the old array form, `to: [Department, Course]`, against an edge whose map only allows the correlated pairs, which would quietly re-admit the cross-pairs the map is there to forbid, so it's rejected. Subclasses are checked against the declared pairs. With `SubTask subClassOf Task` and an edge that allows `Task: [Task]`, every mix of `Task` and `SubTask` works, but a `SubTask` can't borrow a target that only some other source kind is allowed. ## Bulk writes are all or nothing ```typescript await store.edges.assignedTo.bulkCreate([ { from: employeeA, to: departmentA }, // valid { from: employeeB, to: courseA }, // undeclared pair ]); // throws EndpointPairError; store.edges.assignedTo.count() is still 0 ``` A single undeclared pair rejects the whole batch, because committing the valid half and quietly dropping the rest would leave you guessing which rows made it. ## Pairs are stored in the schema The pairs are stored in the serialized schema, and declaring them in a different key order doesn't change the schema hash. Replacing an array `to` with a map removes pairs even though it removes no kinds, so it counts as a breaking change and follows the same migration rules as any other breaking edge change. See [endpoint pair changes](/schema-management#endpoint-pair-changes). ## Try it - [Source-Dependent Targets](/core-concepts#source-dependent-targets): the full reference - [`EndpointPairError`](/errors#endpointpairerror) - [Endpoint pair changes](/schema-management#endpoint-pair-changes) - [GitHub](https://github.com/nicia-ai/typegraph) # Look Before You Write import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; Suppose you tighten the schema on an account directory so that `email` has to be a real address, `plan` has to be one of three values, and `seats` has to be positive. TypeGraph validates on write, so everything written from then on is checked, but rows stored before the change are never looked at again, and you have no idea how many of them fail the new rules. You'd like a background agent (or yourself, in a REPL) to find and fix them while sales reps carry on editing the same accounts. To do that safely the agent needs a cheap overview of what's in the graph, a way to find the rows that don't validate without reading everything back into memory, and a write that won't clobber a change a person made after the agent read the row. 0.54 added `store.describe()`, `store.validateStore()`, and `compareAndSet()` for those three jobs, plus `planCandidateWriteSet()` for reviewing a whole batch before it lands, and 0.64 lets the first two run inside a transaction. It's the same loop [runtime schema evolution](/blog/runtime-schema-evolution) uses when an agent proposes changes to the schema itself: the agent looks at the current state, proposes a change, and the change only applies if nothing has moved since it looked. The graph here is 240 accounts and 90 people created through the store, 60 `memberOf` edges, and three legacy `Account` rows inserted through the backend directly, the way a writer running the old rules would have stored them. The outputs below come from running this code against it on SQLite. ```typescript const Account = defineNode("Account", { schema: z.object({ name: z.string().min(1), email: z.email(), plan: z.enum(["free", "team", "enterprise"]), seats: z.number().int().positive(), ownerId: z.string().optional(), }), }); ``` ## What's in the graph ```typescript const { statistics } = await store.describe(); for (const kind of [...statistics.nodes, ...statistics.edges]) { console.log(`${kind.entity} ${kind.kind}: ${kind.count}`); for (const property of kind.properties) { console.log( ` ${property.path} present=${property.presentCount} coverage=${property.coverage.toFixed(2)}`, ); } } ``` ```text node Account: 243 /email present=243 coverage=1.00 /name present=243 coverage=1.00 /ownerId present=60 coverage=0.25 /plan present=243 coverage=1.00 /seats present=243 coverage=1.00 node Person: 90 /name present=90 coverage=1.00 /title present=30 coverage=0.33 edge memberOf: 60 /role present=30 coverage=0.50 ``` Every declared kind gets a row count, and every declared property gets a presence count and a coverage figure. `/ownerId` at 0.25 says 60 of 243 accounts have an owner, which is enough for an agent to decide where to spend its attention before it reads a single row. The counting happens in the database, and on this graph `describe()` runs two statements, one for all node kinds and one for all edge kinds. The result also records the schema version and hash it was computed under, and TypeGraph checks that they didn't change mid-way, so you never get statistics that straddle a schema change. ## What no longer fits ```typescript let cursor: string | undefined; do { const page = await store.validateStore({ entity: "node", kind: "Account", pageSize: 100, ...(cursor === undefined ? {} : { cursor }), }); console.log( `scanned ${page.scannedCount}, ${page.violations.length} violation(s)`, ); for (const failure of page.violations) { console.log(failure.id, failure.path, failure.reason); } cursor = page.nextCursor; } while (cursor !== undefined); ``` ```text scanned 100, 0 violation(s) scanned 100, 0 violation(s) scanned 43, 3 violation(s) acct_legacy_1 /email Invalid email address acct_legacy_2 /plan Invalid option: expected one of "free"|"team"|"enterprise" acct_legacy_2 /seats Too small: expected number to be >0 ``` There are three violations across two records, because `acct_legacy_2` breaks two rules. Each failure carries the record id, a JSON pointer to the property, the Zod issue code, and the reason. Each page is one bounded statement, and the cursor is tied to the schema it started under. If the schema changes mid-sweep, the next page throws `StoreAnalysisCursorStaleError` instead of quietly checking the rest against different rules. The third legacy row, `acct_legacy_3`, isn't reported. It carries a `salesforceId` the schema doesn't declare, and undeclared properties count as healthy extra data, so you can sweep a graph whose shape is changing at runtime without every extra field showing up as a defect. ## Fix a row without racing anyone The obvious repair is to call `getById()`, decide what to change, and call `update()`, but a sales rep's edit that lands between the read and the write gets overwritten. `compareAndSet()` closes that gap by checking the expected values and applying the update in one statement. ```typescript const applied = await store.nodes.Account.compareAndSet(id, { expected: { name: "Northwind", plan: "team", seats: 12 }, patch: { email: "ops@northwind.example" }, }); ``` ```text applied=true email=ops@northwind.example version 1 -> 2 ``` The interesting case is when someone else gets there first. Say the agent read `acct_legacy_2` and planned a repair, but before it applied, a person fixed the row by hand and changed the email along the way. The agent's guard names the email it read: ```typescript const applied = await store.nodes.Account.compareAndSet(id, { expected: { name: "Contoso", email: "it@contoso.example" }, patch: { plan: "team", seats: 5 }, }); ``` ```text applied=false seats=20 version 2 -> 2 updatedAt changed=false ``` `false` means nothing was written, so the person's `seats=20` stands and the version didn't move. It's an ordinary return value rather than an exception, and the agent should re-read the row and plan again. To require that a property is _missing_, use the exported `compareAndSetAbsent` marker (`undefined` is too easy to lose while an object is being built or serialized). Here are two agents trying to claim the same unowned account: ```typescript const claim = (ownerId: string) => store.nodes.Account.compareAndSet("acct_1", { expected: { ownerId: compareAndSetAbsent }, patch: { ownerId }, }); console.log(await claim("rep-9"), await claim("rep-4")); ``` ```text true false ``` The first agent gets the account and the second finds out immediately, without a lock or a retry loop. Expected values are checked against the property's current schema, so you can't guard on a value that's already invalid, which is why the repair above guards on `name`, `plan` and `seats` rather than the bad email. The row also has to be valid as a whole after the patch, so patching only `plan` on `acct_legacy_2` fails on `seats` before anything is written. ## Review a batch before it lands A guard works for one row, but a partner feed offering a batch of accounts is more of a review problem, where you want to see what would change before anything does. `planCandidateWriteSet()` takes a serializable batch of nodes and edges, stages it on a throwaway copy, and hands back an ordinary merge plan without ever writing to the target. ```typescript const plan = unwrap( await planCandidateWriteSet({ target: store, makeBackend, options, // two Accounts with the same email are the same account writeSet: { formatVersion: 1, sourceId: "partner-feed", target: await captureCandidateWriteSetTarget(store), nodes: [partnerAccount7, partnerGlobex], edges: [], }, }), ); ``` ```text upserts: 2 accounts still 243 acct_7 name: keep "Account 7", feed says "Account 7 Holdings" acct_7 plan: keep "team", feed says "enterprise" acct_7 seats: keep 12, feed says 40 ``` The plan has one new account and one match on `acct_7`, where the feed disagrees on three properties. The existing values win by default and the disagreements go on the review, so "enterprise, 40 seats" becomes a line item for a person to look at instead of a silent overwrite. If something unrelated writes to the graph before you apply, applying the original plan fails: ```text StaleMergePlanError: The target revision changed after this merge plan was created; the plan was not applied. accounts 243 ``` Re-planning and applying takes the count to 244. The plan is tied to the target's revision rather than only the rows it touches, so any write makes it stale, which is cheap to recover from for a bounded batch. When a person's approval has to sit between planning and applying, [durable review](/blog/durable-merge-branches) handles it. ## One consistent snapshot At the root store, each `describe()` statement and each `validateStore()` page is its own read. TypeGraph catches schema changes between those reads but not data changes, so when a sweep needs one consistent view, 0.64 puts both methods on the transaction context: ```typescript const summary = await store.transaction( async (tx) => { const { statistics } = await tx.describe(); // ...page through tx.validateStore() exactly as above }, { isolationLevel: "repeatable_read", accessMode: "read_only" }, ); ``` ```text { accounts: 244, scanned: 244, violations: [] } ``` The violations are gone because both repairs landed earlier in the run. Use `repeatable_read` or `serializable` and consume every page inside the callback. ## Limits - **`describe()` covers directly addressable declared properties.** It doesn't guess through `$ref`, unions, arrays, or conditionals. `validateStore()` is the authority for those. - **Both are current-state only.** There's no "what did the data look like last month" analysis. - **A guard can only protect what it can name.** You can't guard on an invalid value, so if a person and an agent both repair the same bad field, a guard on its neighbors won't tell them apart. The revision-fenced plan is the alternative there. - **Planning copies the whole target,** so its cost scales with the graph. It's for bounded review workflows, not a hot path. ## Upgrading Writing this post turned up two bugs in 0.68.0 and earlier. `planCandidateWriteSet()` fails on a target that holds a row with undeclared properties, like `acct_legacy_3` above ([#733](https://github.com/nicia-ai/typegraph/issues/733)), and `compareAndSet()` throws a raw Zod error on a kind whose schema has an object-level `.refine()` ([#734](https://github.com/nicia-ai/typegraph/issues/734)). Both are fixed in 0.68.1 ([#735](https://github.com/nicia-ai/typegraph/pull/735), [#737](https://github.com/nicia-ai/typegraph/pull/737)), which the examples here assume. `BulkOperationHookContext["operation"]` now includes `"compareAndSet"`, so an exhaustive `switch` over it needs the new case. If you run mixed versions during a rollout, upgrade every process that writes schema versions to 0.54 before using graph-scoped annotations (also new in 0.54: metadata on the graph itself, carried through extensions and returned by `describe()`). Older writers drop fields they don't recognize. ## Try it - [`describe()` and `validateStore()`](/graph-extensions#population-statistics-and-stored-data-validation) - [`compareAndSet()`](/schemas-stores#compareandsetid-params) and `compareAndSetAbsent` - [`planCandidateWriteSet()`](/graph-merge#constraint-aware-ingestion-branches) - [GitHub](https://github.com/nicia-ai/typegraph) # Agent Memory That Knows Why It Believes Things import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; Say a vuln-feed agent in your overnight security pipeline flags 60 services as shipping a vulnerable library, blocks their deploys, opens fix PRs, and pages the owning teams. In the morning someone notices that the feed had a bad package-name mapping, which pinned a real CVE to the wrong package across a whole class of entries. What should happen to those 60 flags? Deleting all of them is wrong, because some were independently confirmed by a SAST run or an SBOM check and those services really are vulnerable. Leaving them is also wrong, because most were only ever the feed talking, and they're blocking shipments and paging people for nothing. What you need is the blast radius: which of the 60 flags depend only on the feed, and which have other evidence behind them. Most agent memory can't answer that, because "memory" today usually means retrieval: embed everything, fetch the nearest neighbors, and let the model sort it out. That works for recall, but a vector store only keeps what was said. It has no record of _why_ something was concluded, or of what should happen to a conclusion when one of its sources turns out to be wrong. You could live with that when an agent was a single chat transcript, but not once a fleet of agents writes into shared memory and acts on it. Over the last few releases I built the layer that answers the question: a second temporal axis (0.33), provenance-backed retraction on top of it (0.34), and the hardening that makes both hold up under real load (0.39 and 0.40). I'm really happy with how this one came out. ## The blast radius, as a query With the provenance layer, the support structure is right there in the graph: ```plaintext VulnFeed ──▶ Vulnerable(svc-14) ──▶ BlockDeploy(svc-14) SASTRun ──▶ Vulnerable(svc-14) two sources, survives VulnFeed ──▶ Vulnerable(svc-22) ──▶ BlockDeploy(svc-22) feed only, dies ``` Retract the feed and `Vulnerable(svc-22)` loses its only source, so it goes non-current. That takes away the premise of `BlockDeploy(svc-22)`, which goes non-current too, and the deploy unblocks. `Vulnerable(svc-14)` survives because the SAST run still backs it, so that block stands. ```typescript const report = await provenance.retract({ kind: "VulnFeed", id: badMappingSnapshotId, }); // report.died: // flags and deploy blocks that had only the feed behind them // // report.survivedVia: // flags still confirmed by SAST or the SBOM check ``` Run that across all 60 services and `report` is your cleanup list: `died` tells you which deploys to unblock and which PRs to close, and `survivedVia` tells you which flags are still real. I wouldn't trust anyone to work that out by hand at 7am. The source doesn't have to be a feed; it can be another agent's run. If a triage agent combined an overnight feed-ingest run, a SAST scan, and an SBOM rebuild into block-and-page decisions, and the feed-ingest run turns out to have trusted bad data, you retract that run, not the whole fleet's memory: ```typescript const report = await provenance.retract({ kind: "AgentRun", id: overnightFeedRunId, }); ``` Flags raised only by that run go non-current, while flags another run also supports survive, so one agent's mistake gets cleaned up without disturbing what the others established. ## How retraction works The `@nicia-ai/typegraph/provenance` subpath maps your existing graph kinds onto four roles (sources, justifications, facts, and the edges between them) and gives you a `retract` that recomputes support instead of deleting blindly: ```typescript import { createRetractionCapability } from "@nicia-ai/typegraph/provenance"; const provenance = createRetractionCapability(store, { source: { kinds: ["ScannerSource", "VendorSource"] }, justification: { kind: "Justification" }, fact: { kinds: ["Vulnerability", "DeployDecision"] }, premiseOf: { kind: "premiseOf" }, derives: { kind: "derives" }, }); ``` A fact stays believed while at least one of its justifications has all its premises still supported. Premises bottom out at sources. Retract a source and every justification that leaned on it stops counting; a fact loses currency only when it runs out of surviving justifications. Two properties make this safe to use. Retraction is **scoped**, touching only facts reachable from the sources that flipped, and it's **reversible**: it changes whether a fact is believed without deleting anything, and the fact's edges stay put, so `unRetract` is an exact inverse of `retract`. None of this is new theory. The storage follows the JTMS shape from Doyle's 1979 _A Truth Maintenance System_: AND-justifications over premises, sources as the base case, a fact believed only if some justification has all its premises supported. The question `retract` answers (which facts keep support after a source drops out) is the ATMS question from de Kleer's 1986 work, and all of it runs on ordinary SQL. It covers the well-founded, monotonic part of classic truth maintenance rather than all of it. The new part is where it lives: your agent keeps writing ordinary graph data and gets retraction semantics without moving to a dedicated reasoning engine. ## Replaying what the agent believed Retraction is much more useful with history. You want "why did the agent block svc-22 at 2am, and why doesn't it anymore?" to be a query rather than a dig through logs, and recorded time is what makes that possible. TypeGraph already had valid time: when a fact was true in the world, set with `validFrom` / `validTo` and read with `store.asOf(T)`. 0.33 adds the second axis, **recorded time**: when TypeGraph captured a fact, what the system knew as of a commit. It's the SQL:2011 `FOR SYSTEM_TIME` axis, or Datomic's system time. Turn it on per store: ```typescript const store = createStore(graph, backend, { history: true }); ``` Writes through that store are captured with a per-graph, monotonic commit anchor. `store.asOfRecorded(T)` gives you a read-only view of the graph as of that anchor, and `store.recordedNow()` hands you the current one. The two axes compose: ```typescript store.asOf(validTime).asOfRecorded(recordedTime); ``` Retraction is a normal write under `history: true`, so the whole before and after replays: ```typescript const before = await store.recordedNow(); await provenance.retract({ kind: "VulnFeed", id: badMappingSnapshotId, }); const after = await store.recordedNow(); await store.asOfRecorded(before).nodes.BlockDeploy.getById("svc-22"); // current await store.asOfRecorded(after).nodes.BlockDeploy.getById("svc-22"); // non-current ``` One naming note, since it trips up people coming from other systems: TypeGraph's `asOf` is valid time, the reverse of SQL:2011's `FOR SYSTEM_TIME AS OF` and Datomic's `d/as-of`, where a bare "as of" means system time. Valid-time reads are the common case here, so they got the short name. ## Why plain SQL, and what it costs There's no portable, engine-native system versioning across Postgres and SQLite. Postgres needs an extension or an application-level pattern, and SQLite has nothing. So TypeGraph stores history in its own tables and reconstructs point-in-time views in the query compiler. One implementation runs on both backends, so the same memory model works from a solo agent's local SQLite file up to a fleet writing into shared Postgres. That isn't free: - **Only TypeGraph-managed writes are captured.** Raw `tx.sql` is disabled on a history store, since it would bypass capture. This is an audit layer for graph writes, not database-level CDC. - **No backfill.** Enable history on a fresh graph. An entity that already existed is first recorded the next time it's written. - **Recorded reads are an audit and replay tool, not a hot path.** They reconstruct from history, so they're slower than current-state reads, and broad reads like `find`, `search`, and vector predicates aren't available at a past anchor because those indexes only reflect the present. - **Writes cost more.** Roughly 2.5–6× an uncaptured write when each write is its own transaction, dropping to about 1–1.5× when writes are batched. ## Hardening it: 0.39 and 0.40 Once real consumers started reading this axis, two gaps showed up. **Commits in the same millisecond.** 0.33 ordered commits by wall-clock timestamp alone, which breaks exactly when writes get fast. Five back-to-back writes: ```text commit 0 anchor: r1:0000000000000001:2026-07-21T17:59:03.914Z commit 1 anchor: r1:0000000000000002:2026-07-21T17:59:03.914Z commit 2 anchor: r1:0000000000000003:2026-07-21T17:59:03.915Z commit 3 anchor: r1:0000000000000004:2026-07-21T17:59:03.915Z commit 4 anchor: r1:0000000000000005:2026-07-21T17:59:03.915Z ``` Commits 0 and 1 share `17:59:03.914Z`. With a timestamp-only anchor, "the graph right after commit 0 but before commit 1" doesn't have an answer. 0.40 anchors are `r1::`: the revision is a strict per-graph counter, so every commit gets its own position no matter how many share a millisecond. The timestamp is still there for display, and it never moves backwards, even if the system clock does. A raw ISO string no longer type-checks as an anchor, on purpose; get anchors from `recordedNow()`. Deployments that adopted the earlier format run `migrateLegacyRecordedTime()` once before opening the upgraded store. **Walking a whole snapshot.** 0.39 adds `scan()` to recorded collections: bounded, deterministic pages (up to 1,000 per call) ordered by id, with a cursor bound to the exact view it came from: ```typescript const view = store.asOfRecorded(anchor); const page1 = await view.nodes.Item.scan({ limit: 2 }); const page2 = await view.nodes.Item.scan({ limit: 2, after: page1.nextCursor }); ``` Together those make "every row as of commit N, in order, one page at a time" a supported operation for a replication tap, an audit export, or a search index rebuild pinned to a point in history. ## Why agent memory needs this A fact stored without provenance is organizational hearsay: the memory can't tell you why it believes something or what should change when a source goes bad. A single chat agent can get away with that because a lot of inconsistency hides inside one transcript, but agents that share memory will read stale feeds, trust bad mappings, and inherit each other's wrong assumptions. Memory that gates deploys, or maintains any long-lived model of a company, needs to answer three questions: what do we believe, why do we believe it, and what changes if this source turns out to be wrong. This layer is built to answer those, and it runs on the SQL database you already have. ## Try it - [Provenance and Retraction](/provenance) and the [Provenance Retraction example](/examples/provenance-retraction/) - [Bitemporal Time Travel example](/examples/bitemporal-time-travel/) - [Logical revision and physical time](/queries/temporal#logical-revision-and-physical-time): the anchor format and diagonal bitemporal reads - [Migrating preview recorded time](/schema-management#migrating-preview-recorded-time) - [GitHub](https://github.com/nicia-ai/typegraph) If you've mapped JTMS/ATMS-style truth maintenance onto relational storage elsewhere, or know of prior art that does, I'd like to hear about it. There's plenty written on truth maintenance systems and not much on doing it directly on ordinary SQL tables. # TypeGraph 0.35: Faster Almost Everywhere import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro"; A 2M-row bulk load into SQLite was still running after 4.5 hours with no sign of finishing, and the cause turned out to be a statistics refresh. After a large batch, `bulkCreate` and `bulkInsert` automatically run SQLite's `ANALYZE` so the query planner has fresh numbers, but that `ANALYZE` was bare and unscoped: it scanned every table in the database file (not just TypeGraph's), with no limit, after every big batch. If you stream a load through repeated `bulkInsert()` calls, each refresh costs more than the last and the total goes quadratic. It's now scoped to TypeGraph's own tables and bounded with `PRAGMA analysis_limit`, the way the Postgres version already behaved. A 100k-row reproduction of the same shape finishes in about 8 seconds, and the last batch costs only about twice what the first one did. 0.35 has about 35 fixes like that one, and below are the ones that matter most, with numbers. ## Bulk writes These four changes stack on top of each other for anything that writes a lot of rows: - **SQLite pragmas at open.** `createLocalSqliteBackend` now turns on `journal_mode=WAL`, `synchronous=NORMAL`, and a busy timeout by default. better-sqlite3's own defaults pay a full fsync per write in rollback-journal mode. Single-operation writes on file-backed databases are roughly **5x faster**. - **Real bind-parameter limits.** The backend used to assume SQLite's old 999-parameter limit on every driver. It now asks the driver (32,766 for better-sqlite3, 100 for Cloudflare D1), so batches on better-sqlite3 use **~33x fewer statements**: 111-row chunks became 3,640-row chunks. - **`bulkCreate` / `bulkInsert` batched end to end.** Existence checks, uniqueness checks, and fulltext/embedding writes each happen once per batch instead of once per row: **~1,600 → ~4,100 rows/s (~2.6x)**. - **`importGraph` batched the same way: ~26k → ~96k entities/s (~4x).** The default `batchSize` also went from 100 to 1,000, and now actually applies. A schema-parsing gap meant the old default was silently ignored. That change alone took a 20k-node, 5k-edge Postgres import from 1,515ms to 781ms. There's also a new `walAutocheckpointPages` option for tuning WAL checkpoints during heavy loads. On its own it cut a 2M-row load's time by more than half at the largest scale tested. ## The regression, and fixing it properly This one's on me. 0.34 fixed a real correctness bug where "current" reads compared against the _database's_ clock instead of the _application's_, so when the two clocks drifted (app and database on separate hosts, which is normal), a row you'd just created could be invisible to the very next read. The fix was to bind the read instant fresh every time a query compiled, which was correct but had two problems I didn't catch until this release. **It was slow.** Every `execute()` recompiled the query from scratch, including the `.prepare()`-once, `.execute()`-many pattern that exists specifically to avoid that. A repeated point query cost about 47µs. **It was also, briefly, worse than the bug it fixed.** The query builders cached their compiled SQL text across calls, so a prepared query froze "now" at the moment it first compiled. Every row created after that had a later `valid_from` than the frozen instant, and stayed invisible to that query forever. It reproduced in a single process on the very next insert after preparing a query, whereas the clock-skew bug at least needed two hosts. 0.35 fixes both the same way: the SQL text is cached, and only the "now" instant is re-bound as a parameter on each call. The repeated point query went from **47µs to 2.4µs (~20x)**, and a row created after `.prepare()` shows up on the next `.execute()`. I'm not going to bury that in a changelog line. If you're on 0.34 and use `.prepare()`, upgrade. ## Point operations and traversals - **CRUD statements reuse the prepared-statement cache.** On synchronous drivers, Drizzle's `db.all()` / `db.run()` re-prepared every statement. They now go through the same cached path the query engine uses. Single creates: **~18.3k → ~28.8k ops/s (~1.6x)**. - **Cascade deletes remove edges in batches** instead of one statement per edge. A 50-edge cascade on local Postgres: **24.4ms → 3.6ms**. - **`degree()` can use the index again.** The direction filter compiled to a shape neither edge index could seek. Now it matches the index prefix: **0.30ms → 0.06ms** on Postgres 18, and on Postgres 17 and earlier it no longer falls back to scanning the whole partition. - **Subgraph extraction is ~4x faster on Postgres** (322ms → 82ms on a depth-3 stress shape). The recursive traversal runs once instead of twice, and the resulting ids go in as a single array parameter. ## Search - **Hybrid search is one SQL statement** instead of two searches plus fusion in JavaScript, with the candidate filter computed once and shared. Filtered hybrid search at 5k documents: **26.5ms → 17.1ms**. - **Postgres fulltext can use its GIN index.** The query referenced the language from a per-row column, which kept the planner off the index. It's now a constant: **12.9ms → 2.3ms** at 5,000 documents. - **Exact `.similarTo()` on SQLite uses sqlite-vec's KNN.** It's brute force in C, so results are identical, just faster: **489ms → 124ms** for top-10 over 50k 384-dimension embeddings. - **Approximate vector search on Postgres uses the ANN index now.** A stray `DISTINCT` kept the planner off the ordered index scan, and the inline path wasn't applying the same pgvector tuning as the search API. Together: **174ms → 2.1ms**, at 0.995 recall. That last area also had a correctness bug. Exact `.similarTo()` was **quietly approximate** whenever a matching ANN index existed, because pgvector will happily answer an exact-looking `ORDER BY ... LIMIT k` from an HNSW or IVFFlat index. Under a selective filter at 50k documents, measured recall dropped as low as 0.000, which means the results were simply wrong. The exact path now forces a true scan regardless of which indexes exist. ## Found by benchmarking against real graph databases Two of these fixes came from running the LDBC Social Network Benchmark against Neo4j and LadybugDB, not from profiling TypeGraph on its own. Edge `bulkCreate` / `bulkInsert` had an N+1 endpoint check that made each batch slower as the graph grew (~90ms → ~630ms per batch). And the default edge traversal indexes were missing columns a join needed to stay index-only, which didn't show up until the table outgrew the page cache and then hit a multi-second latency cliff. The [full benchmark writeup](/blog/benchmarking-typegraph-neo4j-ladybugdb) covers where TypeGraph wins and where it doesn't. **If you're upgrading an existing database**, note that the wider edge indexes only appear on fresh databases. `CREATE INDEX IF NOT EXISTS` does nothing when an index with that name exists, even with a different column list, so an upgraded deployment keeps the narrow index until you rebuild it. [Performance → Indexes](/performance/indexes) has the exact `DROP` / `CREATE INDEX CONCURRENTLY` steps for both backends. ## Also in 0.35 - **`.aggregate({...}).orderBy(key, direction?)`**, so "top N groups by count" no longer means fetching every group and sorting in JavaScript. - **Breaking, ontology:** `implies(edgeA, edgeB)` now checks that the two edges' endpoint kinds are compatible. This also runs when a persisted schema loads, so check saved schemas before rolling out, not just source. ## Try it - [Performance overview](/performance/overview) and [Indexes](/performance/indexes) - [Changelog](/changelog#0350): the full 0.35.0 entry - [GitHub](https://github.com/nicia-ai/typegraph) # Changelog > Release notes for @nicia-ai/typegraph Release notes for [`@nicia-ai/typegraph`](https://www.npmjs.com/package/@nicia-ai/typegraph). Generated from `packages/typegraph/CHANGELOG.md` on every build. ## 0.73.1 ### Highlights `updateWhere()` and `compareAndSet()` now preserve stored values for defaulted properties omitted from a patch, matching `update()`. Previously, changing one property could silently reset unrelated properties to their schema defaults on every matched row. Defaults are also no longer evaluated for omitted compare-and-set expectations. ### Upgrade notes - After upgrading, remove workarounds that restate every defaulted property in `updateWhere()` or `compareAndSet()` patches. Omitted properties retain their stored values; to reset a property, supply the desired value explicitly. - Check rows previously changed by `updateWhere()` or `compareAndSet()` on node kinds with defaulted properties, and restore any unintended resets from application history or backups. Upgrading prevents future resets but does not recover overwritten values. ### Patch Changes - [#783](https://github.com/nicia-ai/typegraph/pull/783) [`d5560e5`](https://github.com/nicia-ai/typegraph/commit/d5560e5dd220e6682799e94e55c26927d7638310) Thanks [@pdlug](https://github.com/pdlug)! - Preserve omitted defaulted properties in `updateWhere()` and `compareAndSet()` patches. Validate only supplied patch and expected-state fields so defaults cannot silently overwrite stored values or run for omitted expectations. ## 0.73.0 ### Highlights The PostgreSQL working-copy manager now covers the whole branch lifecycle. `makeBackend` hands `branch`, `ingestionBranch`, candidate write-set planning, and evolution previews an empty, schema-mutable allocation recorded in the manager's ledger, so they no longer need a hand-rolled backend factory: closing the backend drops it, and `listUnsealedAllocations` and `abortAllocation` recover it after a crash. Durable working copies can now apply a host mutation and commit its operation evidence in one transaction (`operations: { graph, apply }`), with idempotent replay, commit-ordered scans, delivery marking, and a destroy fence that refuses while evidence is undelivered. Every allocation lives in one recorded schema, so a pooled connection whose `search_path` leads elsewhere can no longer scatter tables that cleanup never finds, and removal either drops everything the allocation owns or refuses with a `BranchError` naming where its relations went. Operators can list the graphs in a database with `listGraphIds` and count one graph's rows per relation with `inspectGraphStorage`, on SQLite and PostgreSQL alike, without depending on TypeGraph's physical tables. `inspectGraphStorage` reports whether its counts came from one snapshot. The set of graph-scoped relations now has a single owner shared by `store.clear()`, namespace forks, working-copy cloning, and these reads, which fixed a few places that disagreed about it. On PostgreSQL, a new byte-ordered `graph_id` index keeps a page of `listGraphIds` at about 3 ms with 20,000 graphs in the database, and adds 1 to 2% to writes on the tables it covers. Writes through an ephemeral PostgreSQL working copy of a history-enabled store work again; since 0.72.0 every create, update, or delete on such a copy failed with a `ConfigurationError`. ### Upgrade notes - The base-schema marker advances from 4 to 5. Open each database once with `createStoreWithSchema` under a DDL-capable role, or apply the version-5 migration if you manage DDL externally, before starting workers that use `createVerifiedStore`, `assertSchemaCurrent`, or the DML-only graph-template APIs; until then they throw `BaseSchemaMigrationError`. Earlier releases refuse a database stamped 5, so do not roll back once any process has adopted it. - On PostgreSQL, if `nodes` or `edges` see continuous writes, create the three `_graph_id_bytes_idx` indexes with `CREATE INDEX CONCURRENTLY IF NOT EXISTS` (as `backend-setup` shows) before upgrading. Otherwise the version-5 adoption builds them inline at boot, blocking writes to each table while it builds, even when `systemIndexes: "skip"` is set. - Run the working-copy manager's `control` and every `connect` session as the same database role. A different role, including one that is merely a member of `control`'s role, is now refused with `WORKING_COPY_ROLE_MISMATCH` on every path. Drop by hand any tables that allocations created under a different role left behind. - In `connect`, build the backend's tables with `createPostgresTables(names)` from the object `connect` receives, not a copy of it, use a driver that supports interactive transactions (not `drizzle-orm/neon-http`), and keep the allocation's schema on the connection's `search_path`. - If you wrap the `control` or `connect` backend, forward the transaction `isolationLevel` option. Otherwise destroy, `abortAllocation`, closing ephemeral or `makeBackend` copies, and durable operations fail with `WORKING_COPY_ISOLATION_UNSUPPORTED` whenever the session's default isolation is not READ COMMITTED. - Upgrade every process that shares a working-copy ledger before any of them creates or destroys a durable allocation. An earlier manager destroys allocations without checking their operation evidence and leaves the evidence table behind; the #777 entry below describes how to recover it. ### Minor Changes - [#778](https://github.com/nicia-ai/typegraph/pull/778) [`de9054b`](https://github.com/nicia-ai/typegraph/commit/de9054bd86b13fc01356802203ed1717bc0369eb) Thanks [@pdlug](https://github.com/pdlug)! - Add `listGraphIds(backend, { prefix, after, limit })` and `inspectGraphStorage(store)` so operators can list the graphs in a database and count one graph's rows per relation, including per-field vector tables, without depending on TypeGraph's physical table layout. Both behave identically on SQLite and PostgreSQL: ids page in byte order regardless of database collation, only graphs a default `store.clear()` would empty are listed (a cleared graph stops being listed even though its contribution markers are kept), each page walks graph ids by index seek instead of scanning every row, bounded by the cursor, prefix and page size (about 3 ms a page on PostgreSQL at 20,000 graphs, against about 400 ms for the equivalent walk over the database collation), prefixes match as case-sensitive literal text, and relations that were never provisioned count as empty. `inspectGraphStorage` reports whether its counts share one snapshot in a new `consistency` field, `"snapshot"` or `"per-statement"`. It requests a read-only `repeatable read` transaction, then reads the isolation level the counting session actually ran under inside its first count statement instead of trusting the request: SQLite transactions are always one snapshot, PostgreSQL reports `"snapshot"` only when the session was observed at `repeatable read` or `serializable`, and everything else reports `"per-statement"`: a backend without interactive transactions, a session observed at `read committed` (which is what a wrapper that drops the isolation option gets under a `read committed` default; under a `repeatable read` default the same wrapper still reports `"snapshot"`, because the level is observed rather than requested), or a backend that declares no session isolation read and so cannot be observed. With `"per-statement"`, a write between two counts can make the result describe a state that never existed. A graph with fewer than two count statements reports `"snapshot"`. The evidence proves the isolation of the session that ran the first count, so a backend wrapper that violates the transaction contract by handing the root pool through as its transaction backend can still run later counts on other sessions. It does not refuse on a weaker level. The set of graph-scoped relations now has one owner. `store.clear()`, namespace forks, PostgreSQL working-copy clone policies, the provenance sidecar occupancy probe and the new reads all consume the same declaration, and a ratchet test fails when a bundled table gains a `graph_id` column without being classified. `SqlTableNames` now also carries the `indexMaterializations`, `contributionMaterializations`, `kindRemovals` and `reconciliationMarkers` names (optional, defaulting to the bundled names), and the SQLite backend reports them in `backend.tableNames` as PostgreSQL already did. Inconsistencies the shared declaration exposed are fixed. `store.clear()` now removes the graph's revision-origin row whichever store clears it; before, a store that did not mint origin-namespaced tokens left a row that a revision-tracking store had minted, so the graph still read as occupied. `forkGraphNamespace` and `prepareNamespaceForkTarget` no longer fail with a raw missing-relation error on a PostgreSQL backend created with `fulltext: false`, and refuse a source and target whose fulltext storage differs with a `BranchError`. The provenance sidecar occupancy probe now counts revision-journal rows as occupancy, so a graph id whose only rows are journal entries is no longer treated as free. PostgreSQL gains base-schema version 5: a byte-ordered (`COLLATE "C"`) single-column `graph_id` index on `nodes`, `edges` and `schema_versions`, named `
_graph_id_bytes_idx`, which is what lets the listing bound its walk (SQLite already keeps text indexes in byte order, so its step only advances the marker). The privileged open adopts it on existing databases with a plain `CREATE INDEX IF NOT EXISTS` per relation, which blocks writes to that relation while it builds (about 0.1 second per million `nodes` rows). This happens inline at boot even when `systemIndexes: "skip"` is set, because that option does not defer base-schema adoption; for a large deployment, build the three indexes with `CREATE INDEX CONCURRENTLY IF NOT EXISTS` beforehand, as `backend-setup` describes. The index adds 1 to 2% to writes on the relations it lands on. The index names are reserved like the other system indexes, and a database whose indexes are absent still lists correct pages by de-duplicating the anchor relations instead of walking them. ## Breaking - The base-schema marker advances from 4 to 5, so a database stamped 4 is stale until it is adopted: a store opened with `createStoreWithSchema` adopts it automatically under a DDL-capable role, while `createVerifiedStore`, `assertSchemaCurrent` and the DML-only graph-template APIs throw `BaseSchemaMigrationError` (`BASE_SCHEMA_MIGRATION_REQUIRED`) until it is. Open once with `createStoreWithSchema` before DML-only workers start, or apply the version-5 migration for externally managed DDL. Older library versions refuse the newer marker, so do not roll back to a release that predates version 5 once any process has adopted it. - On PostgreSQL, build the three `graph_id_bytes_idx` indexes with `CREATE INDEX CONCURRENTLY IF NOT EXISTS` before upgrading a database whose `nodes` or `edges` tables see continuous writes, so the adoption step finds them in place and does not block writers. - [#777](https://github.com/nicia-ai/typegraph/pull/777) [`38c3a21`](https://github.com/nicia-ai/typegraph/commit/38c3a2164dd9d8e27510e7ebb3b0baa2147deea2) Thanks [@pdlug](https://github.com/pdlug)! - Implement the durable operation capability on the bundled PostgreSQL working-copy manager. Passing `operations: { graph, apply }` to `createPostgresWorkingCopyManager` makes `durable.operations` apply the host's opaque mutation inside the transaction that commits its evidence, with idempotent replay, digest conflicts, commit-ordered scan cursors, delivery marking, and a destroy fence that refuses with `DurableEvidenceUndeliveredError` while undelivered evidence remains. Evidence coordinates carry `base` and, when the working copy tracks history or revisions, the engine `revision`. Each durable allocation owns an evidence table, created in the allocation's schema and removed with the rest of the allocation (the destroy fence reads it only there, follows the manager's removal rule for allocations whose relations were moved or dropped, and runs only when the evidence relation still exists, so an allocation whose evidence relation is gone is removed like any other relation that exists nowhere), and the working-copy ledger gains an `operation_evidence` column. `operate`, `markDelivered`, destroy, and `abortAllocation` serialize per allocation on a transaction-scoped advisory lock keyed on the allocation id, in a namespace of its own, so the evidence sequence is commit order. Each requests READ COMMITTED and observes the isolation its session actually runs at; any other level is refused with a `ConfigurationError` (`details.code` `WORKING_COPY_ISOLATION_UNSUPPORTED`) before anything is read or written, so a `control` or `connect` wrapper must forward the transaction `isolationLevel` option. Every member checks `operations.graph` against the sealed allocation's graph id and definition hash before opening a connection. `apply` runs under the graph-wide write lock, which blocks tracked writes to the source graph and its sibling working copies for its duration. Allocations created by earlier releases report `unsupported` (`evidenceStore`) for `operate` and no evidence for the read members, answered from the ledger row alone: one read-only ledger `SELECT` through `control`, with no DDL, no lock, and no `connect` call. ### Upgrade notes - Upgrade every process that shares a working-copy ledger before any of them creates or destroys a durable allocation. Only managers on this version take the allocation lock and honor the evidence fence. A manager from an earlier release destroys an allocation without consulting its evidence, so undelivered evidence is lost with it, and it leaves the evidence table behind. - If an earlier manager destroyed a durable allocation created by this version, a later allocation with the same id is refused because `op_evidence` exists without a ledger row; the refusal names the table. Read its undelivered rows (`WHERE NOT delivered`) and deliver them, then `DROP TABLE` the named relation and retry. TypeGraph does not drop it, because it may hold the only copy of undelivered evidence. - A `control` or `connect` backend wrapper that drops the transaction `isolationLevel` option now fails destroy, `abortAllocation`, and durable operations with `WORKING_COPY_ISOLATION_UNSUPPORTED` when the session's default is not READ COMMITTED. Forward the option. - The same check now runs wherever the manager drops an allocation, because dropping takes the allocation lock: closing an ephemeral working-copy store, closing a `makeBackend` backend, and cleanup after a failed allocation. The first two throw the refusal from `close()`. The cleanup swallows it so the allocation's original failure reaches the caller, and the allocation is left behind. List such orphans with `listUnsealedAllocations` and remove them with `abortAllocation` once `control` forwards the option. - [#776](https://github.com/nicia-ai/typegraph/pull/776) [`e4a9419`](https://github.com/nicia-ai/typegraph/commit/e4a941998f9a311fda2189205eda92431b04a821) Thanks [@pdlug](https://github.com/pdlug)! - Add `makeBackend` to the PostgreSQL working-copy manager so `branch`, `ingestionBranch`, `planCandidateWriteSet`, `planCandidateWriteSetReview`, `branchForEvolution`, and `planCandidateWriteSetForEvolution` no longer need a hand-rolled PostgreSQL backend factory. Each call records a fresh, empty, schema-mutable allocation in the manager's ledger; closing the backend drops it, it is listed by `listUnsealedAllocations()` while live, and `abortAllocation()` removes it after a crash. Dropping any allocation now also removes the vector tables created under its reserved physical prefix after the ledger manifest was written, and declared graph indexes on a `makeBackend` allocation are scoped to it so they cannot collide with the source's or another allocation's index names. `connect` always receives the allocation vector strategy for `makeBackend` and must bind it or disable vector support. Backends derived from the returned one with `deriveBackend` inherit that index scoping. Every allocation now lives in one explicit schema, the `control` session's current schema when it is allocated, recorded in a new `schema_name` ledger column (added to an existing ledger on first use; a row written by 0.72.0 carries no schema and is resolved through the removing session). Provisioning fixes its transaction to that schema, the connected backend runs the DDL it issues lazily with that schema leading its search path, the allocation's pgvector strategy creates and drops its tables and indexes schema-qualified, and clone inserts name their target relations through it, so a pooled `connect` connection whose `search_path` leads with another schema can no longer create relations that removal never finds. Removal searches the catalog across every schema for relations named with the allocation's reserved prefixes, drops those in the recorded schema schema-qualified, and deletes the ledger row in the same transaction, so a failed drop keeps the row and the allocation stays listed and recoverable. When such relations exist in another schema, removal refuses with a `BranchError` that keeps the row and carries `allocationId`, the recorded `schema`, the schemas found in `foundIn`, and a recovery `suggestion`: a renamed schema or moved tables read "not in its schema" (move them back or correct the row's `schema_name`); a stale copy left in another schema, such as a backup or restore schema, reads "also has relations" and blocks removal until that copy is dropped, and is only called a stale copy when every relation found elsewhere has a same-named relation in the recorded schema; a partial move (some relations moved, others stayed) reads "is split across schemas" and suggests dropping nothing, because the relations elsewhere may be the only copy; when they exist nowhere (the tables were dropped entirely) there is nothing to recover and the row is removed, so a crashed owner's allocation cannot stay listed forever. The same rule applies to a legacy row that records no schema. A backend built over a caller's own transaction never has that transaction's `search_path` rewritten: its lazy DDL (including extension DDL on the non-lock fence path), schema writes, and schema adoption run only when the session's current schema is the allocation's, and are refused with a `ConfigurationError` (`ALLOCATION_SCHEMA_SESSION_MISMATCH`) otherwise; extension installation under a lock fence makes no session check, running in its own transaction on a pooled backend and as a savepoint inside the caller's transaction on a backend built over one. The connected backend's catalog probes (`tablesExist`, `indexStates`, `columnTypes`) read the allocation's schema rather than the session's current schema, so a history-enabled allocation opens through a connection whose `search_path` leads with another schema. Trusted import drops and recreates the allocation's secondary indexes by the schema the catalog found them in, so a same-named index earlier on the connection's `search_path` is left alone. **Breaking:** the PostgreSQL working-copy manager now supports one deployment shape, in which `control` and every `connect` session run as the same database role. The `ephemeral` and `durable` clone paths and durable reopen, which shipped in 0.72.0 without this check, now compare `current_user` on the two sessions and refuse a difference with a `ConfigurationError` (`details.code` `WORKING_COPY_ROLE_MISMATCH`). The comparison is by role name, so a `connect` role that is merely a member of `control`'s role, which worked before, is now refused too. `makeBackend` applies the same rule and refuses before it writes the ledger or any DDL. The clone paths and reopen run `connect` after the allocation tables exist, as before, so they refuse right after `connect` returns, before any clone or Store write: a refused clone removes the allocation it just created, and a refused reopen leaves the sealed allocation untouched. **Breaking:** because the schema reaches a connected backend through the names `connect` receives, build the backend's tables with `createPostgresTables(names)` from that object, not a copy: a connection built over a copy is refused with a `BranchError` on every path (and a refused clone removes its allocation), and a `connect` driver that cannot hold an interactive transaction (`drizzle-orm/neon-http`) is refused with a `ConfigurationError` (`ALLOCATION_SCHEMA_REQUIRES_INTERACTIVE_TRANSACTIONS`). The connection's `search_path` must still include the allocation schema so it can resolve the allocation's tables. **Migration:** run `control` and every `connect` session as the same database role, which needs `CREATE` on the schema. The reason is ownership: a connected Store creates tables and indexes that only their owner can drop, and `control` removes every allocation. If you connect as a different role today, allocations created that way can leave tables behind on close and `abortAllocation`; after switching roles, drop any such leftovers by hand. ### Patch Changes - [#780](https://github.com/nicia-ai/typegraph/pull/780) [`e0be095`](https://github.com/nicia-ai/typegraph/commit/e0be095310f6c82f96bb36a17a907feaed07b7f8) Thanks [@pdlug](https://github.com/pdlug)! - Fix writes through an ephemeral PostgreSQL working copy of a history-enabled store. The copy's store was built over the allocation store's already capture-wrapped backend, so recorded capture wrapped twice and every create, update, or delete failed with a `ConfigurationError` from the raw-write guard. The ephemeral store is now built over the unwrapped owned backend, and its writes are captured in the copy's own recorded history. ## 0.72.0 ### Highlights PostgreSQL graphs using bundled table, tsvector, and pgvector storage can now use table-backed working copies. Each allocation owns its tables, indexes, and vector sidecars under a recovery ledger; graph-scoped cloning, durable reopen, and cleanup keep copies isolated from the source. Copies use a fixed schema, so migrate the source before allocating one when schema changes are needed. On revision-tracked graphs, candidate merge planning stages the affected rows and their required identity, ontology, cardinality, and durable edge-identity dependencies instead of cloning the whole target. Candidate-scoped durable review is opt-in. Stores without a revision fence and custom backends missing the keyed reads needed for safe staging retain the complete-clone path. Bulk node creation now reports a generated ID collision as `ValidationError` with `ENTITY_ALREADY_EXISTS_CODE` across the bundled SQLite and PostgreSQL drivers, including atomic batches and the portable fallback. ### Upgrade notes - If a custom PostgreSQL table contribution participates in table-backed working copies, declare its `workingCopyClonePolicy`. Use `graphRows` or `graphDocument` only for graph-scoped content; use `freshSeed` or `rebuildAfterClone` for installation or physical status. An absent or unsupported policy now refuses allocation. - If you handle generated ID collisions from `bulkCreate` or `bulkInsert`, branch on `ValidationError` with `ENTITY_ALREADY_EXISTS_CODE` rather than a driver error or `DatabaseOperationError`. ### Minor Changes - [#759](https://github.com/nicia-ai/typegraph/pull/759) [`d3e74f8`](https://github.com/nicia-ai/typegraph/commit/d3e74f8aff5948038351539c3479ed3b3339989f) Thanks [@pdlug](https://github.com/pdlug)! - Plan candidate write sets on revision-tracked graphs, including Operational Identity and ontology graphs, from bounded candidate dependencies instead of cloning the complete target. Unsupported custom backend reads and candidate owners excluded from the clone projection retain full clone staging. Add opt-in candidate-scoped durable review evidence while retaining the existing whole-graph review default. - [#757](https://github.com/nicia-ai/typegraph/pull/757) [`afe153b`](https://github.com/nicia-ai/typegraph/commit/afe153bd720c56ca17c67b527a4d981e8aefab66) Thanks [@pdlug](https://github.com/pdlug)! - Export `generatePostgresDropSQL()` for cleaning up isolated PostgreSQL table sets from the same contribution inventory used by installation DDL. Quote custom PostgreSQL table and index names consistently in generated DDL. - [#762](https://github.com/nicia-ai/typegraph/pull/762) [`d793eff`](https://github.com/nicia-ai/typegraph/commit/d793efff201b78f2aff8fcf0b6be45b6ad9592a8) Thanks [@pdlug](https://github.com/pdlug)! - Support graph-declared PostgreSQL indexes in table-backed working copies with stable allocation-scoped physical names. B-tree, GIN, and trigram indexes retain their logical declaration names and schema hashes while copy allocation, durable reopen, retry, and cleanup use isolated physical indexes. - [#761](https://github.com/nicia-ai/typegraph/pull/761) [`c6130a1`](https://github.com/nicia-ai/typegraph/commit/c6130a1dd05b3a0b19c13623b494069d1792527f) Thanks [@pdlug](https://github.com/pdlug)! - Add a PostgreSQL table-backed working-copy manager for graphs using bundled table and tsvector storage. It owns ephemeral and durable allocation, a persistent recovery ledger, origin-attested reopen and destroy, graph-scoped SQL cloning under source locks, and bounded inventory of unsealed allocations. Inventory rows may still be active, so callers confirm ownership before explicitly aborting one. Copies have a fixed schema and refuse evolution before mutation. Custom fulltext strategies remain available through host-level database forks. - [#769](https://github.com/nicia-ai/typegraph/pull/769) [`ac8c781`](https://github.com/nicia-ai/typegraph/commit/ac8c781d80deb3ff51f061d89d15360d8c5753c7) Thanks [@pdlug](https://github.com/pdlug)! - Bound candidate merge planning and opt-in candidate-scoped review for `oneActive` graphs on bundled backends. An active-only source read excludes ended edge history while preserving the claim rule that an open edge counts even when its `validFrom` is in the future. Custom backends without the optional read continue to use complete-clone candidate planning. - [#766](https://github.com/nicia-ai/typegraph/pull/766) [`c589e00`](https://github.com/nicia-ai/typegraph/commit/c589e00ad3b96960c89c58f8012decdedfdbb10a) Thanks [@pdlug](https://github.com/pdlug)! - Bound candidate merge planning on revision-tracked graphs with `one` or `unique` edge cardinality. The transient working copy now includes only cardinality peers for candidate sources or endpoint pairs, so staging preserves full-clone constraint decisions without reading unrelated edges. A `oneActive` graph uses complete-clone staging when its backend lacks the active-only keyed peer read. - [#756](https://github.com/nicia-ai/typegraph/pull/756) [`ae813a5`](https://github.com/nicia-ai/typegraph/commit/ae813a55056c5eec6c72c7add53b59a1833e160a) Thanks [@pdlug](https://github.com/pdlug)! - Read graph rows across declared kinds in keyset pages for merge planning and review, avoiding an empty query for each unused kind. Reuse one row read when the target is both sides of a diff, and skip statistics refresh for disposable ingestion clones. Custom backends continue using the existing per-kind read path. - [#768](https://github.com/nicia-ai/typegraph/pull/768) [`79bfa87`](https://github.com/nicia-ai/typegraph/commit/79bfa876765b84be6feb30e7f88237c7fbf3e0ca) Thanks [@pdlug](https://github.com/pdlug)! - Bound candidate merge planning on revision-tracked ontology graphs by reading live same-id peers across node kinds. This preserves full-clone disjointness and type-reconciliation decisions without scanning unrelated nodes, and extends opt-in candidate-scoped review evidence to ontology graphs. - [#764](https://github.com/nicia-ai/typegraph/pull/764) [`a3c2d9c`](https://github.com/nicia-ai/typegraph/commit/a3c2d9c718dc6f7ccb4dfcaac7136a036aabe83a) Thanks [@pdlug](https://github.com/pdlug)! - PostgreSQL table-backed working copies now isolate pgvector sidecars under each allocation's ledger-reserved physical prefix, preserve their embeddings through clone and reopen, and remove their owned tables during destroy. Allocation claims and initial table/vector provisioning commit atomically. - [#770](https://github.com/nicia-ai/typegraph/pull/770) [`7f69a44`](https://github.com/nicia-ai/typegraph/commit/7f69a4466ee36084f5866d0e3476794e687c1d4e) Thanks [@pdlug](https://github.com/pdlug)! - Add an optional backend read for exact durable edge match-identity owners, including tombstones. Candidate planning uses the bounded read to seed active owners into sparse working copies and falls back to full cloning for custom backends without the capability or when a durable owner is tombstoned. - [#764](https://github.com/nicia-ai/typegraph/pull/764) [`a3c2d9c`](https://github.com/nicia-ai/typegraph/commit/a3c2d9c718dc6f7ccb4dfcaac7136a036aabe83a) Thanks [@pdlug](https://github.com/pdlug)! - Add `createPgvectorStrategy(namespace)` for allocation-scoped pgvector table and index names while preserving the default strategy's existing names. ### Patch Changes - [#774](https://github.com/nicia-ai/typegraph/pull/774) [`14dacc4`](https://github.com/nicia-ai/typegraph/commit/14dacc4a1b057671fab06e56380448692a330c36) Thanks [@pdlug](https://github.com/pdlug)! - Preserve the typed duplicate-ID error when PostgreSQL rejects claim cleanup after a failed bulk node insert. - [#760](https://github.com/nicia-ai/typegraph/pull/760) [`9198fa8`](https://github.com/nicia-ai/typegraph/commit/9198fa8b1c64e256d30665d913bc4ba82842f303) Thanks [@pdlug](https://github.com/pdlug)! - Compare cloned branch identity assertions with the base's current state so assertions ended before the fork do not appear as new branch retractions. - [#775](https://github.com/nicia-ai/typegraph/pull/775) [`6820b04`](https://github.com/nicia-ai/typegraph/commit/6820b04f2f63b10abd0d17daf414369cf4085c47) Thanks [@pdlug](https://github.com/pdlug)! - Require explicit PostgreSQL working-copy clone policies for table contributions so physical status rows cannot be copied by column-shape inference. - [#765](https://github.com/nicia-ai/typegraph/pull/765) [`d28c8cb`](https://github.com/nicia-ai/typegraph/commit/d28c8cba346b988c724e760a79f758d3c1b9a885) Thanks [@pdlug](https://github.com/pdlug)! - Refuse PostgreSQL working-copy allocation or durable reopen when the target connection changes the bundled fulltext strategy. This prevents copied fulltext projections from being exposed through a backend with different storage or disabled fulltext support. - [#763](https://github.com/nicia-ai/typegraph/pull/763) [`857f578`](https://github.com/nicia-ai/typegraph/commit/857f578b402d35c9ac3f929a18f7d78405c838ba) Thanks [@pdlug](https://github.com/pdlug)! - Add opt-in candidate-scoped V2 merge review evidence for identity-enabled graphs. Revalidation expands the retained endpoint and assertion-ID scope again and detects connected identity changes while leaving V1's global baseline as the default. ## 0.71.1 ### Patch Changes - [#752](https://github.com/nicia-ai/typegraph/pull/752) [`2d40378`](https://github.com/nicia-ai/typegraph/commit/2d403780090d39e0aa8e5ef3e62cf229cfed3a14) Thanks [@pdlug](https://github.com/pdlug)! - Materialize class membership once while paging current identity classes. SQLite could otherwise repeat the full node scan inside the kind filter for every class, making `identity.classes()` much slower than the earlier in-memory read on graphs with many unrelated node kinds. ## 0.71.0 ### Highlights TypeGraph 0.71 makes identity groups easier to inspect. `identity.classes()` pages through visible classes, including singletons, at current or historical read coordinates. `identity.explainSame(a, b)` returns a shortest proof through persisted `same` assertions and eligible same-ID folds, so applications can show why two references belong together. Cursors are bound to the graph, coordinate, and kind scope, and remain compact as the graph schema grows. Historical identity reads and traversals now agree on which kinds belong to the current graph. Explanations at a historical coordinate use only folds that existed then. Schema migration errors also identify changed validators with exact JSON Pointers and before-and-after patterns, making a blocked migration easier to diagnose. ### Upgrade notes - Replace `IdentityReadSurface` with `IdentityReadFacade` and `IdentitySurface` with `IdentityFacade`. The surface aliases are no longer exported. If you implement `IdentityReadFacade` yourself, add `classes` and `explainSame`; these methods are also part of merge callback read contexts. - Before removing a node kind that connects retained identities through `same` assertions, move the needed assertions to retained kinds if those identities should remain joined. Historical identity reads and identity-expanded traversals now exclude kinds absent from the current graph. - To use current-coordinate `identity.classes()` with a custom backend, provide SQL window-function support and declare `capabilities.windowFunctions: true`. A backend profile that declares `false` raises `ConfigurationError` for this read. ### Minor Changes - [#748](https://github.com/nicia-ai/typegraph/pull/748) [`d1c8322`](https://github.com/nicia-ai/typegraph/commit/d1c83228cb4a83c9a99eb6af2c0663dd7eaddc4f) Thanks [@pdlug](https://github.com/pdlug)! - Add `identity.classes({ kinds, cursor, limit })` to read visible identity classes in deterministic pages at the current or a historical coordinate. Pages include visible singleton classes and expose registered visible members of each matching class. Opaque cursors are bound to the graph, read coordinate, and kind filter. - [#748](https://github.com/nicia-ai/typegraph/pull/748) [`d1c8322`](https://github.com/nicia-ai/typegraph/commit/d1c83228cb4a83c9a99eb6af2c0663dd7eaddc4f) Thanks [@pdlug](https://github.com/pdlug)! - Add `identity.explainSame(a, b)` to return a shortest path of persisted same assertions and implicit same-ID folds at the facade's read coordinate. - [#750](https://github.com/nicia-ai/typegraph/pull/750) [`b305ad9`](https://github.com/nicia-ai/typegraph/commit/b305ad9e9b7f890a4d497135ad9fc38422ebf0e0) Thanks [@pdlug](https://github.com/pdlug)! - Make `IdentityReadFacade` and `IdentityFacade` the complete public identity surfaces, including `classes` and `explainSame`. Replace the exported `IdentityReadSurface` and `IdentitySurface` aliases with those facade types. Historical identity reads and traversals now agree on registered kinds, and `explainSame` cites an implicit same-ID fold only when both nodes existed at the requested coordinate. Class cursors keep a fixed size as kind filters grow. Identity invariant errors include graph details and an appropriate current or historical recovery hint. ### Patch Changes - [#748](https://github.com/nicia-ai/typegraph/pull/748) [`d1c8322`](https://github.com/nicia-ai/typegraph/commit/d1c83228cb4a83c9a99eb6af2c0663dd7eaddc4f) Thanks [@pdlug](https://github.com/pdlug)! - Report exact JSON Pointers and before-and-after pattern values in schema migration diagnostics so validator changes can be located and reviewed precisely. ## 0.70.0 ### Highlights TypeGraph 0.70 strengthens durable branch identity and recovery. Each allocation has its own ID, which is checked against the sealed host origin when a branch is reopened, destroyed, or merged. Hosts can persist the branch and allocation IDs before creation to reconcile an uncertain result. Durable merge plans now retain recorded fork points, and strategies can keep older locator formats readable while writing a new format. PostgreSQL namespace forks now support the bundled pgvector storage. Embedding rows join the same verified snapshot as the rest of the graph, and owner-side preparation builds the target's vector tables and eligible indexes before the runtime copy. IVFFlat indexes are deferred until a post-copy `materializeIndexes()` call so they cluster the forked rows. Base-version tokens are now printable and can be stored directly in PostgreSQL text and JSON columns. Incremental merges now preserve a node already committed by the target when another branch proposes the same entity. This prevents a second merge from trying to change the endpoints of existing committed edges. Local `@libsql/client` 0.18 clients also use transaction framing compatible with pooled connections. ### Upgrade notes - Finish or remove durable branches created by an earlier release before upgrading, then create new descriptors and sealed host origins with allocation IDs. Earlier descriptors cannot be reopened, destroyed, merged, or used for evidence access in 0.70. Re-branch other work whose legacy `base@V` token is needed for merge planning; existing plans can still apply when their target fence has not moved. - Update custom durable strategies: accept the allocation ID in `create()`, return `{ operations, cursor, hasMore }` from operation scans, and capture `forkRevision` inside `create()` only when it is atomic with allocation. Remove native `merge` implementations; durable plans now apply through the target Store transaction. Use `readableVersions` if a new strategy version must read older locator formats. - Replace `installNamespaceForkLedger(target)` with owner-side `prepareNamespaceForkTarget(source, target)` before runtime namespace forks. Use matching vector storage on both backends for graphs with embedding fields, and call `materializeIndexes()` on the forked store after copying a graph with IVFFlat indexes. - Re-branch or re-plan work that uses an old `engine:` anchor or untracked content token. New untracked tokens include the complete graph content and active schema version. ### Minor Changes - [#745](https://github.com/nicia-ai/typegraph/pull/745) [`01b8149`](https://github.com/nicia-ai/typegraph/commit/01b814961c61b038fd71f6ad383bfdb0e350914b) Thanks [@pdlug](https://github.com/pdlug)! - Durable branch descriptors now carry a unique allocation ID, independent of the caller's branch ID. Reopen, destroy, and durable merge compare this ID with the host's sealed origin, so two copies using the same branch ID cannot be confused by a swapped locator. Callers may persist a stable branch ID and allocation ID before creation for host-side reconciliation after an uncertain result. A durable branch handle has the `DurableGraphBranch` type, which binds it to its allocation. `applyDurableMergePlan()` also carries the recorded fork point when comparing a branch with its descriptor, allowing plans for history-enabled durable branches to apply. Strategies may declare `readableVersions` alongside the locator format `version` they write, allowing upgraded strategies to continue reading and managing older locator formats. The descriptor version is passed to every read, destroy, and evidence method. A strategy may supply a `forkRevision` captured atomically with allocation; when it cannot, TypeGraph uses a full diff to avoid missing writes between allocation and sealing. Durable operation scans now return `hasMore` and retain their cursor at the end of a page, so callers can resume after later commits; strategies must order evidence by a monotonic commit position. Untracked stores now use the complete graph-content fingerprint and active schema version even when their backend offers lineage. The previous engine anchor could miss identity-only writes and did not read the planned graph state inside the commit transaction. The optional host-native `merge` strategy method is removed; `applyDurableMergePlan()` applies through the target Store transaction until a native merge contract can prove its target fence across the native operation's commit boundary. ### Upgrade notes - Re-create durable branch descriptors and sealed host origins from earlier releases. They lack the required `allocationId` fence and are refused on reopen, destroy, merge, and evidence access. Keep the previous release available to finish or remove those branches before upgrading. - Update `DurableOperationCapability.scan` implementations to return `{ operations, cursor, hasMore }`. The cursor must identify the last observed commit position even when `hasMore` is `false`; an empty page echoes `after`. - Move any strategy revision capture into `create()` and return it as `forkRevision` only when it was captured atomically with the physical fork. Omit it when the host cannot prove that cut. - Update `DurableWorkingCopyStrategy.create` to accept the allocation ID and refuse a duplicate until the host has explicitly reconciled it. Callers that need crash recovery should persist both IDs before calling `branchDurable` and pass them in options. - Re-branch or re-plan work whose base version uses the retired `engine:` anchor or an older untracked content token. New untracked tokens fingerprint current identity assertions and include the active schema version. - Remove `DurableWorkingCopyStrategy.merge` implementations and use `applyDurableMergePlan()`'s transactional apply. Database-native working-copy allocation remains supported. - [#744](https://github.com/nicia-ai/typegraph/pull/744) [`3f7da34`](https://github.com/nicia-ai/typegraph/commit/3f7da34e23265941cffc52cb984aa60d657a2014) Thanks [@pdlug](https://github.com/pdlug)! - `forkGraphNamespace()` now forks graphs that use the bundled pgvector storage. Embedding rows are copied inside the same repeatable-read snapshot, included in the content digest that the copy, retries and `abort()` verify, and removed by `abort()`. A graph with embedding fields forks only between backends with the same vector storage, pgvector on both sides or `vector: false` on both; custom vector and fulltext strategies are still refused. `prepareNamespaceForkTarget(source, target)` is the owner-side step. It installs the retry ledger, creates the graph's pgvector tables, and builds every index the source has materialized for the graph with the DDL the source used. It writes no graph rows and no materialization records, and the runtime fork still issues no DDL. IVFFlat indexes need the copied rows to cluster well, so preparation skips them, the fork does not copy their records, and `fork.store.materializeIndexes()` builds them after the copy. `materializeIndexes()` now rebuilds an IVFFlat index that exists without a materialization record, for example one an aborted fork left behind, instead of keeping it with `IF NOT EXISTS`: it was clustered for other rows. A backend without `dropVectorIndex` keeps the previous behavior. A materialized vector index no longer makes the fork refuse, and indexes whose build never completed on the source are neither built on nor required of the target. ### Upgrade notes - Replace `installNamespaceForkLedger(target)` with `prepareNamespaceForkTarget(source, target)`, run with the schema owner role before the runtime fork. `installNamespaceForkLedger` is removed. - Namespace-fork backends no longer need `vector: false`. For a graph with embedding fields, open source and target with the same vector storage: pgvector on both, or `vector: false` on both. - After forking a graph that declares IVFFlat indexes, run `materializeIndexes()` on the forked store under the owner role to build them. - [#743](https://github.com/nicia-ai/typegraph/pull/743) [`b69ec0b`](https://github.com/nicia-ai/typegraph/commit/b69ec0b9f4dc358b031e39eb8728b6ea11be871a) Thanks [@pdlug](https://github.com/pdlug)! - `base@V` tokens are now printable text. Their two components were joined by a NUL character, which PostgreSQL `text` and `jsonb` columns reject, so an application could not persist a durable branch descriptor, a recorded fork point, a merge plan, or durable operation evidence in PostgreSQL without re-encoding it. The separator is now `|`. ### Upgrade notes - Merge or re-create branches and durable branches minted by an earlier release. Their `base@V` tokens are refused with `BaseVersionMismatchError` and `details.reason: "legacy-token-format"` when `planMerge()`, `merge()`, `planMergeIncremental()`, or `mergeIncremental()` validates the branch's base, and when an incremental plan starts from a persisted `RecordedForkPoint`. Reopening a durable branch still succeeds; planning a merge from it does not. - Existing merge plans are unaffected. Applying a plan, including through `applyDurableMergePlan()`, validates the plan's target fence (graph id, schema, and revision anchor), not the format of the tokens recorded in its anchors. A plan whose target has not moved since planning still applies after the upgrade. - Durable operation evidence stores the coordinates the host supplied and is not compared with newly minted tokens, so existing evidence remains readable. - Code that stored tokens in a re-encoded form (base64, or JSON text in a `text` column) keeps working and may store them directly. ### Patch Changes - [#740](https://github.com/nicia-ai/typegraph/pull/740) [`fe345b0`](https://github.com/nicia-ai/typegraph/commit/fe345b0e08213d414dc71321bc39bfe30c345e07) Thanks [@pdlug](https://github.com/pdlug)! - Update `nanoid` to 6.0, which requires Node.js 22 or later, matching the package's existing `engines` range. `ExportOptionsSchema.signal` keeps its declared `ZodCustom` type, so the published declarations stay valid across the whole `zod ^4.0.0` peer range. - [#746](https://github.com/nicia-ai/typegraph/pull/746) [`6ad8c03`](https://github.com/nicia-ai/typegraph/commit/6ad8c030a5786ad09a73ddbc204bdc6c2d68f730) Thanks [@pdlug](https://github.com/pdlug)! - Incremental merges now keep a node the target committed after the fork point as the survivor when a branch proposes the same entity. Two branches forked from one point that both added an entity could previously fail on the second merge: when the second branch's node had the lexicographically smaller id, it won survivor selection, and the plan tried to repoint the committed edges of the first branch's node, which `applyMergePlan()` refused as an immutable-endpoint change. Merges that resolved through `blockIndex` or a unique constraint were not affected. - [#741](https://github.com/nicia-ai/typegraph/pull/741) [`b8fd08e`](https://github.com/nicia-ai/typegraph/commit/b8fd08e2854d609325928038725c5502027b4b81) Thanks [@pdlug](https://github.com/pdlug)! - Support `@libsql/client` 0.18 local clients. From 0.18 a local client pools its connections and rolls back any transaction a single `execute()` leaves open, so the raw `BEGIN`/`COMMIT` framing used for local clients failed every transaction with "no transaction is active". `createLibsqlBackend()` now probes whether a local client's `execute()` calls share one session and frames transactions through `client.transaction()` when they do not; clients before 0.18 keep raw `BEGIN`/`COMMIT`. ## 0.69.0 ### Highlights TypeGraph 0.69 can copy one graph namespace into a separately allocated PostgreSQL database without discarding its recorded history. `forkGraphNamespace()` verifies a repeatable-read source snapshot against the target before commit and records a durable proof for exact retries. Durable branches can carry recorded fork points, allowing incremental merge planning to use changes since that point when lineage proves them complete. Revision-tracked stores without recorded history can now use a revision-change journal for bounded changed-key lineage. The schema owner installs the journal, while runtime reads verify its readiness without running DDL. Disposable working copies can opt out with `revisionJournal: false`, including clones created by `branchForEvolution()`. PostgreSQL backends opened over transaction handles serialize statements on their pinned connection, and contribution-marker reads inside transactions use that same session. Schema tooling can inspect a graph extension without opening a Store through `introspectGraphExtension()`. Linear traversal queries also carry their final hop directly into the projection. ### Upgrade notes - Adopt base schema version 4 with the schema owner before deploying runtime roles. On existing PostgreSQL databases, run the generated migration or open once with privileged `createStoreWithSchema()` or `createAdapterStoreWithSchema()`. Older library versions refuse the newer base-schema marker. - If a revision-tracked store without history needs journal-backed lineage, call `installRevisionChangesJournal()` once as the schema owner before runtime use. Without a ready journal, lineage raises `REVISION_JOURNAL_NOT_READY`; pass `revisionJournal: false` for a working copy that does not need it. Journal triggers capture every graph writing to their physical tables and rows have no automatic retention, so plan storage and retention before installing them on shared tables. - Install the namespace fork ledger with `installNamespaceForkLedger()` on a private target before calling `forkGraphNamespace()`. Allocate an independent target database and size the worker for the largest copied relation; the source holds one repeatable-read snapshot for the full copy. - `Store.clear()` now preserves graph-local contribution materialization markers by default. Pass `{ preserveContributionMaterializations: false }` when a full cutover purge must remove them. - Custom engine profiles adopting base schema version 4 need revision-change table and index DDL. To enable journal-backed lineage, also provide trigger installation and a readiness probe. ### Minor Changes - [#738](https://github.com/nicia-ai/typegraph/pull/738) [`8f26c78`](https://github.com/nicia-ai/typegraph/commit/8f26c78ec021668281e7f4dd44e743f32d779b73) Thanks [@pdlug](https://github.com/pdlug)! - Add history-preserving PostgreSQL graph namespace forks, store-free graph-extension introspection, recorded fork points for incremental merge, and bounded change enumeration for revision-tracked stores. Linear traversal queries now read their final hop directly. PostgreSQL transaction backends and bare client sessions serialize statements on their pinned connection; transaction marker checks read that same connection. Install the revision-change journal with `installRevisionChangesJournal()` during privileged schema setup. Runtime lineage verifies that its table and triggers are ready without running DDL; short-lived clones, including `branchForEvolution()` working copies, may set `revisionJournal: false`. Install the namespace fork retry ledger with `installNamespaceForkLedger()` on the private target before runtime use. `Store.clear({ preserveContributionMaterializations: false })` also removes graph-local contribution markers for cutover purges. ### Upgrade notes Adopt base schema version 4 with the schema owner before deploying runtime roles. Existing PostgreSQL installations need the generated migration or a privileged `createStoreWithSchema()` / `createAdapterStoreWithSchema()` open; the new revision-change table and index are part of that base schema. Install the optional revision-change function and triggers once with `installRevisionChangesJournal()` under the owner role. Runtime lineage only checks readiness and never runs DDL; a revision-tracked store without history throws `REVISION_JOURNAL_NOT_READY` when the journal is missing. Set `revisionJournal: false` for clones that do not need journal-backed lineage, including the fourth `branchForEvolution()` argument. `Store.clear()` preserves graph-local contribution materialization markers by default; pass `{ preserveContributionMaterializations: false }` to remove them during a full cutover purge. Journal triggers attach to whole physical tables, so on shared tables they record writes for every graph using those tables, and journal rows have no automatic cleanup or retention policy. Avoid enabling the journal on shared tables unless that cross-graph capture and unbounded retention are acceptable. `forkGraphNamespace()` holds one repeatable-read source transaction open for the entire copy, including row reads, target inserts, and digest checks. Long-running copies therefore retain the source snapshot until the copy finishes. Custom engine profiles need revision-change table and index DDL for base-schema version 4 adoption, and trigger DDL plus a readiness probe to enable the change journal. Missing dependencies raise `ConfigurationError` when those operations are requested. ## 0.68.1 ### Patch Changes - [#735](https://github.com/nicia-ai/typegraph/pull/735) [`149db17`](https://github.com/nicia-ai/typegraph/commit/149db17e9042ee461908b9677f7c6c5c7f7bcbd6) Thanks [@pdlug](https://github.com/pdlug)! - `cloneWorkingCopyStrategy` now imports with `onUnknownProperty: "allow"`, so a working copy — and therefore `branch()`, `ingestionBranch()`, and `planCandidateWriteSet()` — can be seeded from live rows that carry undeclared properties `validateStore()` already reports as healthy. Incoming candidate write-set documents remain strict. A streamed interchange abort now names the failing entity and property in the thrown message instead of wrapping only a generic abort. - [#737](https://github.com/nicia-ai/typegraph/pull/737) [`3428764`](https://github.com/nicia-ai/typegraph/commit/34287643aa708ed53caf56c8216c50042aef6441) Thanks [@pdlug](https://github.com/pdlug)! - `compareAndSet()` and `updateWhere()` no longer throw an untyped Zod error on node kinds whose schema has object-level refinements. Early field checks reconstruct a partial schema from `.shape` so refinements stay on the complete after-image, the same document `update()` already validates. ## 0.68.0 ### Highlights TypeGraph 0.68 adds atomic operations to durable graph-merge branches. `operateDurableBranch()` lets a durable host commit an opaque graph mutation and immutable evidence in one host transaction, so a process can recover and deliver committed work after a crash without inventing a second coordination protocol. Canonical request digests make retries exact: the same idempotency key replays its evidence, while a changed mutation or metadata payload conflicts without another write. The new evidence lifecycle is inspectable and bounded. Applications can read or page committed evidence, mark delivery monotonically, and ask whether any evidence remains undelivered. TypeGraph validates every host-returned outcome and preserves a typed destruction fence until downstream delivery is complete; unsupported hosts execute no mutation, and existing durable strategies remain valid without the optional capability. ### Upgrade notes - Existing `DurableWorkingCopyStrategy` implementations require no changes unless they opt into `operations`. To opt in, commit the host mutation and its immutable evidence in one transaction, attest the descriptor's sealed origin, enforce exact idempotency replay and digest conflicts, and serialize operations against destruction. - Treat only `applied` and `replayed` outcomes from `operateDurableBranch()` as committed. An `unsupported` outcome guarantees that the host ran no mutation SQL; do not recreate the atomic guarantee with callbacks or a separate evidence write. - Deliver committed evidence with `scanDurableOperations()` or `getDurableOperation()`, then call `markDurableOperationDelivered()` only after the downstream transaction commits. A first application must remain undelivered, and `destroyDurableBranch()` refuses while any evidence is undelivered. ### Minor Changes - [#731](https://github.com/nicia-ai/typegraph/pull/731) [`f8e800b`](https://github.com/nicia-ai/typegraph/commit/f8e800bca25f7168747624d1ab4dea525ec736c0) Thanks [@pdlug](https://github.com/pdlug)! - Add atomic durable-branch operations. A `DurableWorkingCopyStrategy` may now expose an optional `operations` capability that commits an opaque host mutation and its immutable evidence in one host transaction, keyed by idempotency. New public orchestrators `operateDurableBranch()`, `getDurableOperation()`, `scanDurableOperations()`, `markDurableOperationDelivered()`, and `durableBranchHasUndeliveredEvidence()` wrap it, with typed `DurableOperationError` subclasses for conflicts, unsupported capabilities, malformed evidence, and the undelivered-evidence destroy fence. ## 0.67.1 ### Patch Changes - [#729](https://github.com/nicia-ai/typegraph/pull/729) [`0be1336`](https://github.com/nicia-ai/typegraph/commit/0be1336dbbfb0dedfbe5f473936dbabcd6e4be77) Thanks [@pdlug](https://github.com/pdlug)! - `store.clear()` now deletes the graph's durable contribution-materialization markers along with every other graph-scoped row. A cleared graph no longer leaves graph-local marker rows (full markers for graph-scoped contributions and activation markers for deployment-scoped ones) behind on SQLite or PostgreSQL; deployment-scoped physical markers are preserved because they attest shared storage the per-graph delete never touches, and the next privileged boot re-records the graph-local rows from them. On backends with interactive transactions, the delete runs in the same transaction as the rest of `clearGraph`; it also tolerates the marker table's absence on databases that never materialized a contribution. ## 0.67.0 ### Highlights TypeGraph 0.67 adds durable graph-merge branches for workflows that outlive one process. `branchDurable()` creates a persistent working copy and a JSON-safe descriptor that can cross a queue, deployment, or machine boundary; `reopenDurableBranch()` restores the ordinary `GraphBranch` used by merge planning, and `destroyDurableBranch()` explicitly removes or archives the host allocation. The existing reviewable plan/apply lifecycle remains the source of truth for accepted graph changes. Durable strategies own host allocation and reconnection while TypeGraph validates the sealed graph, schema, branch, and base origin. Creation fences writes racing the allocation and accepts either an exact revision token or a complete semantic match when the persistent copy has its own revision namespace. Strategies can optionally attempt a proven-equivalent native merge; a mutation-free `unsupported` result returns to the complete portable apply path, while uncertain native failures never risk a second application. ### Upgrade notes - Existing `branch()` and portable merge workflows require no changes. Use the durable APIs only when a working copy must survive closing its current backend or move between processes. - Custom durable strategies must return a non-secret, JSON-safe locator and attest the complete sealed origin on reopen and destroy. Each opened working copy must provide either sound cross-client engine fencing or an allocation-wide exclusive writer lease; closing releases access but intentionally leaves the persistent allocation available until `destroyDurableBranch()` succeeds. - Treat `DurableWorkingCopyStrategy.merge()` as an optional optimization. Return `unsupported` only when no merge SQL or host mutation ran. Return `applied` only after proving the complete host diff equals the approved TypeGraph write set and validating the target fence on the merged resource. Throw on failed or uncertain native outcomes; TypeGraph will not fall back after a possibly partial application. Apply callbacks and persisted provenance continue through the portable path. ### Minor Changes - [#726](https://github.com/nicia-ai/typegraph/pull/726) [`595e6d9`](https://github.com/nicia-ai/typegraph/commit/595e6d9e5a68faf5ca118288bf921065b822ff61) Thanks [@pdlug](https://github.com/pdlug)! - Add durable graph-merge branches that can be serialized, reopened in a later process, and explicitly destroyed. Durable strategies attest the complete fork origin, prove the created working copy matches its stamped base, and declare either engine-level fencing or an allocation-wide exclusive writer lease. Approved plans can optionally use a strategy's proven-equivalent native database merge; unsupported native dimensions execute no host mutation and fall back to the complete portable plan application. ## 0.66.1 ### Highlights TypeGraph 0.66.1 fixes store-opening failures when upgrading older SQLite or PostgreSQL databases that never received the recorded-node and recorded-edge tables. Base-schema adoption now creates the missing tables and indexes while preserving existing graph data and custom table names, including when recorded history is not enabled. ### Upgrade notes - If an earlier upgrade failed with a missing recorded-table error, deploy 0.66.1 and retry your normal privileged store-open or `backend.adoptBaseSchema()` path. Adoption resumes from the installed marker, including databases left at base-schema version 2. - After upgrading and verifying successful adoption, remove any extra `bootstrapTables()` call added specifically to work around this failure. Keep bootstrap calls required by your normal provisioning workflow; runtime-only, least-privilege store opening still does not perform adoption. - Roll affected deployments forward. Do not roll back to 0.56.0 after the base-schema marker has advanced beyond version 1: that release refuses the newer marker. ### Patch Changes - [#724](https://github.com/nicia-ai/typegraph/pull/724) [`e3ebddd`](https://github.com/nicia-ai/typegraph/commit/e3ebddda8f37833255dcdd75cd38b5f24a02c2e5) Thanks [@pdlug](https://github.com/pdlug)! - Fix upgrades from legacy databases that lack recorded-node or recorded-edge tables. Version-3 base-schema adoption now creates these tables and their structural indexes before installing changed-since indexes on SQLite and PostgreSQL, preserving existing data and custom table names. Failed upgrades left at base-schema version 2 can retry through normal adoption without an explicit `bootstrapTables()` workaround. ## 0.66.0 ### Highlights TypeGraph 0.66 adds `tx.writeNodeUpsertBatch()` for recorded PostgreSQL transactions. Submit caller-assigned IDs spanning multiple plain node kinds and receive ordered postimages from one statement, while inserts, live updates, resurrections, history capture, and receipt counts remain atomic. The narrow envelope is designed for latency-sensitive heterogeneous node ingestion and preserves the existing per-entity pipeline for constrained writes. Recorded transactions also lease their schema-fence evidence across managed writes and capture flushes. Reusing that evidence removes repeated fence probes from history-enabled adopted transaction loops while retaining conservative per-write fencing for non-history adopted transactions. ### Upgrade notes - Use `tx.writeNodeUpsertBatch(entries)` only inside a recorded PostgreSQL transaction. The batch supports plain node kinds with caller-assigned IDs; it refuses operational-identity graphs, unique or disjointness claims, search or embedding projections, temporal options, unchanged-upsert coalescing, duplicate `(kind, id)` entries, stale schema fences, oversized batches, and unsupported backends. Split oversized inputs before retrying. - Custom backends may implement the optional exact-session heterogeneous upsert capability to support this method. Backends that omit it retain the existing portable write paths, and `tx.writeNodeUpsertBatch()` refuses with a typed unsupported-capability error. ### Minor Changes - [#722](https://github.com/nicia-ai/typegraph/pull/722) [`8ad6da1`](https://github.com/nicia-ai/typegraph/commit/8ad6da18fb352a7f593f0092f006e47b21ed7709) Thanks [@pdlug](https://github.com/pdlug)! - Add `tx.writeNodeUpsertBatch()` for one-statement, caller-ID upserts across plain node kinds inside a recorded PostgreSQL transaction. ## 0.65.0 ### Highlights TypeGraph 0.65 makes reviewed candidate writes evolution-aware. Plan candidate data against the schema produced by a pending evolution, then apply the schema and accepted writes together in one caller-owned transaction and recorded revision. If concurrent schema or data changes invalidate the planning snapshot, TypeGraph now reports an explicit retry-and-replan outcome. Query composition gains cold cursor pages for `batchOnce()`, portable array-membership expressions, and native tuple comparisons for eligible keyset cursors. Cursor pages can execute independently or share one statement with companion reads while preserving the same results and cursor shape, and array membership can compare against candidate-row or correlated outer-row expressions. Eligible existing rows in `bulkUpsertById()` now update as a version-gated batch while retaining history, uniqueness, full-text, and vector synchronization. Other cases continue through the portable row-wise path, while broader caching and set-oriented reads reduce repeated work across identity repair, query execution, candidate-scoped updates, and constrained-edge imports. ### Upgrade notes - For candidate writes planned alongside an evolution, create the target with `captureCandidateWriteSetTargetForEvolution(target, evolutionPlan)` and plan it with `planCandidateWriteSetForEvolution()`. Apply the returned artifact inside the matching evolved transaction. Treat `MergePlanningStaleError` (`GRAPH_MERGE_PLANNING_STALE`) as a concurrency signal: discard the artifact, recapture the target, and replan. - Custom dialect adapters that support `expr.arrayContains()` with expression operands should implement the optional `jsonArrayContainsExpression` hook with the documented JSON-array semantics. Adapters that omit it remain compatible, but compiling this expression refuses with a typed configuration error. - Custom backends may implement the optional `updateResolvedNodesBatch` member to accelerate eligible `bulkUpsertById()` updates. Preserve the expected-version gate across the complete input and return no partial result when any row is ineligible; omitting the member retains the row-wise fallback. ### Minor Changes - [#718](https://github.com/nicia-ai/typegraph/pull/718) [`8eb7ead`](https://github.com/nicia-ai/typegraph/commit/8eb7eada54f38503c129a11ae8d2a79c88ed9b31) Thanks [@pdlug](https://github.com/pdlug)! - Batch distinct existing-row updates in `bulkUpsertById()` while preserving version guards, recorded history, uniqueness claims, full-text indexes, and vector projections. - [#717](https://github.com/nicia-ai/typegraph/pull/717) [`22af384`](https://github.com/nicia-ai/typegraph/commit/22af384cfb4e37f34c56abdcbeee69090836a860) Thanks [@pdlug](https://github.com/pdlug)! - Add cold cursor-page reads that execute independently or compose with other reads in one `batchOnce()` statement. - [#714](https://github.com/nicia-ai/typegraph/pull/714) [`6036f98`](https://github.com/nicia-ai/typegraph/commit/6036f9841da1b3388a684596f15743a83ba721e2) Thanks [@pdlug](https://github.com/pdlug)! - Plan serializable candidate write sets against a pending schema evolution so the schema and accepted data can be applied in one adopted transaction and recorded revision. Document `MergePlanningStaleError` as a retry-and-replan concurrency outcome. - [#715](https://github.com/nicia-ai/typegraph/pull/715) [`8fe206d`](https://github.com/nicia-ai/typegraph/commit/8fe206dfd4b7f3eb62bf8e4de4afeb28b093a1e8) Thanks [@pdlug](https://github.com/pdlug)! - Add expression-level array membership for candidate and correlated row values, and optimize safe keyset cursor comparisons with native row-value tuples. ### Patch Changes - [#719](https://github.com/nicia-ai/typegraph/pull/719) [`a15e382`](https://github.com/nicia-ai/typegraph/commit/a15e382900bef8bd0d3ef84485327fa360a0be6b) Thanks [@pdlug](https://github.com/pdlug)! - Reduce repeated work in identity closure repair, projection and relation query execution, candidate-scoped updates, and constrained-edge imports. Reused queries now cache their compiled SQL templates while preserving fresh temporal bindings, and large identity or import batches avoid duplicate component expansion and per-key cardinality reads. ## 0.64.0 ### Highlights TypeGraph 0.64 lets set-based node updates take their candidates directly from the query DSL. Pass a same-Store or same-transaction query to `NodeCollection.updateWhere({ candidates, patch })` to reuse correlated cross-kind predicates, including relationships that are not stored as edges. TypeGraph projects the query back to root node identities, intersects it with any `where` or relationship selectors, and sends the result through the existing atomic update pipeline so validation, uniqueness, history, full-text, vectors, and revisions still succeed or roll back together. Store analysis can now share the transaction snapshot that gives its results meaning. Transaction contexts expose `describe()` and `validateStore()`, allowing population statistics and every validation page to run on the pinned session. Use repeatable-read or serializable isolation and consume all pages inside the callback when concurrent writes must not change the dataset between statements. Shared backend storage now has an explicit deployment lifecycle. Full-text tables are physically materialized once per deployment and activated independently for each graph, so later graphs can become ready without repeating privileged DDL; vector storage remains graph-scoped. Custom backends also gain the `endpointSetRead` capability bundle and `runEndpointSetReadConformance`, giving bulk endpoint reads one declared support verdict and a portable contract test. ### Upgrade notes - Before serving full-text traffic through a DML-only role, run `createStoreWithSchema(graph, privilegedBackend)` after upgrading so TypeGraph can attest the deployment-scoped full-text table and activate each graph that uses it. `createStore()` remains a zero-I/O attach and does not repair missing markers. Custom table-contribution strategies may set `scope: "deployment"` only when one physical table is shared across graphs; omitting `scope` preserves the existing graph-scoped behavior. - Custom backends that support `store.edges..bulkFindFrom()` or `bulkFindTo()` must expose `findEdgesByEndpointSet` on the executing backend object and should run `runEndpointSetReadConformance` in their adapter test suite. Backends that omit the member remain valid for singleton reads, while set-oriented endpoint reads refuse with `ENDPOINT_SET_READ_UNSUPPORTED`. ### Minor Changes - [#712](https://github.com/nicia-ai/typegraph/pull/712) [`295f646`](https://github.com/nicia-ai/typegraph/commit/295f646c0add46fbd115654790c983ddd50972e7) Thanks [@pdlug](https://github.com/pdlug)! - Add deployment-scoped contribution ownership. Shared full-text storage is physically materialized once and separately activated per graph, allowing subsequent graph opens to run without DDL privileges while vector contributions remain graph-scoped. - [#712](https://github.com/nicia-ai/typegraph/pull/712) [`295f646`](https://github.com/nicia-ai/typegraph/commit/295f646c0add46fbd115654790c983ddd50972e7) Thanks [@pdlug](https://github.com/pdlug)! - Add the `endpointSetRead` capability bundle and the framework-agnostic `runEndpointSetReadConformance` fixture for custom backends. Bulk endpoint reads now resolve one capability verdict and refuse with a typed error when set-oriented reads are unavailable. - [#712](https://github.com/nicia-ai/typegraph/pull/712) [`295f646`](https://github.com/nicia-ai/typegraph/commit/295f646c0add46fbd115654790c983ddd50972e7) Thanks [@pdlug](https://github.com/pdlug)! - Allow `NodeCollection.updateWhere()` to take a same-graph, same-execution-target query as its candidate source. Candidate queries can use correlated cross-kind predicates without stored edges; TypeGraph forces their root-node identity projection and intersects it with existing `where` and relationship selectors before running the ordinary atomic set-update pipeline. - [#712](https://github.com/nicia-ai/typegraph/pull/712) [`295f646`](https://github.com/nicia-ai/typegraph/commit/295f646c0add46fbd115654790c983ddd50972e7) Thanks [@pdlug](https://github.com/pdlug)! - Expose `describe()` and `validateStore()` on transaction contexts so population statistics and validation pages can run through the pinned transaction session. Callers can request repeatable-read or serializable isolation and consume all analysis work inside one callback when they need a stable data snapshot. ## 0.63.0 ### Highlights TypeGraph 0.63 can shape bounded child records for many parents in one query. `relation.topPerPartition({ partitionBy, orderBy, limit })` chooses up to N rows independently for each parent, and `expr.collect({ id, name }, { orderBy, filter })` assembles those rows into ordered, typed record arrays. The result composes with prepared queries and `batchOnce()`, avoiding a separate child query for each parent. Record collections retain scalar codecs, including Date and Boolean fields, and decode admitted SQL NULL fields as `undefined`. Filter missing children inside `expr.collect()` to keep a parent with no matches and return `[]`. Ranking bounds the rows returned per partition; it can still scan and sort candidate rows, so provide a stable ordering key and measure the query on representative data. ### Upgrade notes - Implement `orderedRecordJsonArray` in custom `DialectAdapter` implementations. It must preserve record field names and scalar values, apply collection-local ordering and filtering, and return `[]` for empty input. - Update expression AST visitors that inspect `CollectExpressionNode.operand` to handle both scalar expressions and `CollectRecordOperand` (`kind: "record"`, with named `fields`). - Advertise `capabilities.windowFunctions: true` on custom backends only when the active engine supports the window functions used by `topPerPartition()`; otherwise the new method refuses before SQL. For repeatable winners, include a stable final key in `orderBy`, and add relation ordering if the final result order matters. ### Minor Changes - [#708](https://github.com/nicia-ai/typegraph/pull/708) [`9d4f328`](https://github.com/nicia-ai/typegraph/commit/9d4f3280e1e917ebdbffa8bcd9271c1b82cf34a1) Thanks [@pdlug](https://github.com/pdlug)! - Add ordered record collections with `expr.collect({ field: scalarExpression }, { orderBy, filter })`. Explicit flat record projections retain named scalar fields, decode Date and Boolean values, and preserve admitted SQL NULL fields as `undefined`. Record collections work through relation composition, prepared execution, and `batchOnce()`. Custom `DialectAdapter` implementations must add the required `orderedRecordJsonArray` method when upgrading. This method emits the ordered, optionally filtered JSON record aggregate and returns `[]` for empty input. Consumers inspecting expression ASTs must narrow `CollectExpressionNode.operand`: it can now be a scalar expression or a `CollectRecordOperand` with `kind: "record"` and named `fields`. - [#709](https://github.com/nicia-ai/typegraph/pull/709) [`4a6568e`](https://github.com/nicia-ai/typegraph/commit/4a6568e6db6f69664b02b7d21a3731bbc5dd4a01) Thanks [@pdlug](https://github.com/pdlug)! - Add `relation.topPerPartition({ partitionBy, orderBy, limit })` to retrieve up to N rows per parent in one query. Explicit partition keys and ordering select winners independently for each parent, and the result can feed ordered record collections, prepared queries, and batches without losing scalar codecs. Filters, distinctness, and ranges before the stage select its candidates; filters afterward remove winners without replacement. Include a stable final ordering key for repeatable winners and add relation ordering to control the final result order. Backends must advertise `windowFunctions: true`; unsupported profiles refuse execution before SQL. ## 0.62.0 ### Highlights TypeGraph 0.62 lets schema changes, graph writes, recorded history, and application SQL commit or roll back together in a caller-owned transaction. Prepare the change with `planEvolution()` before opening the transaction, then apply it through `withEvolvedTransaction()`. Planning stays outside the write fence, and ordinary additions of kinds or optional scalar fields avoid entity scans and provisioning DDL. No-op plans can use `withRecordedTransaction()` without acquiring the exclusive evolution fence. Evolved callbacks operate against the resulting schema and return its exact version and hash in the transaction receipt. With TypeGraph-owned history, they also support schema-only recorded checkpoints through `requestRecordedRevision()`, so applications can persist an audit event even when no entities change. After the outer commit, `refreshSchema()` publishes the reconciled Store for subsequent work. Plans can move from a cached Store to a compatible `withBackend()` request Store within the same loaded TypeGraph module. Schema evolution also composes with approved merges. `branchForEvolution()` and `planMergeForEvolution()` prepare work against the resulting schema, allowing a merge that introduces new kinds to join the same atomic commit. Privileged adapters can provision required identity storage and vector slots inside that transaction; the default DML-only policy refuses such work before mutation. ### Upgrade notes - Bootstrap the normal TypeGraph storage before serving adopted evolution requests. Route plans requiring identity or vector provisioning to a privileged adapter configured with `schemaProvisioning: "transactional"`; bundled adapters default to `"dml-only"`. Run generic or concurrent index maintenance explicitly after commit with `materializeIndexes()` on the refreshed Store. - Call `planEvolution()` outside the write transaction and keep its opaque token in memory. Do not serialize, clone, or reconstruct it; transfer it only between compatible Stores from the same loaded module. A `new-kind` requirement describes an addition, not queued removal work. - Enter `withEvolvedTransaction()` before other TypeGraph callbacks on the same native transaction. Propagate callback failures so the caller rolls back, and finish all callback reads and writes before returning: escaped transaction contexts, deferred queries, and prepared batches refuse execution after callback completion. - Handle `SchemaFenceTimeoutError` by rolling back and retrying the entire native transaction. Replan outside the transaction after a stale-baseline refusal. Change plans use a finite fence wait, defaulting to 5,000 ms; pass `waitBudgetMs` only for change plans, since no-op plans refuse an explicit budget. - Treat `receipt.schema` and any recorded anchor as provisional until the outer commit succeeds. Then call `refreshSchema({ ref, minVersion: receipt.schema.version })` on the cached root Store. A matching cache skips SQL and does not probe for newer versions; omit `minVersion` when a fresh lookup is needed. Refresh performs no provisioning. - Build resulting-schema merge plans with `planMergeForEvolution()` before opening the caller transaction, using `branchForEvolution()` when the branch needs new kinds. Apply the merge inside the evolved callback before other target graph writes; old-schema merge plans are refused. - Add `planEvolution()` and `refreshSchema()` to custom `StoreEvolution` implementations. Declare `schemaProvisioning` explicitly on custom `AdapterBackend` and `SqlEngineProfile` implementations, choosing a policy that matches the connection's intended provisioning permissions. - Offer `adoptSchemaWriteTransaction` only when a custom adapter can prove the active caller session and provide bounded schema fencing on that same transaction. Keep PostgreSQL advisory locks transaction-scoped. Native SQLite adoption requires observable transaction state on the actual connection; noninteractive adapters and SQLite drivers without that evidence cannot adopt schema changes. ### Minor Changes - [#705](https://github.com/nicia-ai/typegraph/pull/705) [`acf5b47`](https://github.com/nicia-ai/typegraph/commit/acf5b47bc180520cf372dd1926285d00cbd55935) Thanks [@pdlug](https://github.com/pdlug)! - Plan schema evolution outside a write transaction with `store.planEvolution()`, then apply the version-bound plan alongside graph and application writes through `store.withEvolvedTransaction()`. Plans are opaque, nonserializable capability tokens that can move between compatible Stores from the same loaded module, including `withBackend()` request Stores. They expose `baseline` and `result` schema identities and a discriminated array of schema additions and apply-time requirements. A `new-kind` entry records a graph addition and does not imply a queued removal. No-op plans can use ordinary recorded transactions; metadata-only changes avoid entity scans and provisioning DDL. Schema fence waits are bounded and expose `SchemaFenceTimeoutError` for whole-transaction retry. Evolved callbacks use the resulting schema, support recorded revision requests, and return exact schema version/hash metadata alongside the provisional recorded receipt. Escaped TypeGraph reads and writes refuse after callback completion. Publish root Store changes after outer commit through read-only `refreshSchema({ ref, minVersion })`; a matching cached version needs no reload. Use `branchForEvolution()` to fork an isolated branch with the planned kind set and `planMergeForEvolution()` to prepare a merge for the resulting schema and apply it inside the evolved callback. Old-schema merge plans continue to refuse. Adapters default to a DML-only policy that refuses required identity or vector provisioning before mutation. Privileged adapters configured with `schemaProvisioning: "transactional"` provision identity storage, vector tables, and contribution markers on the caller's fenced transaction session, so outer rollback removes them with the schema and graph writes. Bootstrap storage is still required, and eager index maintenance runs explicitly after commit. SQLite schema adoption requires verifiable native transaction state, and noninteractive drivers remain unsupported. Custom implementations of the `StoreEvolution` interface must add `planEvolution()` and `refreshSchema()`. Custom `SqlEngineProfile` and `AdapterBackend` implementations must declare their schema provisioning policy explicitly; the bundled adapters default to `"dml-only"`. ## 0.61.0 ### Highlights TypeGraph 0.61 expands the query DSL from graph matching into composable SQL result shaping. Schema-aware expressions power `project()`, completed-match filters, grouping, aggregates, and correlated subqueries; projected relations can be combined, filtered, deduplicated, prepared, and batched without returning intermediate rows to application code. `expr.collect(value, { orderBy, filter })` builds ordered scalar lists in SQL, including empty per-parent lists when optional matches are filtered inside the aggregate. `count()`, `exists()`, and selected-query `first()` make common terminal reads direct. Fetching several subgraphs is now a straightforward latency optimization: `store.batchOnce(read => roots.map(root => read.subgraph(root.id, options)))` retrieves independent bounded subgraphs in one statement. Runtime-sized arrays, singleton batches, and empty batches are supported; empty batches submit no SQL. For overlapping roots with substantial shared payloads, opt-in `shareSubgraphs: true` can also share traversal and hydration work while preserving independent results. Benchmark that option against ordinary batching for your workload; disjoint or lightly overlapping roots may not benefit. Queries can start from an explicit list of node kinds, such as `from(["Person", "Company"], "entity")`, and return one ordered stream with compatible shared fields and kind-discriminated results. Pagination now preserves nullable sort partitions and nodes whose IDs overlap across kinds. Native null ordering and directed node index keys let applications align indexes with their actual sort and identity columns. Recursive queries can compose multiple traversal stages, stop expansion explicitly, and return paths containing kind-qualified nodes and directed edge references. Approved merge plans can join graph writes and application SQL in one caller-owned transaction through `applyMergePlanInTransaction()`. Workflows using TypeGraph-owned recorded history can also request a checkpoint with `requestRecordedRevision()` when no entity changes are needed; the completed capture receipt supplies the recorded anchor. These additions let applications commit their own receipts alongside TypeGraph work while retaining control of the outer commit and retry boundary. ### Upgrade notes **Queries and pagination** - Handle `undefined` for empty-input `sum`, `avg`, `min`, and `max` results. Equality and membership predicates require compatible operands, and `countDistinct` accepts scalar string, number, Boolean, or date values; project an explicit scalar key when replacing structured JSON or array distinct counts. - Move cross-alias conditions out of staged `whereNode()` / `whereEdge()` predicates into completed-match `where()`, accounting for optional-row filtering. Use only compatible shared properties for polymorphic predicates, grouping, and ordering; query a specific kind when a field is not shared. Low-level composed `resultPredicate` ASTs must use database-expression predicates, optionally combined with AND/OR/NOT. - Add an explicit query `limit()` when a ranked query must cap completed rows. Candidate `k` now bounds ranked candidates only; traversal fan-out can produce more than `k` result rows, including inside set operations. - Restart saved multi-kind cursors that lack the new `kind` identity column, including cursors from subclass-expanded sources. For traversal fan-out, include traversed-row identities in the ordering when one source node produces multiple rows. Remove query-level `limit()` / `offset()` before cursor pagination and pass only one pagination direction; conflicting options are now refused. - Keep one source per query, use unique aliases across nodes, edges, and recursive outputs, and pass non-negative safe integers for limits and offsets. Use integer subgraph depths from 0 through 1000 and supported traversal directions and cycle policies; invalid inputs are refused rather than ignored. - Build batched reads and set-operation operands from the executing Store or transaction context. Split requests explicitly when a batch exceeds its planning or bind budget. Custom operands exposing only `toAst()` must also supply execution provenance; prefer library-created queries. Keep legacy `select()` callbacks pure because they may be probed; use `project()` for database expressions and `map()` for transformations of decoded projected rows. **Transaction composition** - Call `applyMergePlanInTransaction(target, tx, plan)` inside an active callback from the same Store, before other writes to the target graph. Propagate its thrown merge error so the caller rolls back, and retry the entire native transaction if needed; the helper opens no nested transaction and performs no local retry. Use a plan with `persistProvenance: false`: persisted merge provenance cannot join this atomic unit, while report-only provenance remains available. - Request recorded checkpoints inside a writable callback with TypeGraph-owned history and obtain the anchor from its completed capture receipt; engine-native history and read-only transactions refuse explicit allocation. For atomic application receipt persistence, use `withRecordedTransaction()` and write the receipt through the same still-open native transaction before its outer commit. Do not infer a recorded revision number before capture flush. **Custom backends and dialects** - Implement the new `DialectAdapter` members `safeNumericConversion`, `unboundedLimit`, `textJsonArray`, `appendTextJsonArray`, and `orderedScalarJsonArray`. The last accepts one required `{ value, valueType, orderBy, filter }` argument: admit only SQL TRUE filter results, preserve included NULL elements, and return `[]` for empty input. Handle the dedicated `kind: "collect"` node in expression visitors; collection-only options are not ordinary aggregate-node fields. - Enable `capabilities.orderedAggregates` only after verifying ordered and filtered aggregate support on the active engine. Bundled PostgreSQL declares support; supported SQLite factories probe it. An unprobed custom or remote SQLite connection must explicitly declare verified support before using `expr.collect()`. Existing non-collection reads do not require this capability. - Supply engine serialization or session-bound READ COMMITTED isolation evidence through the write fence for composed merge application. Managed merge callbacks and adopted application now refuse missing or unsuitable evidence; a caller-serialization assertion alone is insufficient. Update custom backends and transaction test doubles that participate in these paths. ### Minor Changes - [#702](https://github.com/nicia-ai/typegraph/pull/702) [`81fa27b`](https://github.com/nicia-ai/typegraph/commit/81fa27b26e182b4be276255452bc4b27f3f366b7) Thanks [@pdlug](https://github.com/pdlug)! - Add `applyMergePlanInTransaction()` so applications can apply an approved merge plan, record graph receipts, and write application SQL under one caller-owned transaction and recorded-time receipt. Merge callbacks and adopted application now refuse custom backends without engine serialization or session-bound read-committed isolation evidence. Custom backends must expose that evidence through their write fence. Fix constrained writes in adopted SQLite history transactions by acquiring the writer slot through the internal transaction-control path while retaining capture lifetime checks. - [#701](https://github.com/nicia-ai/typegraph/pull/701) [`141deb4`](https://github.com/nicia-ai/typegraph/commit/141deb436af2818ca45288a647ede1fe7f61a6ff) Thanks [@pdlug](https://github.com/pdlug)! - Add directed node index keys that can interleave property and system columns, enabling B-tree indexes such as `(createdAt DESC, id ASC)` while keeping covering fields last. Export `NODE_SYSTEM_COLUMN_NAMES` as the readonly runtime companion to `NodeSystemColumnName` for config generation and validation. - [#703](https://github.com/nicia-ai/typegraph/pull/703) [`0e41ee3`](https://github.com/nicia-ai/typegraph/commit/0e41ee3703df914154e04aef815c6c354643ab87) Thanks [@pdlug](https://github.com/pdlug)! - Query a nonempty explicit list of node kinds with `from(["Person", "Company"], "entity")`. Shared fields support the existing query composition APIs, and full-node results retain kind-discriminated properties. Multi-kind cursor pagination and streaming now use both kind and ID to preserve rows when IDs overlap across kinds. Polymorphic sources refuse predicate, grouping, and ordering fields that are missing or incompatible across their kinds. Existing multi-kind cursors may need to be restarted because their identity columns now include kind. - [#694](https://github.com/nicia-ai/typegraph/pull/694) [`3ba5ff4`](https://github.com/nicia-ai/typegraph/commit/3ba5ff4926b3ac224f95ad4dea9c6fe7dd0e845b) Thanks [@pdlug](https://github.com/pdlug)! - Add an optional database-expression `filter` to `expr.collect(value, { orderBy, filter })`. SQL TRUE includes an element, while false and SQL NULL exclude it. Aggregate-local filtering preserves parent groups from optional traversals, so missing children can produce `[]` without removing the parent row; included NULL operands still decode to `undefined` and keep the collection's inferred element type. Keep scalar values and explicit nonempty ordering as the collection contract. `distinct` and aggregate-local `limit` remain unsupported, and collection expressions retain their dedicated `kind: "collect"` node. Change custom dialect adapters to accept one required `{ value, valueType, orderBy, filter }` argument in `orderedScalarJsonArray`. Apply the optional filter inside the aggregate before empty-input coalescing, preserve included NULL operands, and return `[]` for empty input. The existing `orderedAggregates: true` capability remains the declaration for filtered collections. - [#693](https://github.com/nicia-ai/typegraph/pull/693) [`6eb34ad`](https://github.com/nicia-ai/typegraph/commit/6eb34ad68c5b854a4f2af021fc52932392299f49) Thanks [@pdlug](https://github.com/pdlug)! - Add `expr.collect(value, { orderBy: [...] })` for ordered scalar collection aggregation. Project a relation, group by its parent columns, and collect string, number, Boolean, or date values into typed readonly arrays. Collection ordering is explicit and independent of result-row ordering; duplicates and nullable elements are preserved. Empty ungrouped collection aggregates return `[]`. Export `CollectOptions` for reusable helpers and expose collection expressions as their own `kind: "collect"` node. Ordinary aggregate nodes do not carry collection-only options. Collection results compose with preparation, projection, and one-statement batching. Structured equality restrictions continue to apply, and collections are materialized without implicit truncation. Object elements and aggregate-local limits are outside this scalar API. Custom dialect adapters must implement `orderedScalarJsonArray()` with ordering and empty-input semantics. Collection reads require `capabilities.orderedAggregates: true`; bundled PostgreSQL declares support, and supported preparable synchronous SQLite clients and the async libSQL factory probe for support at construction. Other unprobed SQLite connections remain unsupported unless their capability is explicitly declared after verification. Existing reads are unaffected. - [#691](https://github.com/nicia-ai/typegraph/pull/691) [`d80a10a`](https://github.com/nicia-ai/typegraph/commit/d80a10a21c10186e4286c0867890b28940304558) Thanks [@pdlug](https://github.com/pdlug)! - Add query `count()` and `exists()` terminals and selected-query `first()`. Scalar terminals count or test the current SQL relation, including grouping, limits, and offsets, without invoking result selectors. Chained `having()` conditions now accumulate with AND. Offset-only queries compile consistently on SQLite and PostgreSQL. Allow `batchOnce()` to accept runtime-sized readonly arrays, singleton tuples, and empty arrays. Nonempty batches execute one statement or refuse before execution; empty batches execute no statement. Independent subgraphs retain their own roots, projections, traversal windows, and results. Batches now validate graph and execution-target provenance, window-function support, the request count, and the backend's declared bind budget. Introduce schema-aware database expressions for SQL projection, predicates, ordering, grouping, and aggregates. `project()` builds SQL once, while `map()` transforms decoded rows; legacy `select()` keeps its compatibility behavior, including callback probing; keep selectors pure. Expressions include nested JSON paths, metadata, parameters, arithmetic, coalescing, conditions, and typed correlated `$exists()` / `$scalar()` subqueries with scope and temporal validation. Explicit projections support scalar terminals, preparation, and one-statement batching. Document `batchOnce(read => roots.map(root => read.subgraph(root.id, options)))` as the recommended pattern for reducing round trips across several independent, bounded subgraphs. Add explicit SQL relation composition for projected and aggregated results. Combine visible columns with set operations, filter and aggregate derived results, deduplicate whole projections, order output columns, and execute prepared or batched relations through shared infrastructure. Typed preparation declarations preserve binding names and values across composition boundaries. Grouped relations apply input distinctness, ordering, limits, and offsets before grouping; repeated `groupBy()` calls accumulate. Identity-only `distinctNodes()` deduplicates node identities, and relation paging and streaming require a proven unique order. Add scoped `where()` filters for completed graph matches, independently of optional-match and recursive hop constraints. Add `stopExpansion()` with an explicit stopping-node emission policy. Preserve these stages in prepared queries, batches, and logical plans, and document ranked candidates, fanout, and distinct-entity counting. Ranked candidate `k` no longer implicitly caps completed rows after traversal fanout, including set-operation operands. Use an explicit query `limit()` to bound the final row count. Add opt-in shared subgraph hydration with `batchOnce(build, { shareSubgraphs: true })`. Compatible reads share a multi-root traversal and hydrated entities while preserving per-request membership, projections, temporal coordinates, edge windows, and independent result objects. Default batching retains independent plans; benchmark overlapping, payload-heavy roots before enabling sharing. All current-time reads built inside a batch use one pinned instant. Add qualified recursive paths with `path: { format: "qualified", alias: "route" }`. The output alternates kind-qualified node references and edge references with traversal direction. Existing `path: true` and string aliases still return node-ID arrays. Compose multiple recursive traversal stages with separate depth, path, cycle, and stop state. Later stages expand upstream source identities and preserve prior row multiplicity; final filters and ranges apply after composition. Fixed-hop stages compose before and after recursion, retaining fixed-edge properties. The first recursive stage can be optional and preserves roots without eligible endpoints. Scalar recursive-edge projections remain unsupported. Ordered recursive reads now retain their sort columns when embedded in `batchOnce()`. Add transaction-bound `query()`, `neighbors()`, `countNeighbors()`, `subgraph()`, and `batchOnce()` reads. Every read executes through the open transaction and observes earlier writes in the callback; `tx.subgraph()` and `tx.batchOnce()` each execute as exactly one statement. ### Upgrade notes Equality and membership predicates now require compatible operands, and invalid dynamic literals are rejected before SQL execution. Aggregate results preserve scalar field types and represent empty-input SQL NULL as `undefined`; handle absent sum, average, minimum, and maximum results. `countDistinct` accepts only string, number, Boolean, and date operands; replace structured JSON or array distinct counts with an explicit portable scalar projection. Query sources cannot be replaced mid-chain, aliases must be unique across nodes, edges, and recursive outputs, and limits and offsets must be non-negative safe integers. Cursor pagination refuses query-level limits/offsets and conflicting direction options instead of ignoring them. Subgraph depths must be integers from 0 through 1000; unsupported traversal directions and cycle policies are refused. Build batch reads from the executing Store or transaction context, and split requests explicitly if a single statement exceeds its planning budget. Custom objects exposing only `toAst()` are no longer accepted as legacy set-operation operands: operands must also supply execution provenance. Use queries created by the same Store or transaction so graph and execution-target compatibility can be verified. Staged `whereNode()` and `whereEdge()` predicates now refuse cross-alias references instead of compiling incorrect comparisons; use completed-row `where()` for those conditions, accounting for its optional-row filtering behavior. Raw composed `resultPredicate` ASTs must use database-expression predicates, optionally combined with AND/OR/NOT. - [#699](https://github.com/nicia-ai/typegraph/pull/699) [`f7f7376`](https://github.com/nicia-ai/typegraph/commit/f7f7376ae27a516d93816f70815b46d0d267c835) Thanks [@pdlug](https://github.com/pdlug)! - Add `requestRecordedRevision()` to history transaction contexts so applications can create a durable recorded-time checkpoint even when a transaction makes no entity changes. Repeated requests and entity changes in the same transaction allocate a single revision, exposed through the terminal receipt. ### Patch Changes - [#700](https://github.com/nicia-ai/typegraph/pull/700) [`872ee09`](https://github.com/nicia-ai/typegraph/commit/872ee0997883eca0153c01d30c2eb9a1cb2430e2) Thanks [@pdlug](https://github.com/pdlug)! - Emit native null placement for field ordering so matching B-tree expression indexes can satisfy the primary sort without an added null-check key. - [#698](https://github.com/nicia-ai/typegraph/pull/698) [`cefe0b1`](https://github.com/nicia-ai/typegraph/commit/cefe0b1b45b6250c9f42d6331456e78b721d162e) Thanks [@pdlug](https://github.com/pdlug)! - Fix cursor pagination and streaming across nullable sort values. Forward and backward pages now preserve rows on both sides of a NULL partition, including tied values and queries that omit the sort field from their selected result. Existing ordering defaults remain unchanged. ## 0.60.0 ### Highlights TypeGraph 0.60 brings the set-oriented read APIs introduced in 0.59 into transaction callbacks. `TransactionContext` now provides `query()`, `neighbors()`, `countNeighbors()`, `subgraph()`, and `batchOnce()`, all bound to the open transaction so read-modify-write workflows can observe earlier writes without leaving their atomic boundary. The transaction forms preserve the physical guarantees that matter on a held connection: `tx.neighbors()` and `tx.countNeighbors()` each execute as one statement, while `tx.subgraph()` and `tx.batchOnce()` use exact-one-statement plans. Transaction-bound `batchOnce()` composes fluent reads from `tx.query()` with neighbor, count, and subgraph reads from its callback builder, returning independently typed results in tuple order. The surface is consistent across managed, adapter, receipt-enabled, recorded-time, measurable, and adopted transaction contexts. `withCheckedReads()` remains a root-store API: transaction reads already execute against the transaction's bound session, and callers choose the required snapshot behavior through the transaction isolation options. ### Upgrade notes - Update hand-built `TransactionContext`, adapter transaction-context, or measurable transaction-context mocks and wrappers with `query`, `neighbors`, `countNeighbors`, `subgraph`, and `batchOnce`. Contexts created by TypeGraph provide these methods automatically. - Replace runtime feature checks and edge-read-plus-node-load fallbacks inside transaction callbacks with the typed transaction APIs. Type callback parameters as `TransactionContext` or the appropriate adapter/measurable variant rather than as `Store`; `withCheckedReads()` is intentionally unavailable on transaction contexts. - On PostgreSQL, request `isolationLevel: "repeatable_read"` or `"serializable"` when several separate transaction reads must observe one stable database snapshot. `tx.subgraph()` and `tx.batchOnce()` remain one statement regardless of isolation level. ### Minor Changes - [#689](https://github.com/nicia-ai/typegraph/pull/689) [`956fd56`](https://github.com/nicia-ai/typegraph/commit/956fd560d270dc58fab687f810b2c63abd42694a) Thanks [@pdlug](https://github.com/pdlug)! - Add transaction-bound `query()`, `neighbors()`, `countNeighbors()`, `subgraph()`, and `batchOnce()` reads. Every read executes through the open transaction and observes earlier writes in the callback; `tx.subgraph()` and `tx.batchOnce()` each execute as exactly one statement. ## 0.59.0 ### Highlights TypeGraph 0.59 adds an exact-one-statement read batch for latency-sensitive request assembly. `store.batchOnce((read) => [...])` embeds two or more independent fluent queries, set operations, and batch-scoped neighbor, neighbor-count, or subgraph reads into one SQL statement, then restores their independently typed results in tuple order. There is no sequential fallback: if a read shape cannot be embedded, TypeGraph refuses it before execution instead of weakening the statement-count contract. Relationship reads no longer require applications to load every edge or hand-roll edge-plus-node joins. `store.neighbors()` returns each visible edge with its adjacent node in one statement, supports incoming, outgoing, and bidirectional reads, and can order by edge metadata or an adjacent-node property before applying a deterministic limit. `store.countNeighbors()` performs the matching aggregate without hydrating entities. `subgraph()` gains per-edge-kind windows with their own direction, ordering, and limit, so one traversal can follow different relationship kinds in different directions and retain only the top N edges for each oriented source. Direct and composable graph reads now share one public model without parallel `*Query` APIs. Direct `store.neighbors()`, `store.countNeighbors()`, and `store.subgraph()` execute eagerly; the corresponding `read.*` forms inside `batchOnce()` defer the same logical read so it can be embedded. Direct subgraph extraction keeps its backend-tuned plan of two statements on SQLite and three on PostgreSQL, while the batch-scoped form uses one statement on both backends. Query hooks report every submitted statement for these paths. `store.withCheckedReads(expectedSchemaVersion, fn)` extends schema-checked reads from one query to a fluent-query block. Every `.execute()` created through the scope checks the same expected active schema version, and a mismatch escapes through one callback boundary so an application can reload its schema and retry the whole read block. ### Upgrade notes - Update hand-built `Store`, history-store, recorded-read-store, and adapter-store mocks or wrappers that expose the complete store surface with `batchOnce`, `neighbors`, `countNeighbors`, and `withCheckedReads`. Stores created by TypeGraph provide these methods automatically. - Use `store.batchOnce()` only for two or more independent embeddable reads. Prepared queries, queued collection reads, pagination, streaming, and writes are intentionally excluded; keep using `store.batch()` for mixed queued reads and `store.transaction()` for atomic multi-operation work. - When adopting `withCheckedReads`, catch `SchemaChangedError` outside the callback, reload the reconciled schema, and rebuild the whole block before retrying. The scope accepts ordinary fluent `.execute()` reads; aggregates, set operations, prepared queries, pagination, streaming, and `batchOnce()` refuse rather than run without the version check. ### Minor Changes - [#687](https://github.com/nicia-ai/typegraph/pull/687) [`542e6a3`](https://github.com/nicia-ai/typegraph/commit/542e6a3de8e8aac001563fbd082bae4abdba1007) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.batchOnce()` for exact-one-statement independent reads, with a batch-scoped builder for composable neighbor, neighbor-count, and subgraph reads. Add one-statement `store.neighbors()` and `store.countNeighbors()` APIs with edge- or adjacent-node ordering, limits, and aggregates. Add per-edge-kind direction, ordering, and limits to `subgraph()` traversal and hydration. Direct `store.subgraph()` and batch-scoped `read.subgraph()` share result semantics while choosing backend-tuned and exact-one-statement physical plans, respectively. Add `store.withCheckedReads()` to bind an expected schema version once across a fluent-query read block. ## 0.58.0 ### Highlights TypeGraph 0.58 reduces database round trips in graph read paths. `store.bulkFindEdgesTo` and its pinned-view counterpart resolve inbound edges for a set of targets across node and edge kinds, mirroring `bulkFindEdgesFrom`. Callers can replace per-target lookups with a set-oriented read while retaining input order, repeated and empty target buckets, temporal visibility, and per-input limits. For a single edge kind, the existing `edges.Kind.bulkFindTo` remains available. Whole-node and whole-edge selections now choose a full-row fetch before executing SQL, including nested and spread selections detected during planning. Previously, a fresh query instance could issue a projected query, discover that the selector needed the complete entity, and fetch again. These selections now avoid that extra statement without requiring applications to retain query instances between requests. Selectors whose field needs depend on row values keep the existing fallback. `executeChecked(expectedSchemaVersion)` combines a relational read with an active schema-version check in one statement snapshot. It offers an explicit alternative to probing the committed version before fetching data: a mismatch raises `SchemaChangedError` before the selector runs, even when the query returns no rows. Applications can then reload the schema and rebuild the query before retrying. The check covers that statement; it does not pin later reads in the request or replace write fences. ### Upgrade notes - Update hand-built `Store` mocks and wrappers exposing the full store surface with `bulkFindEdgesTo`; query wrappers exposing the full executable-query surface must also forward `executeChecked`. Library-created stores, pinned views, and queries provide the new methods automatically. - When adopting `executeChecked`, catch `SchemaChangedError`, reload the reconciled schema, and rebuild the query before retrying. Start a new transaction if the old transaction holds a repeatable-read snapshot. An expected version of `undefined` means no active schema and is distinct from version zero. - Use checked reads for relational queries with ordinary bound values. They fetch full rows and support traversals, ordering, offsets, and limits; recursive and relevance-ranked queries require a separate schema probe. Replace named `param()` references with bound values before building a checked query. - Custom backends adopting checked reads must supply `tableNames.schemaVersions`, naming a relation with `graph_id`, `version`, and `is_active` columns whose active row agrees with `getActiveSchema`. A missing binding raises `ConfigurationError` before SQL execution. Bundled SQLite and PostgreSQL backends supply it automatically; this release requires no database migration. ### Minor Changes - [#683](https://github.com/nicia-ai/typegraph/pull/683) [`ce35043`](https://github.com/nicia-ai/typegraph/commit/ce35043aa9356317229d85d0c4b9998284a36e1a) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.bulkFindEdgesTo` and its pinned-view counterpart for set-oriented inbound reads across edge kinds. Detect whole-node and whole-edge selections before issuing a projected query, avoiding a redundant fetch for fresh query instances. Add `executeChecked(expectedSchemaVersion)` for a relational read and committed-schema check in one statement, with `SchemaChangedError` on mismatch, including empty results. ## 0.57.1 ### Patch Changes - [#681](https://github.com/nicia-ai/typegraph/pull/681) [`d34dd51`](https://github.com/nicia-ai/typegraph/commit/d34dd515a75ab2c34451a77b113a70ca3939ad7c) Thanks [@pdlug](https://github.com/pdlug)! - Fix upgrades from older databases that lack recorded identity-assertion storage. Base-schema adoption now creates the missing table and its structural indexes before installing the version-3 changed-since indexes, preserving existing data and honoring custom table names. ## 0.57.0 ### Highlights TypeGraph 0.57 opens the backend boundary. Until now the library shipped two backends and hardcoded what it knew about them, so reaching a third engine meant editing the query compiler. This release replaces that with a declared contract. `createPostgresBackend` and `createSqliteBackend` are now the same `createSqlBackend` applied to a bundled `SqlEngineProfile`, and a new entrypoint, `@nicia-ai/typegraph/adapters/drizzle/engine`, exports both profile builders alongside `deriveEngineProfile` — which produces a variant of a bundled profile with a bounded set of fields replaced and a typed refusal for anything else. Decisions the code used to infer from a dialect comparison are now facts a backend states: `catalog` for physical-schema introspection, `fenceSql` for the lock a fence spells, `writeFence` for the exclusion primitive the engine actually provides. Concurrency is the part of that contract with the longest reach. `capabilities.writeFence` is a discriminated union on `mechanism` — `advisory`, `engine-serialized`, `caller-serialized`, and the new `row`, a portable exclusion for an engine with no advisory-lock primitive, backed by a small `typegraph_fences` relation. An engine that resolves write conflicts at commit rather than by blocking declares `conflict: "commit-time"`, and TypeGraph then runs under an optimistic-retry tier: every store-owned transaction that takes a fence row replays as one whole unit on a real commit-time conflict, up to three attempts, invisible to hooks and to the caller. Conflicts everywhere now have one classifier and one typed error, `TransactionConflictError`, and `store.transaction()` accepts `retry: { attempts }` so an application can ask for the same replay on its own callback — under a documented replay contract, since such a callback runs more than once. The third thread is for engines that already implement, in the database, what TypeGraph otherwise implements in software. A backend can declare `lineage` — an opaque whole-database revision, plus the rows of a graph that changed since it — and graph-merge prunes its diff to that set instead of scanning. A backend can declare `recordedTime` and answer temporal reads from its own system-versioned tables, in which case TypeGraph builds no capture relations and runs no clock at all; every recorded read in the library now resolves through a single `RecordedReadSource` seam, so TypeGraph's own capture, an externally bound relation, and an engine's native history are three interchangeable bindings rather than three spellings of the same interval predicate. And `branch()` gains a second bundled strategy: where `cloneWorkingCopyStrategy` streams a base through public interchange into a fresh backend, `forkedWorkingCopyStrategy` hands off to a host that can copy a database itself — a file copy, `CREATE DATABASE ... TEMPLATE`, a provider's branch API. Because a fork is the same physical database rather than a replay, it carries what interchange cannot: soft-delete tombstones, `created_at`/`updated_at`, the `version` column, and the base's recorded history, so a fork can answer `asOfRecorded` for instants from before it was taken. Two fixes land regardless of which backend you run. `Store.clear()` now rotates the graph's durable revision-origin nonce in the same transaction as the clear. Previously a graph repopulated to look the same could mint a `base@V` token byte-identical to one from before the clear, and a branch forked against that older epoch would silently pass the merge precondition against entirely different content. And a PostgreSQL availability defect in constrained edge writes is repaired: the atomic edge-claim program built one predicate arm per proposed row, so a `bulkCreate` of a few thousand rows on a kind declaring a cardinality could run for minutes, grow past two gigabytes of server memory, and ignore cancellation. Those statements now drive from a single relation and are planned once. Both bundled backends emit the same SQL, advertise the same capabilities, and behave exactly as they did in 0.56. Every new backend member above is optional, and neither bundled profile declares `recordedTime`, so recorded time on SQLite and PostgreSQL stays TypeGraph-owned; the bundled backends continue to derive their own `lineage` from their recorded relations. ### Upgrade notes **Applications** - `store.transaction()` and `store.transactionWithReceipt()` now throw `TransactionConflictError` (code `TRANSACTION_CONFLICT`) instead of the raw driver error on a serialization failure or deadlock. Replace matches on SQLSTATE, driver message, or a driver error class with `instanceof TransactionConflictError`, and read the original off its `cause`. - `MergeError` raised once merge retries are exhausted now carries a `TransactionConflictError` as its `cause`, one link deeper than before. Code matching `mergeError.cause` against a driver error must match `mergeError.cause.cause`. - Before passing `retry: { attempts }` to `store.transaction()`, check the callback against the replay contract: it must await all of its own work, read and write only values it creates fresh on each call, cause no effect outside its own transaction, and tolerate running more than once. - A branch forked before `Store.clear()` now fails `merge()` with `BaseVersionMismatchError` — including when the graph was repopulated to look identical. Re-branch from the post-clear store rather than reusing a pre-clear branch. **Graph-merge and working copies** - `WorkingCopyStrategy.create` takes a second required parameter: `create(baseStore)` becomes `create(baseStore, base)`. `branch()` passes it automatically; a direct caller passes `await computeBaseVersion(baseStore)`. - `GraphBranch` gains a required `close()`. A hand-built branch object — a structural mock or test fixture — must supply one. - `Store` gains a required `workingCopyOptions` getter. A structural `Store` mock or wrapper not built through `createStore` / `createAdapterStore` / `createStoreWithSchema` must implement it. **Recorded time** - `RecordedInstant` admits a second anchor form, `e1:` (engine-native), beside `r1:`. `recordedInstantRevision()` now throws a `ValidationError` on an `e1:` anchor, and `compareRecordedInstants()` throws when handed two anchors of different ownership forms. Use `recordedInstantWallTime()` wherever an instant may be of either form. Two engine-native anchors minted in the same millisecond compare equal; a TypeGraph anchor's per-commit counter is strict. - `ExternalRecordedReadSource` and `TypeGraphRecordedReadSource` renamed their string discriminant from `source` to `kind` (`"external"` / `"typegraph-capture"`), freeing `source` for the seam method. Update any pattern match on the old name. `RecordedReadBinding` now names the three-member binding union; `RecordedReadSource` names the shared seam the three implement. **Custom backends and engine profiles** - Replace `capabilities.recordedTimeOwnership` with `EngineProvisioning.recordedTime`: ownership is now derived from that member's presence rather than hand-declared. `ENGINE_NATIVE_RECORDED_TIME_NOT_IMPLEMENTED` is removed with no replacement. A profile declaring `recordedTime` must also declare `lineage`, or `createSqlBackend` refuses with `ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE`. - Rename resolved-plan lock calls: `sql.advisoryLock(...)` becomes `sql.acquireKeyed(...)`, and `sql.advisoryLockWithIsolation(...)` becomes `sql.acquireKeyedWithIsolation(...)`. `sql.isolationFact(...)` is unchanged. - `FenceSql`'s three members are now all optional, since a `row`-mechanism target supplies a different subset than an `advisory` one. Reach a lock through the resolved plan's accessors rather than the members directly, or narrow for `undefined` first. - Add `fences` to any `ResolvedSqlTableNames` object literal built by hand. A caller that only overrides names through `createSqlSchema` or a bundled factory is unaffected. - Exhaustive switches gain new cases: `WriteFencePlan["kind"]` gains `"row"`, and `capabilities.execution.unitOfWork` gains `"optimistic-retry"`. - A profile that omits `provisioning.catalog` produces a backend with no `catalog`, and `store.materializeIndexes()`, `store.materializeSystemIndexes()`, the recorded-time schema check, and the recorded-time migration's column read each refuse with a `ConfigurationError` naming it. Supply `catalog`, or keep off those paths. - Trusted import on a custom PostgreSQL backend now refuses before any statement runs when the resolved write fence is `unfenced` or carries `drain: "none"` — which now includes an advisory-only declaration that previously took the table lock anyway. Declare `{ mechanism: "advisory", drain: "table-lock" }` to restore the lock. - `SqlEngineProfile` drops `firstParty` and replaces `buildOperations` / `lateMembers` with one opaque `assembly`. Build a profile through a bundled builder or `deriveEngineProfile`; a profile literal is no longer constructible, and first-party standing is bound to the object a bundled builder returned rather than to a field. - Supply `SqlExecutionAdapter.serializationFailure` if the engine's commit-conflict shape is not PostgreSQL's `40001` / `40P01` SQLSTATE. A registered classifier adds to the standard rules rather than replacing them. - An `"optimistic-retry"` backend requires `AsyncLocalStorage` to tell a nested write apart from an independent one. A runtime without `node:async_hooks` is refused with `OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` at the first retried unit; interactive backends are unaffected. **Operators** - The base schema gains a `typegraph_fences` relation on both dialects. Run the regenerated migration SQL (`generateSqliteMigrationSQL` / `generatePostgresMigrationSQL`) against an existing database before declaring `writeFence.mechanism: "row"` against it. The two bundled mechanisms, `advisory` and `engine-serialized`, need no migration and keep working unmigrated. ### Minor Changes - [#626](https://github.com/nicia-ai/typegraph/pull/626) [`9b17a68`](https://github.com/nicia-ai/typegraph/commit/9b17a689e109c84500b6faf1e087db001c6b780f) Thanks [@pdlug](https://github.com/pdlug)! - `@nicia-ai/typegraph/adapters/drizzle/engine` now exports `buildPostgresEngineProfile` and `buildSqliteEngineProfile`, the bundled `SqlEngineProfile` builders, so a caller can derive a variant of one instead of only consuming a finished backend. It also exports `deriveEngineProfile` (with `DerivableEngineProfileOverrides`, `DerivableEngineProfileKey`, and `DERIVABLE_ENGINE_PROFILE_KEYS`), which builds a fresh profile from a bundled one with a bounded set of fields overridden — a lock spelling, a declared capability, a resource-audit verdict, or a runtime dependency bag — refusing any other field with a typed error. `SqlEngineProfile.firstParty` is removed; first-party standing is now bound to the exact profile object a bundled builder returned rather than to a field, so a copy or derived profile never carries it forward. `SqlEngineProfile.buildOperations` and `.lateMembers` are replaced by one opaque `assembly` field, constructible only by the two bundled builders. `BackendResourceAudit` is now public on the engine entrypoint. A derived profile's overridden `fenceSql` now also backs PostgreSQL's fused schema-version + recorded-graph-write statement, not only its standalone lock sites. `FenceSql` itself shrinks to three author-supplied members — `advisoryLockExpression`, `isolationFactExpression`, and `lockTables` — with the standalone-statement forms every ordinary lock site calls (`advisoryLock`, `advisoryLockWithIsolation`, `isolationFact`) now derived by TypeGraph from the two expressions, so a backend author never spells both forms separately. The bags a derived profile shares with its base by reference (`declaredCapabilities`, `resourceAudit`, `autocommit`, `tableNames`, `fenceSql`) are frozen so mutating one through the derived profile can no longer corrupt the base's own. This entrypoint is unreleased, so none of the above is a breaking change; the two bundled backends' emitted SQL, capabilities, marks, and behavior are unchanged. See [Authoring an engine profile](https://typegraph.dev/backend-authoring) for the derivable-field table, the refusals a custom profile can hit, and a worked example. - [#625](https://github.com/nicia-ai/typegraph/pull/625) [`e34d53c`](https://github.com/nicia-ai/typegraph/commit/e34d53cbddc8d5872a778b1e473f44fb6c75b019) Thanks [@pdlug](https://github.com/pdlug)! - `GraphBackend` and `TransactionBackend` gain an optional `catalog` member (`BackendCatalogProbes`): `tableExists`, `tablesExist`, `indexStates`, `dropInvalidIndex`, `columnTypes`, and an `indexBehavior` bag (`concurrentBuilds`, `hasInvalidIndexState`, `supportsGinFamily`). `columnTypes` reports each column as a `CatalogColumn`, `{ name, kind, declaredType }`; `declaredType` is required, and every custom `columnTypes` implementation must populate it alongside the normalized `kind` a comparison classifies against. `dropInvalidIndex` is a root-backend operation on an engine with an invalid-index state: a `transaction()`-scoped PostgreSQL catalog refuses it with `CATALOG_DROP_INVALID_INDEX_REQUIRES_ROOT_BACKEND`, since PostgreSQL refuses `DROP INDEX CONCURRENTLY` inside a transaction block, while SQLite has no invalid-index state and stays a no-op in both scopes. `catalog` is the one physical-schema introspection surface a store path consults directly instead of compiling a portable query, and four call sites across three modules require it: `store.materializeIndexes()` refuses only once its empty-candidate short circuit and the status-table ensure step have already run; `store.materializeSystemIndexes()`, which has no candidate short circuit, refuses only once that same status-table ensure step has run; the recorded-time schema check and the recorded-time migration's column read likewise need it. `EngineProvisioning` gains a matching optional `catalog` field; a profile that builds one populates the backend's member, and a profile that omits it produces a backend with no `catalog` — those four call sites then refuse with a `ConfigurationError` naming `catalog` instead of reaching engine-specific SQL with nothing to spell it. `createPostgresBackend` and `createSqliteBackend` both supply `catalog`, each transaction reading its own session's uncommitted state rather than the root connection's. `DialectCapabilities` gains `subgraphMembershipStrategy` (`"materialized-ids" | "inline-cte"`), naming the plan-shape decision `store.subgraph()`'s reachable-node filter already made per dialect: fetch the traversal closure once and filter against a fixed id list, or embed the recursive closure in each fetch. This capability replaces an inline dialect comparison in `store/subgraph.ts`; emitted SQL, round-trip counts, and the resulting query's prepared-plan shape are unchanged for both bundled backends. The dialect-literal ESLint ban (previously scoped to the query compiler) now also covers `src/backend` and `src/store`, behind a named, ratcheted exemption inventory (`DIALECT_LITERAL_EXEMPTIONS` in `eslint.config.mjs`) asserted against the tree in both directions by `tests/dialect-literal-inventory.test.ts`. Every remaining exemption is a decision that is not query compilation (error classification, one-shot migrations, a driver-specific resource audit, a SQLite-only transaction write-lock flag, or the write-fence planner's own dialect-keyed lock semantics) and carries a reason and a site count. No bundled backend's emitted SQL, capabilities, or behavior changes. **Behavior change:** trusted import's PostgreSQL table lock now resolves the same write-fence plan every other lock site does, instead of unconditionally taking `LOCK TABLE ... ACCESS EXCLUSIVE`. Trusted import now refuses up front, before any statement runs, when a custom PostgreSQL backend's `writeFence` declaration resolves `unfenced` (no declaration present) or resolves a `lock` plan with `drain: "none"` — this now also catches an advisory-only declaration (`{ mechanism: "advisory", drain: "none" }`), which previously took the table lock anyway. Every refusal names the drain that could not be satisfied. `createPostgresBackend` itself rejects a `writeFence. mechanism: "engine-serialized"` capability override at construction (`ConfigurationError`, 'PostgreSQL backend capability overrides cannot declare writeFence.mechanism: "engine-serialized"'), so a declaration resolving `engine-serialized` is reachable only through a custom `SqlEngineProfile` or a hand-built PostgreSQL-dialect backend for an engine that genuinely serializes writers; for one, trusted import now takes no relation lock at all, where it previously took `LOCK TABLE ... ACCESS EXCLUSIVE` — the declaration states the engine serializes writers, so trusted import's own transaction is fence enough on its own. The `WRITE_FENCE_SQL_UNAVAILABLE` code applies only to the narrower case of a `mechanism: "advisory"` declaration with no `fenceSql` to spell the lock; every other refusal above is `WRITE_FENCE_UNAVAILABLE`. Declare `writeFence: { mechanism: "advisory", drain: "table-lock" }` — the bundled `createPostgresBackend` default, which also supplies `fenceSql` — to restore the lock. **Author-facing:** `CommonOperationStrategy` no longer carries `dynamicEdgeConvergence`. The flag it carried — whether a convergent edge create's non-durable match may inspect JSON match fields — moved onto `OperationFusionHooks.dynamicEdgeConvergence`, which the bundled dialect factories pass to `buildCommonOperationOptions`. Neither `OperationFusionHooks` nor `buildCommonOperationOptions` is exported from any entrypoint. No action is required of a backend author: a `CommonOperationStrategy` is not author-supplyable in this release. `strategy` is absent from `DERIVABLE_ENGINE_PROFILE_KEYS`, so `deriveEngineProfile` refuses it, and `SqlEngineProfile.assembly` — which replaced the `buildOperations`/`lateMembers` pair, see the derivable-profiles entry below — is branded with a non-exported symbol, so a profile cannot be built from a literal either. The bundled builders are the only source of a strategy. `SqlEngineProfile.graphTemplateRuntime.instantiateStatement` is a required builder: given a template and target graph's ids and schema hashes (`InstantiateGraphTemplateSqlParams` — `templateId`, `templateSchemaHash`, `graphId`, `schemaHash`, and the three physical table names it reads), it must return the statement that inserts the target graph's `schema_versions` row from the template's stored document and copies the template's contribution-marker rows into the target graph, taking the target graph's write lock — the same key the schema-commit fence takes — co-atomically with the insert on an engine that fences with locks. An engine whose dialect can compose a data-modifying CTE beside the schema INSERT (PostgreSQL) folds the marker copy and the lock into that one statement; an engine that cannot (SQLite) instead supplies the optional `copyContributionMarkers` dep, which runs the marker copy as a second statement once the schema row is confirmed. The bundled `postgresInstantiateGraphTemplateStatement` and `sqliteInstantiateGraphTemplateStatement` builders (`graph-template-sql.ts`) are what `createPostgresBackend` and `createSqliteBackend` supply to their own profiles; neither is exported, so a custom profile reaches the same shape only by copying a bundled profile and adapting its statement, the same as every other engine-owned SQL a profile supplies. The `adapters/drizzle/engine` authoring entrypoint that carries `SqlEngineProfile` and `CommonOperationStrategy` ships for the first time in this release, so neither the removed `dynamicEdgeConvergence` field nor the required `graphTemplateRuntime.instantiateStatement` builder ever appeared in a published version; the notes above only affect authors building a custom profile against `main`. - [#656](https://github.com/nicia-ai/typegraph/pull/656) [`b32b7fc`](https://github.com/nicia-ai/typegraph/commit/b32b7fc51b626aadf8a76cb6d0b3774c9ba33b8a) Thanks [@pdlug](https://github.com/pdlug)! - `GraphBackend` gains an optional `recordedTime` member (`EngineRecordedTimeMembers`): an engine that tracks recorded (system) time itself, rather than through TypeGraph's own capture relations and clock. `source(table, revision)` names the table expression `"nodes"` / `"edges"` / `"identityAssertions"` reads its recorded rows from AS OF an opaque `EngineRecordedRevision` (`{ revision, recordedAt }`) — the engine's own temporal-table syntax, with the interval already folded in — and `revisionNow(session)` reads `session`'s own recorded-time revision: the current COMMITTED revision on a root backend, or the PENDING revision an open `transaction()` handle's writes will land at once it commits (the position `TransactionReceipt.recorded` is stamped from). `requireRecordedTime` is the typed refusal for a caller that needs it and finds it absent, in the same style as `requireLineage`. `TransactionBackend`/`EngineProvisioning` gain the matching optional member, threaded onto every `transaction()` handle both bundled dialects build, exactly parallel to `lineage`. A profile that declares `recordedTime` must also declare `lineage` (engine-native history keeps no recorded relations for TypeGraph to derive a graph-merge change delta from); `createSqlBackend` refuses otherwise (`ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE`). Neither bundled Drizzle profile declares `recordedTime`, so `resolveRecordedTimeOwnership` derives `"typegraph-relations"` for both today, and every recorded-time integration suite and the parity snapshot are unchanged — the engine-native path is proven by a PostgreSQL-family simulation (`tests/backends/postgres/engine-native-recorded-time.test.ts`, `pglite-engine-native-recorded-time.test.ts`) that dresses TypeGraph's own recorded relations as a temporal-table expression, labeled as a simulation rather than a real third engine, since no bundled backend implements one. Every recorded read — the query compiler's recorded arm, `recorded-read-service.ts`'s point reads and scans, the historical identity readers — now goes through one `RecordedReadSource` seam (`source(table, revision)` / `predicate(prefix, revision)` / `carriesInterval`) instead of each spelling the recorded relation swap and the `recorded_from <= r AND r < recorded_to` interval itself. TypeGraph's own capture binding and the external `recordedRelation({ schema })` binding both implement it as the recorded relation plus the interval predicate (`carriesInterval: true`); a new third binding kind, built only for a store whose backend declares `recordedTime`, implements it as the engine's own `source` with `predicate` always `undefined` (`carriesInterval: false`) — the engine's own expression already scopes every row to exactly one revision. Emitted SQL for both bundled backends is unchanged: no query-compiler behavior differs for a `typegraph-relations` or external-binding store, proven by the untouched parity snapshot and the full recorded-time integration and property-law suites. `RecordedInstant` widens to a two-form grammar: TypeGraph's own `r1:<16-digit revision>:`, and a new engine-native `e1::` minted internally from a `revisionNow` result. `recordedInstantWallTime` works on either form; `recordedInstantRevision` and `compareRecordedInstants` are narrower — see Breaking below. Store construction derives `recordedTimeOwnership` once (`resolveRecordedTimeOwnership(backend)`, `"engine-native"` exactly when `backend.recordedTime` is declared) and branches only where engine-native genuinely differs from TypeGraph-owned capture: `history: true` builds the engine-native read binding and leaves the backend unwrapped — no capture relations, no clock, no write-fence-gated clock allocation; `revisionTracking: true` is refused regardless of whether `history` is also requested (`ENGINE_NATIVE_REVISION_TRACKING_UNSUPPORTED` — there is no TypeGraph clock for it to advance, and the engine's own revision is available only under `history: true`); an external `recordedRead` binding is refused (`ENGINE_NATIVE_RECORDED_READ_UNSUPPORTED`); `store.recordedNow()`, `store.revisionNow()`, and both transaction-commit sites that stamp `TransactionReceipt.recorded` now read the engine's revision through one owner, `#engineRecordedInstant(session)`, called once per transaction on the actual committing handle — never once per graph, and never unless a graph node/edge/identity write inside the transaction actually changed a row (a mutation witness watches the write surface itself, not the collection-level write-intent counters `receipt.writes` is built from, so a delete of a missing id, a found-not-created `insertNodeIfAbsent`, or a coalesced no-op upsert all leave `recorded` undefined); and `store.asOfRecorded(instant)` refuses an instant minted under the OTHER ownership form (`RECORDED_INSTANT_OWNERSHIP_MISMATCH`) before any read compiles. `migrateLegacyRecordedTime` refuses under engine-native ownership (`ENGINE_NATIVE_MIGRATE_RECORDED_TIME_UNSUPPORTED`): it rewrites TypeGraph's own recorded relations, which an engine-native backend does not have. Reconstructing identity at a recorded coordinate — `store.identityAtCoordinate` at a past instant, and the query compiler's historical identity traversal — is refused under engine-native ownership (`ENGINE_NATIVE_RECORDED_IDENTITY_UNSUPPORTED`): identity history reads TypeGraph's own recorded relations directly, which an engine-native backend does not populate. `resolveLineage` under engine-native ownership always answers with the backend's own `lineage` (the co-required member above), never the recorded-relations one, since there are no recorded relations to derive it from. Public exports beside `LineageMembers`: `EngineRecordedTimeMembers`, `EngineRecordedRevision`, `RecordedTimeSession`, `RecordedTimeBackend`, `RecordedReadSource`, `RecordedSourceTable`. `Store` gains a readonly `recordedTimeOwnership` property, the store-level reader of the derived ownership. Documentation: [Engine-native recorded time](/queries/temporal#engine-native-recorded-time) covers the reader-facing contract and the `e1:`/`r1:` rule; [Supplying `recordedTime`](/backend-authoring#supplying-recordedtime) covers what a profile implements; the [SQLite ↔ PostgreSQL parity matrix](/backend-setup#sqlite--postgresql-parity) and [Engine-native recorded-time codes](/errors#engine-native-recorded-time-codes) round it out. ## Breaking - `capabilities.recordedTimeOwnership` is removed. It was hand-declared and could fall out of sync with what a backend actually implemented; ownership is now derived from `backend.recordedTime`'s presence. Declare `EngineProvisioning.recordedTime` instead — its presence alone makes `resolveRecordedTimeOwnership(backend)` answer `"engine-native"`. - `ENGINE_NATIVE_RECORDED_TIME_NOT_IMPLEMENTED` is removed. There is no replacement code: the interim refusal it named no longer applies to any reachable construction path now that engine-native construction is implemented. - `RecordedReadBinding` widens from a two-member union (`ExternalRecordedReadSource | TypeGraphRecordedReadSource`) to three members, adding `EngineRecordedReadSource`. `RecordedReadSource` is repurposed and newly exported: it no longer names the binding union (that role moved to `RecordedReadBinding`) and instead names the shared seam shape (`source` / `predicate` / `carriesInterval`) all three binding kinds implement. - `ExternalRecordedReadSource` (the type `recordedRelation({ schema })` returns) widens: it now carries the `RecordedReadSource` seam's `source` / `predicate` / `carriesInterval` members alongside its existing `schema` and brand, and its string discriminant is renamed from `source` to `kind` (`"external"`) — the `source` name was freed for the seam method. The binding is brand-gated and built only by `recordedRelation({ schema })`, so this affects only code that pattern-matched the old `source` discriminant on a value it produced. - `TypeGraphRecordedReadSource` (the type `history: true` binds internally) gets the same two changes: it widens with the `RecordedReadSource` seam's members, and its string discriminant is renamed from `source` to `kind` (`"typegraph-capture"`). The binding is brand-gated and built only internally, so this affects only code that pattern-matched the old `source` discriminant. - `RecordedInstant`'s grammar widens to admit the `e1:` form alongside `r1:`, and `RecordedInstantParts` becomes a discriminated union (`kind: "typegraph" | "engine"`) instead of a flat `{ revision: number; recordedAt: string }`. `recordedInstantRevision(instant)` now throws a `ValidationError` for an `e1:` anchor — there is no TypeGraph numeric revision to return; use `recordedInstantWallTime(instant)` for a value that works on both forms. `compareRecordedInstants(a, b)` now throws when the two anchors were minted by different ownership forms, and compares two engine-native (`e1:`) anchors by `recordedAt` only — document the same-millisecond tie as a caveat in your own code if you compare engine-native anchors: two distinct engine revisions minted within the same millisecond compare equal, unlike a TypeGraph-owned anchor's strict per-commit counter. - [#633](https://github.com/nicia-ai/typegraph/pull/633) [`f6d5387`](https://github.com/nicia-ai/typegraph/commit/f6d5387e5fecbea08db627c17aaeb3226e9e1db9) Thanks [@pdlug](https://github.com/pdlug)! - `@nicia-ai/typegraph/graph-merge` now exports `forkedWorkingCopyStrategy`, `ForkedWorkingCopyOptions`, and `ForkHandle` — a second bundled `WorkingCopyStrategy` for `branch()`, alongside the existing `cloneWorkingCopyStrategy`. Where the clone streams the base through public interchange into a fresh backend, `forkedWorkingCopyStrategy({ fork, connect })` targets a fork-capable host: `fork(baseStore)` calls the caller's own host-level fork API (a file copy, `CREATE DATABASE ... TEMPLATE`, a hosting provider's branch call) and returns a `TFork extends ForkHandle` (an optional `dispose`), and `connect(fork)` opens a `GraphBackend` on the result. The connected backend's `close` is composed with `dispose` so the working copy's single `close()` releases both the connection and the fork, and a `connect` failure disposes the fork before rethrowing. `Store` gains a `workingCopyOptions` getter (returning the new `WorkingCopyOptions` type, also exported) — the one place a working-copy strategy reads a store's own hooks, upsert coalescing, SQL schema, auto-refresh-statistics threshold, query defaults, and externally-bound recorded-read relation, without re-deriving them from private state. A fork inherits the base's WHOLE such option set through it, plus `history`/`revisionTracking` matched to the base's own `historyEnabled`/`revisionTrackingEnabled`. This is safe because a fork is the SAME physical database as the base, so every one of those options names something the fork also carries. The clone strategy keeps its narrower, already-documented subset (`revisionTracking` only): its fresh backend is a distinct, empty database, so a schema naming the base's tables or an externally-bound recorded-read relation would misdirect it. `WorkingCopyStrategy.create` gains a second parameter, `base: BaseVersion` — the token `branch()` already stamped off the base store, passed through so a strategy that needs to re-validate its working copy (the fork strategy) compares against the caller's own token instead of computing a second one. This is an additive parameter on a callback type callers implement; existing implementations that ignore the second argument are unaffected. Code that invokes a strategy's `create` directly (rather than going through `branch()`) must now pass the base token too, e.g. `strategy.create(baseStore, await computeBaseVersion(baseStore))`. The strategy asserts `computeBaseVersion(forkStore) === base` right after attaching the store, and refuses with a typed `BranchError` — carrying `forkVersion`/`baseVersion` in `error.details`, closing the backend first — when they disagree; `branch()` returns that error as the `cause` of the `BranchError` it resolves with. This proves base-token equality (schema plus a revision anchor, or a live-content fingerprint) at the instant the fork was taken, not byte-for-byte physical identity — the untracked fingerprint deliberately omits tombstones, `created_at`/`updated_at`, the `version` column, and recorded history, and providing those unchanged is the fork mechanism's own contract, not something re-verified on every branch. That is still the right fence: the merge's lost-update guard reads `version` and the diff reads tombstones/timestamps straight off the fork, so a `fork` that is not a true physical copy breaks them regardless of what the content fingerprint agrees on. Unlike a clone, a fork is never rebuilt through `exportGraphStream`/`importGraphStream`, so it preserves soft-delete tombstones, `created_at`/`updated_at`, the `version` column, and — with `history: true` — the base's recorded relations, letting a fork answer `asOfRecorded` for instants before the fork was taken. `create()` also refuses, before ever attaching a store, when `connect()`'s backend aliases the base's own backend: the same backend object, one derived from the other through backend derivation, or two wrappers sharing one underlying connection. Only the fork is disposed in that case — never the aliased backend, which the base still owns — and the refusal is a typed `BranchError` naming `connect()`. This cannot detect a fresh backend built over the base's own connection pool when that pool audits as independent (a default-size `pg.Pool`, for example); a pooled checkout genuinely is a different connection from the pool's own perspective. See ["Forked working copies"](https://typegraph.dev/graph-merge#forked-working-copies) for a worked strategy and the suspend hazard on hosts that reclaim idle compute. ## Breaking - `WorkingCopyStrategy.create` now takes a second, required parameter: `create(baseStore)` becomes `create(baseStore, base)`. `branch()` passes it automatically, so this only affects code that calls a strategy's `create` directly (rather than through `branch()`) — pass `await computeBaseVersion(baseStore)` for `base`. - `GraphBranch` gains a required `close: () => Promise` member, releasing the branch's working-copy backend (composed with a forked working copy's host-level fork, when applicable). `branch()` and `ingestionBranch()` populate it; a hand-built `GraphBranch` object (a structural mock or test fixture) must now supply one too. - `Store` gains a required `workingCopyOptions` getter (see above). A structural `Store` mock or wrapper — one that is not built through `createStore`/`createAdapterStore`/`createStoreWithSchema` — must now implement it too. - [#617](https://github.com/nicia-ai/typegraph/pull/617) [`fa6b468`](https://github.com/nicia-ai/typegraph/commit/fa6b4683e4f39b140089fe369747d76c52620092) Thanks [@pdlug](https://github.com/pdlug)! - `createPostgresBackend` and `createSqliteBackend` accept `fulltext: false`, mirroring the existing `vector: false` option. The backend then advertises no `capabilities.fulltext` and omits the fulltext CRUD/search members (`upsertFulltext`, `deleteFulltext`, `upsertFulltextBatch`, `deleteFulltextBatch`, `fulltextSearch`) along with `hybridSearch` and `fulltextStrategy` instead of stubbing them, and the generated DDL and runtime contributions never create a fulltext table for that backend. A fulltext predicate, a `searchable()` field, `store.search.fulltext`, and hybrid search all refuse with `UnsupportedBackendCapabilityError` (reason `fulltext_unsupported`) against a fulltext-off backend, instead of compiling SQL against a table that does not exist. `hardDeleteNode`'s cascade skips the fulltext delete for such a backend rather than issuing a statement against a missing table; every other cascade step, and every configuration that still has a fulltext strategy, is unchanged. This refusal is not limited to the bundled backends: any `GraphBackend` — including a third-party one — that omits the optional fulltext members now refuses a write to a node kind with `searchable()` fields with the same typed error, rather than silently skipping the fulltext index sync as it did before. The read path keys off `capabilities.fulltext` rather than the optional members: a third-party `GraphBackend` that implements `fulltextSearch` and/or sets `fulltextStrategy` but never declares `capabilities.fulltext` now has a fulltext predicate and `store.search.fulltext`/hybrid refuse with `UnsupportedBackendCapabilityError`, where before they compiled and ran. This also replaces the `ConfigurationError` those two call sites previously threw against a backend with no fulltext strategy at all — callers catching on the old class or error code should switch to `UnsupportedBackendCapabilityError`. `capabilities.contributions.rebuild` on a fulltext-off backend tracks only the transactional-fence condition — with no fulltext contribution to fail the check, that condition is vacuously satisfied — and a `rebuildContribution` call naming the fulltext contribution on such a backend refuses with a typed `ContributionRebuildUnsupportedError` instead of running DDL against nothing. Soft-deleting or hard-deleting a `searchable()` node now succeeds on a fulltext-off backend instead of refusing: a delete removes data rather than accepting a write the backend cannot index, and this backend maintains no fulltext sidecar to issue that removal against, so there is nothing to do. Only create and update of a `searchable()` kind refuse. `createLocalSqliteBackend`, `createLocalPgliteBackend`, `createLocalSqliteStore`, and `createLocalPgliteStore` accept the same `fulltext?: FulltextStrategy | false` option and forward it to both their installation DDL and the underlying backend factory, so a batteries-included store can skip the fulltext table the same way a hand-wired one can. **Upgrade note:** `fulltext: false` stops creating and maintaining the fulltext table; it never drops one. On a database that already has fulltext rows, disabling fulltext leaves them in place and unmaintained — a hard delete performed while fulltext is off leaves an orphaned row behind in the fulltext table, because `hardDeleteNode`'s cascade has no active strategy to build a delete statement from. Re-enabling fulltext later therefore requires the destructive contribution rebuild, `store.rebuildContribution("fulltext")` (which drops and recreates the fulltext table), not `store.search.rebuildFulltext()`: that method pages live nodes to recompute their content, and a hard-deleted node has no row left for it to page, so it never revisits — and therefore never clears — the orphan. - [#630](https://github.com/nicia-ai/typegraph/pull/630) [`7bb743a`](https://github.com/nicia-ai/typegraph/commit/7bb743a5a4ad6ea1b3baf2066b3298ab07189842) Thanks [@pdlug](https://github.com/pdlug)! - A backend whose `capabilities.execution.unitOfWork` is `"batch"` (Cloudflare D1's `batch()`, Neon HTTP's `transaction(queries)`) fixes every statement before the first one runs and commits them together with no session in between. `resolveBatchWriteVerdict` in `src/backend/capabilities/batch-write-verdict.ts` is the one place that classifies a schema-managed write's fitness for that tier: given a write's already-proven need (an interactive callback, a probe-then-write constraint check, Operational Identity, history, or a schema commit), it either defers (any other tier) or returns a refusal carrying a stable `BATCH_WRITE_UNSUPPORTED` code, the reason, and a canonical explanation. Every enforcing gate that used to word its own batch-engine limitation independently — the constrained-write fence, `store.transaction`, Operational Identity's atomic-backend checks, recorded-time capture's transactionability guards, and each dialect's schema-commit refusal — now asks this one verdict for its phrasing and nests `{ code: "BATCH_WRITE_UNSUPPORTED", reason }` under `details.batchRefusal`, so every refusal on a batch-tier backend names the same reason in the same words. The portable schema-version fence keeps its plain, reasonless `SCHEMA_WRITE_FENCE_UNSUPPORTED` limitation for everything it reaches that isn't one of those five proven needs — an ineligible write kind, a derived backend, a provenance mismatch — rather than guessing which reason, if any, applies. A singleton node `create` with a caller-supplied id now fuses its schema fence on a batch-tier backend exactly as a generated id already did, provided the kind carries no declared unique constraint: the id-generation gate that existed for an interactive root's autocommit durability no longer excludes a batch program, which commits its one statement as a unit regardless of which id it carries. `isAutocommitSingleStatementWrite` — the separate, stricter classifier for a bundled root's transaction-free write — is deliberately not relaxed the same way: the fused supplied-id create instead proves `insertNodeIfAbsentWithSchemaFence` through the ordinary hooked write plan, which already selects the correct fenced statement per id. The tombstone-resurrection write a supplied id can fall through to is fenced immediately before it runs, so it refuses on a batch-tier target rather than writing the row unfenced. `tests/batch-engine-harness.ts` adds a fake D1 client and a fake Neon HTTP client, each backed by a real engine (better-sqlite3, PGlite) wrapped in a real transaction, so batch atomicity — a rollback on a failing statement, a stale schema version writing nothing, every refusal reason reaching its gate — is now proven against real engine behavior instead of a mocked response. Bundled interactive behavior, emitted SQL, and the engine-profile-parity snapshot are unchanged. - [#638](https://github.com/nicia-ai/typegraph/pull/638) [`d57098a`](https://github.com/nicia-ai/typegraph/commit/d57098a8b5fa8416842fb493de91dbd60d5ca9e6) Thanks [@pdlug](https://github.com/pdlug)! - `GraphBackend` gains an optional `lineage` member (`LineageMembers`): an opaque, whole-database `revision(session)` an engine can report and compare, plus `changesSince(session, revision, graphId)`, which names every node and edge of one graph that changed — inserted, updated, deleted, or resurrected — since that revision, or admits `{ kind: "unbounded" }` when it cannot bound the answer. It is a query surface over graph rows: a backend's own `lineage` writes nothing at all, and the one carve-out on the bundled recorded-relations derivation below is a graph-identity row, not a graph row. `requireLineage` is the typed refusal for a caller that needs it and finds it absent, in the same style as `requireCatalog`. `TransactionBackend` gains the same optional `lineage` member (through the new `LineageBackend` member type, mirroring `CatalogBackend`), so a profile-supplied `lineage` is visible on a `transaction()` handle exactly as `catalog` already was, not only on the root backend. `EngineProvisioning` gains a matching optional `lineage` field, forwarded onto the backend unchanged; neither bundled Drizzle profile supplies one, so a store's own recorded-relations derivation backs the capability instead (below); the `lineage` member itself emits no SQL. The recorded relations it derives from are already part of the schema regardless of `history`, and a DDL-running boot (`createStoreWithSchema`, unless `systemIndexes: "skip"`) now materializes two new system indexes on them, history on or off, plus a third structural index on the recorded identity-assertions relation. A caller that opted out with `systemIndexes: "skip"` gets the two system indexes on the next explicit `store.materializeSystemIndexes()` call instead of at boot — see the parity-snapshot note below for exactly what moves. `recordedRelationsLineage(store)` derives `lineage` from a store's own recorded relations for any store constructed with `history: true`. `revision()` reports `:` — the graph's durable, random revision-origin nonce (`typegraph_revision_origins`, minted on demand through `Store.revisionOriginNow()`) joined to its recorded-time clock — never the bare clock value alone: two independently created stores that happen to share a `graphId`, or the SAME store across a `Store.clear()` boundary, mint numerically comparable clock values, and only the origin tells them apart. `changesSince` refuses (`unbounded`) outright on an origin mismatch against the graph's LIVE origin row, before comparing anything numeric. Otherwise it covers every write shape a recorded relation can express — inserts, updates, soft deletes, hard deletes, and resurrections — deduplicated, and proves completeness directly: every integer revision between the requested one and the graph's current clock must carry direct evidence (a `recorded_from` or a non-sentinel `recorded_to`) in one of the three recorded relations; anything short of that — one OTHER writer sharing the graph's clock, a `revisionTracking`-only `Store` with no `history`, having advanced it without capturing a row, at ANY point in the span, not only the most recent one — reports `unbounded` rather than a delta missing that writer's rows. `resolveLineage(store)` is the one place graph-merge (and any other caller) picks a `lineage` source: the backend's own when declared, else this recorded-relations one when history is on, else `undefined`. A new system index, `since_idx (graph_id, recorded_from)`, backs `changesSince`'s completeness scan on the two recorded relations, and the recorded identity-assertions relation gains a matching `since_idx` of its own (structural, created with the table, since it is not a `materializeIndexes`-managed system index) — that same completeness scan folds the identity-assertions relation in alongside the two recorded relations, so a graph whose earliest captured commit only asserted an identity is never mistaken for a gap. A database already open when this ships adopts all three indexes on its NEXT open, through the base-schema release-3 adoption step below (`"lineage-since-index"`) — the same lazy backfill machinery a missing system index already goes through for any OTHER caller (the two recorded-relation indexes; the identity-assertions index is adopted only through the base-schema step, never through `store.materializeSystemIndexes()`), and immediately for the two recorded-relation indexes when a caller opted out of boot-time materialization with `systemIndexes: "skip"`, via its own explicit `store.materializeSystemIndexes()` call. The parity snapshot moves by exactly these three index declarations, plus one extra version-marker `INSERT`/`SELECT` round trip on each of four capture scenarios on both bundled backends (bootstrap publishing the new base-schema release below) — no other statement, and no graph-data write SQL, changes. `GraphBackend` adopters that ship their own `EngineProvisioning` gain a required base-schema release: `CURRENT_BASE_SCHEMA_VERSION` advances from 2 to 3, id `"lineage-since-index"`, adopting the three `since_idx` indexes above through `CREATE INDEX IF NOT EXISTS` (idempotent, safe to run concurrently, and a no-op on a fresh install whose generated DDL already carries them). The bump is one-way — there is no downgrade path — and deployment-visible: a database already stamped 3 is untouched, one stamped 2 is caught up in place on next open, and a store built against an `EngineProvisioning` whose adoption-step registry stops at 2 fails to construct (`CompilerInvariantError`, "adoption registry must end at the current version"). A zero-DDL `createVerifiedStore` attach against a database still stamped 2 refuses with `BaseSchemaMigrationError` until `adoptBaseSchema()` runs. A custom SQL engine profile must register a version-3 adoption step (or accept the three indexes into its own fresh-install DDL and mark the step `bootstrap: "covered-by-generated-ddl"`) before upgrading past this release. `base@V`'s anchor gains a third form, `engine::`, chosen when a store has no `revisionTracking`/`history` but its backend declares `lineage` directly (a capturing store's recorded-relations lineage never reaches this form — capture also turns revision tracking on, so the per-graph anchor wins first). `` is the SAME durable per-graph revision-origin nonce the revision anchor carries (`typegraph_revision_origins`, ensured at mint time on the store's own backend); the engine's own revision is whole-database, not per-graph, so pairing it with the per-graph origin is what keeps two independent databases whose engines coincidentally report the same bare revision string from minting indistinguishable anchors — without it, a branch forked from one database could satisfy the base-version precondition of an unrelated database. The precedence — revision anchor, then engine anchor, then the compatibility content fingerprint — is documented once, in `base-version.ts`. Re-validating an engine anchor checks the live origin row first (`revisionOriginMatch`, the same predicate the revision anchor's guard uses) and refuses with `BaseVersionMismatchError` ("forked from a different store") on a mismatch before ever consulting `changesSince`; once the origin matches, a raw revision mismatch is confirmed through `changesSince` before refusing, since the engine's revision is whole-database and an unrelated graph's commit must not fail this graph's merge — an empty delta is tolerated as unchanged, and a non-empty delta or `unbounded` raises `BaseVersionMismatchError` with `details: { expectedRevision, liveRevision, changedKeys? }`, where `changedKeys` (when present) is capped to the first 20 node keys and first 20 edge keys plus each list's own total count, never the raw unbounded delta. One known gap: `changesSince` names only node and edge keys, so a commit touching only a graph's current identity assertions is invisible to an engine-anchored guard and tolerated as unchanged — the content-fingerprint and revision-anchor forms do not share this gap. `GraphBranch` gains an optional `forkRevision`, the fork's own `lineage.revision(session)` captured by `branch()` right after the working copy is created, with the working copy's own root backend as the session — an origin-bearing token for the recorded-relations source, so clearing and repopulating the FORK itself to the same revision count `forkRevision` held is caught the same way a cleared BASE store already is, rather than looking unchanged. `diffAgainstBase` takes an optional `pruneTo` lineage delta: when present, each node/edge kind is read by id set instead of a full keyset enumeration, restricted to the union of what changed on the fork since `forkRevision` and on the base since its own `base@V` anchor. A key absent from both deltas cannot have moved since the fork point, so pruning cannot miss a change — it only narrows how much is read. Pruning applies only when both sides can supply a bounded delta; a hand-built branch, a store with no `lineage`, an `unbounded` answer on either side, or either side's `changesSince` REJECTING falls back to the full diff exactly as before. Pruning is a pure optimization: it never changes what a merge decides, only how much of the store it reads to decide it. `LineageMembers`' `revision`/`changesSince` each take a **session** as their first argument — the narrowest existing execution-target type a root backend and a `transaction()` handle both satisfy (`LineageSession`, `Pick`). An implementation MUST run its read on the session it is given, never on a connection it closed over instead: `assertTargetUnchanged` (`graph-merge/merge.ts`) is the concrete caller this exists for — it reads `lineage` off the pinned transaction handle and passes that SAME handle as the session, so the read observes the transaction's own snapshot. The one documented exception is the bundled recorded-relations derivation's origin resolution, which is a graph-identity row rather than a transaction-scoped fact and is deliberately resolved off `session` entirely (see the `lineage` capability's own doc). `requireLineage` now refuses with a `ConfigurationError` (`LINEAGE_UNAVAILABLE`) when the transaction handle carries no `lineage` of its own, with no fallback to the root backend's `lineage`; a `lineage` a custom backend wants honored at commit time must be threaded through `EngineProvisioning.lineage` so it reaches every `transaction()` handle, not attached only to the root object after construction. This is not listed under Breaking below: `lineage` shipped on this same unreleased branch, so its signature has never been part of a published release. ## Breaking - `BaseSchemaRuntime` (and the `CreateBaseSchemaMembersDeps` it is derived from) gains a newly required `sinceIndexDdl` field: `readonly [string, string, string]`, three `CREATE INDEX IF NOT EXISTS` statements in `(recordedNodes, recordedEdges, recordedIdentityAssertions)` order, built from a dialect's own physical table names via `sinceIndexAdoptionDdl` (`src/indexes/system.ts`). A custom `SqlEngineProfile` that builds its own `baseSchemaRuntime` must supply this field. - [#638](https://github.com/nicia-ai/typegraph/pull/638) [`d57098a`](https://github.com/nicia-ai/typegraph/commit/d57098a8b5fa8416842fb493de91dbd60d5ca9e6) Thanks [@pdlug](https://github.com/pdlug)! - `Store.clear()` now rotates the graph's durable revision-origin nonce (`typegraph_revision_origins`) in the same transaction as the rest of the clear, for every store able to mint either origin-namespaced `base@V` anchor form — a store with `revisionTracking` or `history` enabled (the TypeGraph revision anchor), AND an engine-anchored store whose backend declares `lineage` directly with tracking off (the engine anchor). Previously `clear()` reseeded (or, under `history`, left unseeded) only the recorded clock and left an engine-anchored store's origin untouched entirely, so a graph repopulated after `clear()` to look the same — the same revision COUNT for a tracked store, or a coincidentally-matching engine revision for an engine-anchored one — could mint a `base@V` token byte-identical to one minted before the clear, and a branch forked before the clear would silently pass the merge precondition against a base whose entire content had been replaced. `computeBaseVersion` and `Store.revisionOriginNow()` also now read that origin row fresh on every call instead of caching it per `Store` instance. Two live `Store` objects can legitimately observe the same graph, and only one of them runs `clear()` at a time; the removed cache previously let the OTHER instance keep minting anchors from its pre-clear origin until it happened to be recreated, so every merge into it failed at commit for no reason visible to the caller. This closes the BASE-side half of the epoch gap; the FORK side had an equivalent one of its own — `recordedRelationsLineage`'s `revision()` used to report the bare recorded-clock value with no origin, so `GraphBranch.forkRevision` carried nothing to catch a cleared-and-repopulated FORK either. That half is closed the same way, by embedding the origin directly in the bundled `EngineRevision` token every `revision()`/`changesSince()` call now compares — see the `lineage`-capability changeset for the token format. ## Breaking - A branch forked from a store BEFORE `Store.clear()` now correctly fails `merge()`'s `base@V` precondition (`BaseVersionMismatchError`) once that store has been cleared, even when the branch is later merged against a graph repopulated to look the same — for a revision-tracked store, the same revision count; for an engine-anchored store, a coincidentally-matching engine revision. This was always the documented intent — a cleared store is a new epoch a pre-clear branch cannot merge into — and is now enforced for BOTH anchor forms. Re-branch from the post-clear store instead of reusing one forked before the clear. - [#628](https://github.com/nicia-ai/typegraph/pull/628) [`f74582c`](https://github.com/nicia-ai/typegraph/commit/f74582c99fb3587725255959665b27da7abc9a43) Thanks [@pdlug](https://github.com/pdlug)! - Transaction conflicts (PostgreSQL serialization failures and deadlocks) are now classified by one shared predicate everywhere the store recognizes them, and reported through a new typed error, `TransactionConflictError` (code `TRANSACTION_CONFLICT`, `details: { operation, attempts }`, `cause` the driver error), exported from the package root alongside its sibling `VersionConflictError`. **Behavior change:** `store.transaction()` and `store.transactionWithReceipt()` now throw `TransactionConflictError` — not the raw driver error — when the backend reports a conflict, with `attempts: 1`. A caller that matched the previous driver-shaped error (by SQLSTATE, message, or `instanceof` on a driver error class) must instead match `TransactionConflictError` and read the same driver error off its `cause`. `store.transaction()` and `store.transactionWithReceipt()` accept a new option, `retry: { attempts: number }`, to have TypeGraph itself re-run the whole callback on a conflict, up to `attempts` times total, with no delay before the second attempt and a short capped, jittered backoff after. A retried callback must satisfy a replay contract — await all of its own work, read and write only values it creates fresh on each call, perform no effect outside its own transaction, and tolerate being invoked more than once — documented on the option and on the transactions guide. `HookContext` gains an optional 1-based `attempt` field (absent means `1`, so a hook context built outside the store still typechecks) so `onOperationStart` / `onBulkOperationStart` / `onQueryStart` / `onError` can tell a replay from a new operation; a rolled-back attempt's completed operations report neither `onOperationEnd` nor `onError` of their own, and a retried `transactionWithReceipt()`'s receipt reflects only the committed attempt's writes. Graph-merge's three commit paths (the public `apply`, its incremental variant, and internal plan commits) now go through the same retry owner as `store.transaction()`, with their existing budget of three attempts. `MergeError` on exhaustion now carries a `TransactionConflictError` as its `cause` (which itself carries the driver error), one link deeper than before — a caller matching `mergeError.cause` against the driver error directly must instead match `mergeError.cause.cause`. `capabilities.execution` gains an optional derived field, `unitOfWork?: "interactive" | "batch" | "none"`, naming how a backend groups a multi-statement write: `"interactive"` when it can hold an open callback transaction, `"batch"` when it cannot but exposes a native atomic program (an HTTP-only driver such as `drizzle-orm/neon-http`), otherwise `"none"`. Both bundled backends derive and populate it, overwriting anything a profile declared; a custom `GraphBackend` may leave it absent. Nothing in the store consumes it yet. - [#631](https://github.com/nicia-ai/typegraph/pull/631) [`aa0d599`](https://github.com/nicia-ai/typegraph/commit/aa0d5993a3f1feb4c20ae6a85a30bc13f4f71d40) Thanks [@pdlug](https://github.com/pdlug)! - `capabilities.writeFence` gains a third keyed mechanism, `"row"`: a portable exclusion for an engine with no advisory-lock primitive, backed by a new, never-dropped base-schema relation, `typegraph_fences(key TEXT PRIMARY KEY, generation BIGINT NOT NULL)` (`INTEGER NOT NULL` on SQLite). `{ mechanism: "row"; drain: "table-lock" | "quiescent" | "none"; conflict: "wait" | "commit-time" }` carries the same `drain` fact `"advisory"` already declares, plus `conflict` — the engine fact for two writers of one fence row: `"wait"` for a lock-based engine (the second acquirer's statement blocks, exactly like an advisory lock), `"commit-time"` for an optimistic-concurrency engine (both acquirers proceed and the loser's COMMIT fails). Every keyed lock site now shares one `case "lock": case "row":` body, spelling the acquisition through the resolved plan's `sql.acquireKeyed` / `sql.acquireKeyedWithIsolation` regardless of mechanism; the fences relation's key reuses each site's existing advisory namespace verbatim (`${namespace}:${key}`), so the lock-order contract carries over unchanged. The two bundled backends still resolve `"advisory"` / `"engine-serialized"` by default — emitted write SQL for both is unchanged, and only the bootstrap DDL gains the fences relation's `CREATE TABLE`. Existing databases add it by re-running the generated migration SQL (`generateSqliteMigrationSQL` / `generatePostgresMigrationSQL`), which now include it. `capabilities.execution.unitOfWork` gains `"optimistic-retry"`, derived (never hand-set) when the backend is interactive and its resolved write-fence plan is `"row"` with `conflict: "commit-time"`. Under that tier, every TypeGraph-owned transaction that acquires a fence row replays a real commit-time conflict as one whole unit (open, prelude, reads, writes, commit) through the retry owner introduced for `store.transaction()`, up to `OPTIMISTIC_RETRY_ATTEMPTS` (3) attempts. That covers every store-owned write — collection create/update/delete, bulk paths, `importGraph`, identity maintenance, contribution rebuild, index materialization — as well as the two backend-owned transactions that acquire the schema-commit fence row directly: graph-template instantiation and a schema commit (`commitSchemaVersion` and its three siblings). A nested write running inside an existing transaction never retries on its own (it cannot restart a transaction it does not own), so its conflict propagates unchanged to the outermost store-owned write or to `store.transaction` itself. Under `"interactive"` this changes nothing: one attempt, as before. An `"optimistic-retry"` backend requires `node:async_hooks`' `AsyncLocalStorage` to tell a nested unit apart from an independent one; a runtime without it is refused with `OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` at the first retried unit, while interactive backends are unaffected. `SqlExecutionAdapter` gains an optional `serializationFailure?: (error: unknown) => boolean` for an engine whose commit-conflict shape is not PostgreSQL's `40001` / `40P01` SQLSTATE (or its fixed message fallback). `isSerializationFailure(error, target?)` — the one predicate every retry owner consults — checks a classifier registered against `target` first, but the classifier only ever ADDS to the standard SQLSTATE/message rules: a registered classifier that recognizes `error` wins outright, while one that declines still falls through to those rules rather than having the final word, so there remains one predicate rather than a second inline check per engine. `onOperationStart`'s `attempt` field counts caller-owned retries of `store.transaction` only. A store-owned unit's own internal replay under `"optimistic-retry"` (a create, an update, a bulk write, `importGraph`, and the rest) is invisible to that count: `onOperationStart` / `onOperationEnd` / `onError` each fire exactly once for the outer unit, for its one committed attempt, no matter how many attempts the retry owner spent internally to reach it. ## Breaking - `FenceStatements`'s derived standalone-statement members are renamed to what a lock site actually asks for: `advisoryLock` → `acquireKeyed`, `advisoryLockWithIsolation` → `acquireKeyedWithIsolation`. `isolationFact` is unchanged. A caller consuming a resolved plan's `sql.advisoryLock(...)` / `sql.advisoryLockWithIsolation(...)` must call `sql.acquireKeyed(...)` / `sql.acquireKeyedWithIsolation(...)` instead. - `WriteFencePlan` gains a `"row"` arm: `{ kind: "row"; drain; conflict; sql }`. An external exhaustive `switch` on `WriteFencePlan["kind"]` (or its `default` branch, if any) now sees this case too. - The base schema gains the `typegraph_fences` relation on both dialects, and `ResolvedSqlTableNames` — the fully-resolved table-name set `createSqlSchema` returns and `GraphBackend.tableNames` exposes — gains its required `fences` member. Code that builds a complete `ResolvedSqlTableNames` object literal by hand, rather than through `createSqlSchema` or a bundled backend factory, must add it; the corresponding `SqlTableNames` input field stays optional and defaults, so a caller only overriding table names is unaffected. A database migrated before this release needs the regenerated migration SQL run against it before declaring `writeFence.mechanism: "row"` (the two bundled mechanisms, `"advisory"` and `"engine-serialized"`, do not need it and keep working unmigrated). - `capabilities.execution.unitOfWork` gains `"optimistic-retry"` as a possible value — an external exhaustive `switch` on it now sees this case too. - `SqlExecutionAdapter` gains an optional `serializationFailure` member — additive, but an object satisfying this interface structurally (rather than by declaring it) may need updating if it re-implements the full member list explicitly. - `FenceSql`'s three members — `lockTables`, `advisoryLockExpression`, `isolationFactExpression` — are now all optional: which ones a `row`-mechanism target supplies differs from an `advisory`-mechanism one. A caller that reads one of these members directly, rather than through a resolved plan's `sql.acquireKeyed` / `sql.acquireKeyedWithIsolation` / `sql.isolationFact` / `sql.lockTables`, must narrow for `undefined` before calling it; the resolved-plan accessors already refuse with a named-member error when a mechanism does not supply one. Bisect note: the intermediate commit adding the `row` write-fence mechanism is red on one construction-inventory ratchet, fixed by the commit that follows it in the same PR; bisecting between the two will find that known-red state. - [#607](https://github.com/nicia-ai/typegraph/pull/607) [`e966b30`](https://github.com/nicia-ai/typegraph/commit/e966b3049755fb4945a7120319f5a9726abab1b0) Thanks [@pdlug](https://github.com/pdlug)! - Add a new entrypoint, `@nicia-ai/typegraph/adapters/drizzle/engine`, exporting `createSqlBackend` and the `SqlEngineProfile` types. `createPostgresBackend` and `createSqliteBackend` are now each `createSqlBackend` applied to a profile built by `buildPostgresEngineProfile` / `buildSqliteEngineProfile`. Emitted SQL, capabilities, marks, transaction framing, and error paths are unchanged for every configuration the two factories accepted before. Two construction-time narrowings apply to the bundled factories as well as to third-party profiles, because both now run through `createSqlBackend`: - A backend whose resolved capabilities carry no `writeFence` declaration is refused with a `ConfigurationError` (`ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION`) that prints the one declaration line to add. Omitting `writeFence` from a `capabilities` override is unaffected (the factory's own declaration applies); passing `capabilities: { writeFence: undefined }` explicitly, which previously built a backend that resolved every write fence through a dialect fallback, now throws at construction. - A backend whose declaration resolves `unfenced` no longer earns the schema-fenced-insert eligibility mark, so a schema-managed first write on it now refuses with `WRITE_FENCE_UNAVAILABLE` instead of fusing the insert. Schema commits on such a backend already refused, so a working configuration is unaffected. - [#629](https://github.com/nicia-ai/typegraph/pull/629) [`02152da`](https://github.com/nicia-ai/typegraph/commit/02152da57cf26a00cf23c96e4e0a95e218fd1d06) Thanks [@pdlug](https://github.com/pdlug)! - `capabilities.writeFence` is the write-fence declaration, a discriminated union on `mechanism`: `{ mechanism: "advisory"; drain: "table-lock" | "quiescent" | "none" }` | `{ mechanism: "engine-serialized" }` | `{ mechanism: "caller-serialized" }`. `mechanism` is the exclusion primitive a backend provides; `drain` — a field of the `"advisory"` shape only — is the separate fact of whether a caller that already excluded other writers can additionally take a relation-wide lock on a resource a few sites protect. Both bundled backends declare it directly (`SQLITE_CAPABILITIES`: `{ mechanism: "engine-serialized" }`; `POSTGRES_CAPABILITIES`: `{ mechanism: "advisory", drain: "table-lock" }`), so nothing built against them changes: same emitted SQL, same resolved plan. `resolveWriteFencePlan` validates a declared `writeFence` at runtime — an unrecognized `mechanism`, an unrecognized `drain`, or a `drain` attached to a serialized mechanism — and refuses with `WRITE_FENCE_DECLARATION_INVALID` naming the field and (where applicable) the accepted values, since a plain-JavaScript backend author is not held to the discriminated-union type the way a TypeScript caller is. A new arm, `{ kind: "caller-serialized" }`, joins `WriteFencePlan`'s union for a deployment-level promise that no other client writes to the backend's database while it is open. `createPostgresBackend` accepts `writeFence: { mechanism: "caller-serialized" }` — a claim about the deployment, not the engine — while continuing to refuse `mechanism: "engine-serialized"` outright, since that claims the engine itself serializes writers. The promise splits into two halves. In process, TypeGraph enforces its own half: every root member the backend classifies in a mutation-capable class — graph-entity and sidecar writes, backend-owned bulk import, derived-data maintenance, schema commits, table/DDL provisioning, `clearGraph`, and the raw-SQL members that can carry an arbitrary write (`execute`, `executeRaw`, `executeStatement`, `executeTemporaryStatement`) — plus `transaction` and `transactionWithNative`, runs through one per-backend serialized queue, so two concurrent calls through the same pool cannot race each other; a root write awaited from inside a `store.transaction` callback is refused (`SERIALIZED_QUEUE_REENTRANT_SUBMISSION`) rather than left to deadlock. Adopting an externally owned transaction (`adoptTransaction`, backing `store.withTransaction(externalTx)`) is refused outright (`CALLER_SERIALIZED_REFUSES_ADOPTION`): its lifetime belongs to the caller, not to this backend's queue, so there is no honest way to hold a queue slot open for it. Outside the process, the deployment still has to hold up its half (no other client writing to the same database) since TypeGraph cannot observe that. `requireWriteFence` takes `requires: "keyed" | "drain"`: `"keyed"` is satisfied by every non-`unfenced` arm; `"drain"` refuses only when the resolved plan's `drain` is `"none"`, and `"engine-serialized"` / `"caller-serialized"` satisfy it without consulting `drain` at all. ## Breaking - `capabilities.pessimisticLocks` and its `PessimisticLockCapabilities` type are removed — declare `capabilities.writeFence` instead. - `requireWriteFence`'s `requires` parameter is renamed: `"advisory-lock"` becomes `"keyed"`, `"table-lock"` becomes `"drain"`. - `WriteFencePlan`'s `lock` arm drops `tableLocks` and `advisoryLocks` — read `drain` instead (`"table-lock"` means what `tableLocks: true` used to). - `WriteFencePlan`'s `unfenced` arm drops `reason` — declaring `writeFence` leaves no shape that resolves `unfenced` for a reason other than an absent declaration, so there is nothing left to distinguish. - `WriteFencePlan` gained the `caller-serialized` arm as a permanent part of the union — an external exhaustive switch on `WriteFencePlan["kind"]` must add a case for it (or its `default` branch, if any, now sees it too). - `WRITE_FENCE_DECLARATION_CONFLICT` is removed — `writeFence` is the only declaration, so no two declarations can conflict. - `WriteFenceDeclaration` is now a discriminated union on `mechanism`, not one flat shape: `drain` is a field of `{ mechanism: "advisory" }` only. Declaring `drain` alongside `mechanism: "engine-serialized"` or `mechanism: "caller-serialized"` — accepted (and ignored) by earlier commits on this same feature branch — is now refused with `WRITE_FENCE_DECLARATION_INVALID`. Read `declaration.drain` only after narrowing `declaration.mechanism === "advisory"`. - [#623](https://github.com/nicia-ai/typegraph/pull/623) [`04ce22c`](https://github.com/nicia-ai/typegraph/commit/04ce22cc87ffbf4f3cc44c47408d0ff18198dd31) Thanks [@pdlug](https://github.com/pdlug)! - `GraphBackend` gains an optional `fenceSql` member: the lock spelling a backend supplies alongside `capabilities.writeFence`, as `FenceSql` — three builders, `advisoryLockExpression`, `isolationFactExpression`, and `lockTables`. `resolveWriteFencePlan`'s `lock` arm carries `sql: FenceStatements`: those three plus the standalone `advisoryLock`, `advisoryLockWithIsolation`, and `isolationFact` statements, which `resolveFenceStatements` derives from the two expressions so the portable lock sites and the fused recorded-write fence always spell the same key. Every write-fence lock site consumes `fence.sql.(...)` instead of hand-writing PostgreSQL lock syntax inline. The bundled PostgreSQL spelling is exported as `postgresFenceSql` from `@nicia-ai/typegraph/adapters/drizzle/postgres`. `createPostgresBackend` supplies it automatically; `createSqliteBackend` supplies no `fenceSql` since its fence is `engine-serialized` and takes no lock. A backend that declares `capabilities.writeFence.mechanism: "advisory"` but supplies no `fenceSql` is now refused at construction with a typed `ConfigurationError` (`WRITE_FENCE_SQL_UNAVAILABLE`) naming the member to supply, rather than reaching a lock site with nothing to spell the statement. For both bundled backends the locks taken, their order, and their modes are unchanged, and the PostgreSQL statement text is equivalent: two advisory-lock sites now bind the lock namespace as a parameter instead of an inline string literal (`hashtext` hashes the value either way), and insignificant whitespace in three statements changed with the move. Two behavior changes reach custom `dialect: "postgres"` backends. A backend declaring `writeFence.mechanism: "advisory"` without `fenceSql` is refused at construction (above). A backend declaring only `mechanism: "engine-serialized"` and no `fenceSql` is refused when a history-capturing transaction reads its isolation level, which previously ran a hard-coded `current_setting('transaction_isolation')` read; supply `fenceSql` (or `postgresFenceSql`) to restore it. ### Patch Changes - [#662](https://github.com/nicia-ai/typegraph/pull/662) [`5fc8e84`](https://github.com/nicia-ai/typegraph/commit/5fc8e84efc9faa45a10ba49375bd211b1b05d4a7) Thanks [@pdlug](https://github.com/pdlug)! - Fix a PostgreSQL availability defect in constrained edge writes: a `bulkCreate` or `bulkInsert` on an edge kind declaring `cardinality: "one"`, `"oneActive"`, or `"unique"` could run for minutes, grow past two gigabytes of server memory, and ignore cancellation. The three statements of the atomic edge-claim program were built with one predicate arm per proposed row — each arm carrying its own two `EXISTS` and one `NOT EXISTS` subquery — so a chunk the bind budget permits (thousands of rows) asked the executor to initialize and evaluate thousands of subplans, with per-arm cost that was not constant. A 200-arm statement took seconds to plan and execute against an empty match set; a 2000-arm statement did not finish. Each statement now drives from a single `proposed` relation of the chunk's rows, so the engine plans it once and the per-row work is an index probe. The same shape lands on both bundled backends. The stale-claim release additionally becomes robust to the plan degradation reported under stale statistics after a bulk load: its driving relation is now the bounded proposed set joined to the claim relation on its primary key, rather than a claim scan whose correlated `NOT EXISTS` the planner could demote to a whole-graph nested-loop anti-join. The claim predicates — what a competing live edge is, and whether a recorded claim holder still satisfies its axis — now have one owner each, rendered from an explicit value source so the single-row and batched statements cannot drift. The conditional takeover statement, which previously spelled the holder-liveness predicate a second time inline, resolves it through that owner. **Author-facing:** `CommonOperationStrategy.buildDeleteStaleAtomicEdgeClaims`, `.buildAcquireAtomicEdgeClaims`, and `.buildAssertAtomicEdgeClaimsOwned` now return `readonly SQL[]` rather than `SQL`. A chunk renders one statement per distinct declared cardinality it contains — one statement in the ordinary single-kind case — because the endpoint terms an axis key covers and the liveness a holder must still satisfy are predicate shape, not values, and folding them into the relation as guarded terms would give back the index probes this change exists to gain. No action is required of a backend author: `CommonOperationStrategy` is visible on the engine entrypoint as part of `SqlEngineProfile`'s shape, but it is not author-supplyable — `strategy` is absent from `DERIVABLE_ENGINE_PROFILE_KEYS`, so `deriveEngineProfile` refuses it, and `SqlEngineProfile.assembly` is branded with a non-exported symbol, so a profile cannot be built from a literal either. The bundled builders are the only source of a strategy, and theirs changed with the code. - [#627](https://github.com/nicia-ai/typegraph/pull/627) [`4cec2b0`](https://github.com/nicia-ai/typegraph/commit/4cec2b0096d5fad833d32ce8b5ed6fb3978d1dc2) Thanks [@pdlug](https://github.com/pdlug)! - The PostgreSQL backend's database-extension install now resolves the write-fence plan and spells its advisory lock through the backend's `fenceSql`, like every other lock site, instead of hardcoding `pg_advisory_xact_lock(hashtext(...), 0)` inline. A custom or derived PostgreSQL profile whose resolved plan is `engine-serialized` or `unfenced` installs extensions without taking that lock and relies solely on the duplicate-key retry, which was already the fence's correctness owner in that case. The bundled PostgreSQL backend's behavior and emitted SQL are unchanged. ## 0.56.0 ### Highlights TypeGraph 0.56 adds a durable review workflow for candidate write sets. `planCandidateWriteSetReview()` captures immutable, digest-checked evidence for the candidate, original plan, policy context, normalized options, and target baseline. After an application records approval, `revalidateCandidateWriteSetReview()` compares that evidence with current graph state and returns a structured compatible, changed, or incompatible result together with a fresh revision-fenced plan when execution remains safe. Reviewed plan application can now share one protected transaction with application-owned checks and writes. Optional `beforeApply` and `afterApply` callbacks run after the target fence is validated, use transaction-bound typed collections, and roll back with the merge if any step fails. Transaction-conflict retries replay the complete operation, including both callbacks. Generic code can dispatch edge operations by kind through `DynamicEdgeCollection` while preserving the selected edge's property and result types and validating endpoint pairs at runtime. Transactions gain `getEdgeCollection` and `getEdgeCollectionOrThrow`, valid-time views support pinned dynamic edge reads, and generic traversal factories now preserve declared target kinds across array targets, source-dependent maps, and unions. ### Upgrade notes - Replace generic `tx.edges[kind]` access and uncorrelated edge collection casts with `tx.getEdgeCollectionOrThrow(kind)`. Concrete `.edges.` access retains compile-time endpoint-pair checking. - Known-kind dynamic edge lookups now enforce their property schema at compile time. Parse unvalidated records before passing them, and update hand-authored transaction or view mocks with the new lookup methods. - `beforeApply` and `afterApply` callbacks may run again after a transaction conflict, so callback behavior must be safe to retry. Callback failures roll back the combined transaction; optional provenance persistence remains a separate post-commit operation. - Candidate review evidence is application-authenticated and V1 uses a conservative whole-target baseline. Applications must version opaque policy and callback dependencies and decide whether a compatible revalidation still satisfies their approval policy. ### Minor Changes - [#615](https://github.com/nicia-ai/typegraph/pull/615) [`4574be7`](https://github.com/nicia-ai/typegraph/commit/4574be7f5b7fb91428639a9bf3056801f8f15bbf) Thanks [@pdlug](https://github.com/pdlug)! - Add durable candidate merge reviews with `planCandidateWriteSetReview()` and `revalidateCandidateWriteSetReview()`. Persist immutable review and approval evidence in the target graph, then compare the retained candidate against current state before applying a fresh revision-fenced plan. Structured compatibility results expose changes requiring review without weakening atomic apply-time concurrency or constraint checks. - [#611](https://github.com/nicia-ai/typegraph/pull/611) [`cf126e3`](https://github.com/nicia-ai/typegraph/commit/cf126e3542f585eeb18d8bc784e769b224de4ed5) Thanks [@pdlug](https://github.com/pdlug)! - Support generic edge dispatch through `DynamicEdgeCollection` and graph-aware dynamic lookups. Edge property and result types are preserved while endpoint pairs are validated at runtime. Transactions now expose `getEdgeCollection` and `getEdgeCollectionOrThrow`, including scoped receipt accounting; valid-time views expose pinned dynamic edge reads. Migration: replace generic `tx.edges[kind]` calls or uncorrelated collection casts with `tx.getEdgeCollectionOrThrow(kind)`. Concrete `.edges.` calls retain compile-time pair checking. Known-kind dynamic lookups now enforce their property schema at compile time; parse unvalidated records before passing them. Broad `EdgeRegistration` annotations continue to permit array or map targets; narrow the target shape when inspecting it. Hand-authored transaction and view mocks must provide the new lookup methods. Fix outgoing traversal target inference in graph factories that accept generic node types. Preserve the declared target kinds for both array targets and source-dependent maps, including unions of edge kinds. - [#614](https://github.com/nicia-ai/typegraph/pull/614) [`5232be5`](https://github.com/nicia-ai/typegraph/commit/5232be5dd1cc465dd7fc6569d391d9fd52e517eb) Thanks [@pdlug](https://github.com/pdlug)! - Compose reviewed merge-plan application with application-owned graph checks and writes using optional `beforeApply` and `afterApply` callbacks. Prechecks receive transaction-bound read-only collections after the target fence is validated; post-apply work uses typed graph operations before the same transaction commits. Failures roll back the combined operation, and transaction-conflict retries replay both callbacks. ## 0.55.0 ### Highlights TypeGraph 0.55 adds source-dependent edge targets to the existing `from`/`to` syntax. One edge kind can now permit Employee → Department and Student → Course without admitting the cross-pairs. TypeScript write inference preserves that relationship, runtime validation rejects invalid pairs with `EndpointPairError`, and restrictions survive schema serialization, imports, graph merges, and runtime graph extensions. Store node and edge creation and upsert inputs now accept `validFrom: null` to explicitly request an open-left validity window. Snapshot and incremental graph merges preserve those windows through canonicalization, edge repointing, and serialized merge plans instead of narrowing them to the merge commit time. Bulk-upsert coalescing also distinguishes a confirmed absent lower bound from a creation timestamp that has not yet been assigned. Startup identity repair now checks the schema version used to derive its registry before rebuilding closure. If a concurrent migration has committed newer semantics, the stale repair raises `StaleVersionError` and leaves the newer closure intact. This covers both existing-relation repair and creation plus population of missing derived relations. ### Upgrade notes - Existing array-valued `to` declarations retain their Cartesian-product semantics. Removing allowed pairs from a source-dependent target map is classified as a breaking schema change. - Omitted `validFrom` retains its existing defaults. Explicit `null` requests no lower bound, while returned metadata still uses `undefined` for an absent bound. Live-row upserts retain their immutable-bound rules; `validTo` continues to use its existing set/clear protocol. - If startup identity repair raises `StaleVersionError` after a concurrent migration, reopen with the current graph definition before retrying. ### Minor Changes - [#609](https://github.com/nicia-ai/typegraph/pull/609) [`39a65ef`](https://github.com/nicia-ai/typegraph/commit/39a65ef8eb544942f35f97fdbfbc2609d6284bba) Thanks [@pdlug](https://github.com/pdlug)! - Accept `validFrom: null` on Store node and edge creation and upsert inputs to explicitly request an open-left validity window. Omitted lower bounds keep their existing default behavior; live-row upserts still refuse changes to an immutable lower bound. Preserve open-left staged node and edge windows through snapshot and incremental graph merges, edge repointing, and serialized merge plans instead of narrowing them to the merge commit time. Keep repeated bulk-upsert coalescing from confusing an unknown creation timestamp with a confirmed open-left bound. - [#605](https://github.com/nicia-ai/typegraph/pull/605) [`f7dcd17`](https://github.com/nicia-ai/typegraph/commit/f7dcd1708be8dd3ef67ce2b32a1c2ee22536d2f6) Thanks [@pdlug](https://github.com/pdlug)! - Define source-dependent edge targets with a `to` map, such as `from: [Employee, Student]` with `to: { Employee: [Department], Student: [Course] }`. This allows one edge kind to connect specific source/target pairs without admitting every combination. Array-valued `to` declarations retain their existing Cartesian-product behavior. Allowed pairs are preserved in typed writes, runtime validation, schema serialization, imports, graph merges, and runtime graph extensions. Invalid pairs produce `EndpointPairError`, removing allowed pairs is a breaking schema change, and ontology compatibility checks account for the pair relationship. See [source-dependent targets](https://typegraph.dev/core-concepts#source-dependent-targets) for examples and lifecycle rules. ### Patch Changes - [#608](https://github.com/nicia-ai/typegraph/pull/608) [`5b2dcc0`](https://github.com/nicia-ai/typegraph/commit/5b2dcc07bd84d3fe97959a3c86cab4e29449c5f5) Thanks [@pdlug](https://github.com/pdlug)! - Prevent startup identity repair from overwriting a newer schema migration with closure data derived from an older schema. Repair now checks the observed schema version inside its write transaction and raises `StaleVersionError` if a concurrent migration advanced it, preserving the newer closure. ## 0.54.0 ### Highlights TypeGraph 0.54 makes runtime-evolved schemas first-class across metadata, typing, analysis, conflict planning, and guarded recovery. Graph-scoped annotations now flow through `defineGraph`, `defineGraphExtension`, `SerializedSchema`, and introspection. Schema-bound runtime-kind tokens preserve extension node and edge types through collections, bulk edge reads, and one-hop traversals without consumer-side widening adapters. New current-state Store analysis APIs keep discovery work bounded as vocabularies and datasets grow. `store.describe()` computes per-kind population and declared-property coverage in SQL, splits node and edge work, and chunks wide schemas under a fixed result-column budget. `store.validateStore()` uses keyset-paginated scans to report declared-schema violations with record ids, JSON-pointer paths, and reasons while treating undeclared fields as healthy semi-structured data. Both APIs bracket their work with a stable schema coordinate; data is live, so concurrent writes can affect different statements or pages. `planCandidateWriteSet()` brings the existing source-attributed merge planner to branch-free candidate writes without applying them. Node collections also gain scalar `compareAndSet()` and `compareAndSetAbsent()` guards that preserve ordinary validation, hooks, history, and version bookkeeping while enforcing the expected-current predicate atomically. ### Upgrade notes - `BulkOperationHookContext["operation"]` now includes `"compareAndSet"`; exhaustive consumers must handle the new case. - Store analysis is current-only and intentionally absent from `StoreView` and transaction callback facades. Calling `store.describe()` through an enclosing root Store inside a transaction callback does not enlist the analysis in that transaction. - During a mixed-version rollout, upgrade every process that can write schema versions to 0.54 before enabling graph-scoped annotations. TypeGraph 0.54 preserves unknown top-level serialized-schema fields during reconciliation; older writers do not provide that guarantee. ### Minor Changes - [#601](https://github.com/nicia-ai/typegraph/pull/601) [`f7d1ac4`](https://github.com/nicia-ai/typegraph/commit/f7d1ac44d105f6b1aedabb4c70917dbc57570c1a) Thanks [@pdlug](https://github.com/pdlug)! - Add first-class graph annotations, Store population statistics and validation, branch-free candidate write-set conflict planning, schema-bound runtime-kind tokens for typed collections and traversals, and guarded node compare-and-set. This release also preserves unknown top-level serialized-schema fields across reconciliation, reports no-op schema diffs without a phantom version delta, ships the package changelog in the npm tarball, and clarifies base-schema adoption and transaction-scoped mutation performance. The bulk-operation hook context now also reports `"compareAndSet"`; exhaustive consumers of its `operation` field must handle the new case. ## 0.53.0 ### Highlights TypeGraph 0.53 substantially reduces database exchanges for write-heavy applications running near Neon, Cloudflare D1, libSQL, or PostgreSQL. Compared with the former 5–6-exchange managed-write shape, eligible singleton updates and deletes now use one authoritative read or gate plus one atomic mutation submission, about 60–67% fewer exchanges. Eligible bulk creates, inserts, deletes, complete-document replacements, and durable edge convergence submit their complete mutation unit once; eligible upserts use one batched preimage read plus one atomic submission. These reductions apply to root writes outside `store.transaction()`; transaction-scoped writes remain on the interactive path and see no exchange-count change. These are transport-submission counts rather than wall-clock benchmarks. See the [performance guide](https://typegraph.dev/performance/overview) for the exact eligibility envelopes and fallback behavior. The atomic programs now carry uniqueness and disjointness claims, full-text and vector projections, endpoint and schema fences, postimage assertions, and typed rollback diagnosis in the same submission. Large D1 upserts remain atomic across bind-sized statements, with supported ceilings raised to 512 nodes and 187 edges per call. PostgreSQL transaction sessions use the same programs through a savepoint, while unsupported shapes continue through the complete portable path. This release also adds [`nodes..bulkReplaceById()`](https://typegraph.dev/schemas-stores#bulkreplacebyiditems) for read-free, complete-document replacement; makes graph-template registration and instantiation safe across serverless isolates; repairs externally provisioned pre-0.52 base storage through a numbered deployment-wide lifecycle; and publishes transport plus semantic conformance runners for custom atomic backends. ### Upgrade notes - Existing or externally provisioned databases must be opened once with privileged `createStoreWithSchema()` / `createAdapterStoreWithSchema()`, or receive TypeGraph's generated base-schema migration, before any DML-only runtime opens. Applications that already use either privileged opening path adopt the base schema automatically during that open; no separate bootstrap command is needed. The deployment invariant is ordering: privileged adoption must finish before DML-only workers start. See [Upgrading deployment-wide base storage](https://typegraph.dev/backend-setup#upgrading-deployment-wide-base-storage). - The removed top-level `capabilities.transactions` override is now refused instead of ignored. Use `capabilities.execution.interactiveTransactions`. - `store.transaction()` and `store.transactionWithReceipt()` now refuse backends without interactive transaction support. Call ordinary Store write methods directly when the application intentionally owns partial-failure recovery. - Eligible singleton updates now converge optimistically and retry a moved preimage up to four times. Under sustained same-row contention they can throw `DatabaseOperationError` where an interactive backend previously serialized the writers. - A PostgreSQL-dialect backend that declares neither usable pessimistic locks nor serialized writers can no longer open a schema-managed store. TypeGraph now reports `WRITE_FENCE_UNAVAILABLE` rather than accepting a declaration it cannot honor. First-party PostgreSQL, Neon, PGlite, SQLite, D1, and libSQL configurations retain their existing fence posture. See [Write fence declaration](https://typegraph.dev/backend-setup#write-fence-declaration-pessimisticlocks). ### Minor Changes - [#574](https://github.com/nicia-ai/typegraph/pull/574) [`e60f77c`](https://github.com/nicia-ai/typegraph/commit/e60f77cd2f2af0653193021a9dec7e6a6bf35aaf) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible durable `bulkGetOrCreateByEndpoints()` calls as one schema-fenced native atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. The eligible envelope is a schema-declared `matchIdentity` with `cardinality: "many"`, the declaration's match fields, default `ifExists: "return"`, and no temporal mutation. The program owns endpoint validation, identity arbitration, input-order restoration, and whole-call rollback. A tombstoned winner rolls the native attempt back and refuses transactionless convergence with the typed `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` (`edgeMatchKeyConvergence`) error; use a transactional backend when schema-aware resurrection is required. Dynamic match fields, update mode, constrained cardinality, temporal options, caller-owned transactions, derived/custom backends, and history/revision stores retain their existing path. The existing read-only fast path remains available outside the native envelope: an all-live default-`"return"` batch completes from one set-oriented root read without opening a confirmation transaction. Inside the native envelope, the authoritative upsert program establishes the same logical `"found"` result in one exchange but can take incumbent-row locks and produce write amplification. The libSQL transport inventory measures one `batch` submission and zero `execute` calls for a multi-item eligible call. This is a transport submission-count measurement, not a wall-clock RTT benchmark. - [#588](https://github.com/nicia-ai/typegraph/pull/588) [`a8cd7b0`](https://github.com/nicia-ai/typegraph/commit/a8cd7b07c09e706542f0845d144f5a618b5f9488) Thanks [@pdlug](https://github.com/pdlug)! - Add `nodes..bulkReplaceById()` for complete-document replacement by ID. Eligible bundled SQLite, D1, libSQL, Neon HTTP, and PostgreSQL transaction-session roots execute missing-row creation, live-row replacement, tombstone resurrection, claims, and search projections as one read-free atomic submission; unsupported shapes retain the complete portable path. - [#571](https://github.com/nicia-ai/typegraph/pull/571) [`b57fb90`](https://github.com/nicia-ai/typegraph/commit/b57fb90576349aa65586e8d439e251bc8661d636) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible mixed create/update `bulkUpsertById()` calls as one atomic mutation exchange after their batched authoritative read, preserving whole-set rollback on transactionless bundled roots. - [#562](https://github.com/nicia-ai/typegraph/pull/562) [`57245c2`](https://github.com/nicia-ai/typegraph/commit/57245c22cb12da8e806b650c63c0caeedbc6928d) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible plain schema-managed `nodes.bulkInsert` and `nodes.bulkCreate` batches as one schema-fenced native atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. Claim-free nodes accept generated IDs, caller-supplied IDs, or mixed batches. A claimed node is eligible only for generated-ID batches with exactly one same-kind (`scope: "kind"`) uniqueness constraint, subject to the backend's claimed-member budget. `bulkCreate` restores rows in input order. Operational Identity, projections, history, revision, caller-supplied or mixed IDs in claimed batches, multiple or non-kind uniqueness constraints, disjointness, and other unsupported forms retain their existing transaction or fallback path. The eligible program preserves live-duplicate rollback, tombstone resurrection validity semantics, bind-budget chunk rollback, and schema-fence errors. The libSQL transport inventory measures one `batch` submission and zero `execute` calls for both generated claim-free and generated single-claim node batches; this is a submission-count measurement, not a wall-clock RTT benchmark. - [#582](https://github.com/nicia-ai/typegraph/pull/582) [`83b8bd0`](https://github.com/nicia-ai/typegraph/commit/83b8bd07fe0bece1a721c9c0418e6b39be1f9a59) Thanks [@pdlug](https://github.com/pdlug)! - Fold single disjointness claims and owner-side node claim cleanup into bundled atomic mutation programs. Eligible `nodes.bulkInsert()` and `nodes.bulkCreate()` calls for a `disjointWith` kind now retain one schema-fenced transport submission on Neon HTTP, Cloudflare D1, and libSQL, including legacy live-row detection and typed `DisjointError` rollback. Restricted node deletes release owned uniqueness and disjointness claims inside the same atomic program, and update-only upserts no longer fall back merely because their kind participates in disjointness. Custom mutation executors advertise claim support explicitly through `claimSupport.families`, the per-member `claimSupport.maxInputCostPerEntry` bound, and `releasedClaimFamilies`; omitted claim families remain on the portable path and an empty family list with a zero bound is an honest opt-out. - [#583](https://github.com/nicia-ai/typegraph/pull/583) [`0abfdc6`](https://github.com/nicia-ai/typegraph/commit/0abfdc6e5c05c5575fe5fe558ae9889450a0c29e) Thanks [@pdlug](https://github.com/pdlug)! - Fold eligible node fulltext and vector projection transitions into the same schema-fenced atomic programs as `bulkInsert()`, `bulkCreate()`, and resolved node mutations. Bundled Neon HTTP, Cloudflare D1, and libSQL roots keep projected creates and updates at one mutation exchange, while PostgreSQL session programs keep the complete row-and-sidecar unit on their pinned transaction. The atomic program proves the exact durable contribution markers in that same submission, including on a newly constructed backend. Custom atomic node executors can opt in per projection family through validated `projectionSupport` metadata; omitted support fails closed to the portable path. - [#581](https://github.com/nicia-ai/typegraph/pull/581) [`2c8f97f`](https://github.com/nicia-ai/typegraph/commit/2c8f97ff2e108df33caa5985b23ca54d50b36007) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible mixed node and edge `bulkUpsertById()` mutation sets through atomic programs bound to the exact collection-opened, caller-supplied, or adopted PostgreSQL transaction. The operation now returns an explicit `applied | unsupported` verdict, where `unsupported` is emitted before program SQL and safely re-enters the complete portable path. Transaction-session programs use a savepoint so typed refusal diagnosis remains available without committing or poisoning the caller's surrounding transaction. - [#578](https://github.com/nicia-ai/typegraph/pull/578) [`7c7e328`](https://github.com/nicia-ai/typegraph/commit/7c7e3281721443ee580ba113b3b2407f5a1999aa) Thanks [@pdlug](https://github.com/pdlug)! - Route eligible singleton node and edge updates and soft deletes through the same exact-root atomic mutation programs as resolved bulk writes. These operations preserve per-item hooks and fallback semantics while replacing explicit managed transactions with one authoritative read or gate plus one guarded atomic mutation exchange on bundled serverless roots. Eligible singleton updates now converge optimistically rather than holding a write transaction across their read and write. They retry a moved preimage up to four times before throwing `DatabaseOperationError`; under sustained same-row contention this can refuse an update that a transaction-capable backend previously serialized. Transaction-scoped calls and ineligible shapes retain the interactive transaction path. - [#574](https://github.com/nicia-ai/typegraph/pull/574) [`e60f77c`](https://github.com/nicia-ai/typegraph/commit/e60f77cd2f2af0653193021a9dec7e6a6bf35aaf) Thanks [@pdlug](https://github.com/pdlug)! - Add independent execution capability declarations and a framework-agnostic atomic transport conformance runner for backend authors. It verifies ordered result slots, exact statement and parameter forwarding, empty-batch no-op behavior, later-statement rollback without primary or sidecar leakage, and caller-supplied exact-root provenance checks. Generic registration certifies transport mechanics but does not by itself authorize bundled node or edge mutation programs; semantic eligibility remains a separate fail-closed contract. Bundled factories now refuse the removed top-level `capabilities.transactions` override with migration guidance instead of silently retaining it as inert data; use `capabilities.execution.interactiveTransactions`. - [#585](https://github.com/nicia-ai/typegraph/pull/585) [`2e878bc`](https://github.com/nicia-ai/typegraph/commit/2e878bcd740a6fc36856deb0745c24890293bd10) Thanks [@pdlug](https://github.com/pdlug)! - Batch committed-state diagnosis after a refused atomic node-claim program. Bundled backends now recover typed uniqueness and disjointness errors with set-oriented reads instead of sequential per-member probes, while custom scalar fallbacks use bounded windows, stop once the earliest refusal is known, and retain deterministic error selection. - [#586](https://github.com/nicia-ai/typegraph/pull/586) [`51da9c0`](https://github.com/nicia-ai/typegraph/commit/51da9c05eeaddbdd45ee0283e48b81b8322f5929) Thanks [@pdlug](https://github.com/pdlug)! - Keep eligible large node and edge `bulkUpsertById()` sets on the atomic mutation-program path when they exceed one statement's bind budget. Resolved updates, mixed creates and updates, projection sidecars, per-chunk postimage assertions, and ordered result reads now execute as bind-sized statements inside one bounded atomic transport submission. On Cloudflare D1 this raises the previous 17-node and 6-edge batch-wide ceilings to 512 nodes and 187 edges without weakening rollback: a moved member in any chunk aborts every sibling chunk. Larger sets fail closed to the portable path rather than constructing an unbounded transport request. - [#584](https://github.com/nicia-ai/typegraph/pull/584) [`69d5930`](https://github.com/nicia-ai/typegraph/commit/69d5930ef50a0db74c86db725e2c342b9bd478a3) Thanks [@pdlug](https://github.com/pdlug)! - Compose eligible node claim sets and fulltext/vector projections inside the same schema-fenced atomic `bulkInsert()` and `bulkCreate()` program. Multiple uniqueness constraints, hierarchy-wide uniqueness scopes, disjointness claims, caller/generated ID mixtures, and claim-plus-projection members now retain one Neon HTTP, Cloudflare D1, or libSQL transport submission and the same pinned PostgreSQL transaction program. Legacy claim-axis conflicts keep their typed errors and roll back every row and projection sidecar. Claimed node programs now chunk row statements by actual per-member claim work inside one atomic submission instead of imposing a batch-wide claim ceiling. The custom-backend claim contract is correspondingly simplified from family-scoped batch ceilings to `claimSupport.families` plus `claimSupport.maxInputCostPerEntry`; backend authors use the exported `atomicNodeClaimInputCost()` owner rather than reproducing the compiled SQL cost model. Claim refusal is enforced by a terminal database assertion so every row and projection chunk rolls back before failure-only committed-state reads recover the portable path's typed diagnostic. - [#563](https://github.com/nicia-ai/typegraph/pull/563) [`589a38d`](https://github.com/nicia-ai/typegraph/commit/589a38d10d2c2d1a94002a5a12b4c19f5ce1fb32) Thanks [@pdlug](https://github.com/pdlug)! - Make graph-template registration DML-only after normal TypeGraph bootstrap, allow embedding-bearing templates on backends configured with `vector: false`, and clone durable runtime-contribution markers during instantiation so targets can be reopened through verified stores from a later serverless isolate. Vector-enabled backends continue to refuse schema-only templates that would require graph-scoped vector storage. - [#577](https://github.com/nicia-ai/typegraph/pull/577) [`967c098`](https://github.com/nicia-ai/typegraph/commit/967c0985757bd774e420ba7b4f679744e8af43f8) Thanks [@pdlug](https://github.com/pdlug)! - Strengthen custom-backend atomic conformance so transport registration is immutable, runners bind the exact registered transport and semantic profile, observe dispatch inside property-preserving executor wrappers, verify author-created backend lineage, report inapplicable transaction checks as skipped, validate case bindings before fixture preparation, and distinguish native from pre-dispatch semantic refusals. Bundled libSQL now self-certifies every semantic mutation variant against complete committed database state: direct root families on the backend its public interactive factory returns, and resolved update or mixed families on a transactionless root where the Store can actually dispatch them. - [#579](https://github.com/nicia-ai/typegraph/pull/579) [`8f890c7`](https://github.com/nicia-ai/typegraph/commit/8f890c7f3e55a595f5cba16c9c8bd27ca0b5fa25) Thanks [@pdlug](https://github.com/pdlug)! - Enable the registered atomic SQL and mutation-program profile on recognized interactive PostgreSQL drivers. TypeGraph now executes eligible bulk creates, bulk deletes, singleton updates/deletes, and durable edge convergence on one pinned transaction instead of falling back to the multi-step portable write plan. Neon HTTP keeps its existing native transaction-batch path, while unrecognized PostgreSQL drivers continue to fail closed to the portable implementation. - [#565](https://github.com/nicia-ai/typegraph/pull/565) [`677ee91`](https://github.com/nicia-ai/typegraph/commit/677ee9161057b28e58ac23cc1c4515d53038e826) Thanks [@pdlug](https://github.com/pdlug)! - Existing databases must be opened once through privileged `createStoreWithSchema` / `createAdapterStoreWithSchema`, or receive the published base-schema migration, before zero-DDL verified and graph-template runtime paths are used. Those paths now fail early with `BaseSchemaMigrationError` until deployment-wide base storage is stamped at version 1. Version deployment-wide base storage independently of per-graph schemas. A privileged open adopts the graph-template table and durable edge match-identity storage once, then stamps the marker; later warm opens perform only one marker read. This repairs externally provisioned 0.51 databases even when their graph schema is unchanged. Plain edge writes also classify legacy missing-column failures as `EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE` instead of leaking driver errors. Fresh SQLite and PostgreSQL installation SQL now publishes the current base-schema marker as its final statement, so zero-DDL verified stores and graph-template APIs can attach immediately after applying TypeGraph's generated migration. Concurrent adoption accepts a marker already advanced beyond the step it completed and never downgrades it. - [#570](https://github.com/nicia-ai/typegraph/pull/570) [`5432598`](https://github.com/nicia-ai/typegraph/commit/5432598bc3ec851ac53c535f89d0db517336d09d) Thanks [@pdlug](https://github.com/pdlug)! - Reduce eligible update-only node and edge `bulkUpsertById()` calls on bundled serverless backends to one batched preimage read plus one guarded atomic set update. Consolidate fallback updates under one write plan and batch their authoritative reads when transaction semantics allow it. - [#590](https://github.com/nicia-ai/typegraph/pull/590) [`c374a21`](https://github.com/nicia-ai/typegraph/commit/c374a210813e164ccd783add0c96e713fa37ba67) Thanks [@pdlug](https://github.com/pdlug)! - The PostgreSQL schema fence's three consumers now resolve a `WriteFencePlan` instead of emitting their locks unconditionally, bringing the last lock sites in `createPostgresBackend` into the model the rest of the codebase already reads: the per-graph schema-commit fence (`acquireSchemaWriteFence`: `pg_advisory_xact_lock` plus `SELECT ... FOR UPDATE` on the active schema row), the managed writer's `FOR SHARE` on that same row (`lockActiveSchemaVersion`), and the copy of that `FOR SHARE` the fused managed-insert programs carry inside their own statement (`schemaFenceInsertLockClause`). They are one `FOR UPDATE`/`FOR SHARE` contract, so they must never disagree about whether the engine honors row locks, and the inventory ratchet pins `resolveWriteFencePlan` at 14 call sites rather than 11. `createPostgresBackend` already accepted and validated a `pessimisticLocks` override declaring no locks — it rewrites `serializedWriters` and throws only when that field claims a writer slot — and the eight sites consolidated previously already refuse under it. The schema fence emitted its locks regardless, so the backend accepted a stated capability and then contradicted it. **The two cross-statement fences refuse an `unfenced` backend** rather than running without the lock, because both fence a read-then-write sequence that spans statements. `commitSchemaVersion` reads the row for the incoming version, reads the active version, and only then runs its `deactivateAll` / `activateVersion` pair; its own comment names this fence as what serializes that. A managed write HOLDS that `FOR SHARE` for the remainder of its transaction, which is what makes the version it just asserted binding through to the writes that follow — a concurrent schema commit's `FOR UPDATE` blocks on it. Skipping either lock does not leave a slower-but-correct path, it leaves a check-then-write window: the version is asserted, and then the flip the assertion was checking for lands before the write. The partial unique index on `(graph_id) WHERE is_active` still refuses a second active row, but it cannot order two commits that each read a version the other is about to replace. `engine-serialized` is exempt through `requireWriteFence` as everywhere else: the writer slot is the fence, which is why SQLite's `lockSchemaVersionForWrite` has always run as an ordinary read. **The in-statement clause degrades**, and it is the only one that may, because its predicate is evaluated inside the INSERT that depends on it. One statement cannot race itself: with an empty clause the fence subquery still yields no row when the expected version is no longer active, so the INSERT still writes nothing. That is the posture SQLite has always run this path in. It is emitted as an empty clause rather than dropped, because dropping it (`undefined`) means "this backend has no schema-fenced insert program at all" and sends the Store down the unfused fallback path. No first-party configuration changes behavior: `POSTGRES_CAPABILITIES` declares `{ advisoryLocks: true, tableLocks: true, serializedWriters: false }`, which resolves `lock`, and the emitted SQL is byte-for-byte what it was; SQLite already passed an empty clause. What changes is that a PostgreSQL-dialect backend declaring no usable write fence is now refused at the schema commit with `WRITE_FENCE_UNAVAILABLE`, naming the operation, instead of silently running a schema-managed store without the serialization that store's correctness assumes. - [#575](https://github.com/nicia-ai/typegraph/pull/575) [`5e30ee4`](https://github.com/nicia-ai/typegraph/commit/5e30ee4ec05e40f5fd7419df55261d26fef29837) Thanks [@pdlug](https://github.com/pdlug)! - Open exact-root atomic Store mutation programs to custom backends through an explicit per-family semantic registration profile. Transport-only registration continues to enable no Store fast path; derived and transaction-scoped backends inherit neither proof, and malformed or out-of-order registrations fail with typed configuration errors. - [#576](https://github.com/nicia-ai/typegraph/pull/576) [`5622ff4`](https://github.com/nicia-ai/typegraph/commit/5622ff48490208ae471fc8196b2ea6d7037d6f72) Thanks [@pdlug](https://github.com/pdlug)! - Add a framework-agnostic semantic conformance runner for custom atomic mutation programs. The runner requires ordered Store results, independently observed committed state, stale-fence no-write behavior, typed semantic-refusal rollback, exact executor dispatch, and exact-root provenance for every registered family variant. - [#574](https://github.com/nicia-ai/typegraph/pull/574) [`e60f77c`](https://github.com/nicia-ai/typegraph/commit/e60f77cd2f2af0653193021a9dec7e6a6bf35aaf) Thanks [@pdlug](https://github.com/pdlug)! - Make `store.transaction()` and `store.transactionWithReceipt()` fail closed on backends without transaction support instead of invoking callbacks with non-atomic write semantics. Applications that intentionally relied on the old D1 or Neon HTTP fallthrough must call ordinary Store write methods directly and own partial-failure recovery explicitly. - [#569](https://github.com/nicia-ai/typegraph/pull/569) [`ed66e47`](https://github.com/nicia-ai/typegraph/commit/ed66e4703e9a4f9cd6094e8887a0a2e0b8a8e693) Thanks [@pdlug](https://github.com/pdlug)! - Unify bundled root create and delete optimizations behind one exact-root mutation-program profile. Eligible node and edge `bulkDelete()` calls now run as one schema-fenced atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots; the programs preserve edge collection identity, node restricted-delete semantics, stale-schema refusal, bind-budget chunking, and whole-call rollback, while node kinds that owe unique, disjointness, identity, projection, or capture sidecars retain the transactional path. Portable edge bulk deletion now replaces per-ID reads and writes with one batched authoritative read and set-based soft-delete chunks when the backend exposes the existing batch ports. ### Patch Changes - [#580](https://github.com/nicia-ai/typegraph/pull/580) [`43f2e34`](https://github.com/nicia-ai/typegraph/commit/43f2e3427e5f3dc35c0407c563deb5727695f70b) Thanks [@pdlug](https://github.com/pdlug)! - Harden atomic edge-write refusal diagnosis. Bundled backends now diagnose missing endpoints with set-oriented reads instead of a sequential read per edge; custom backends without the batch-point-read capability use bounded-concurrency windows that still cover the complete input. If endpoint or cardinality state changes after the atomic rollback and no current violation can be found, TypeGraph preserves the driver cause in a typed `DatabaseOperationError` instead of leaking an unclassified transport error. Native singleton edge deletes now report the same authoritative `written` hook outcome as node deletes. - [#587](https://github.com/nicia-ai/typegraph/pull/587) [`a3a4315`](https://github.com/nicia-ai/typegraph/commit/a3a4315f2fd87f39d3fc328085f99bf512123f5d) Thanks [@pdlug](https://github.com/pdlug)! - Eliminate the cold contribution-marker read before eligible atomic node projection writes. Bundled atomic programs now prove the exact fulltext/vector marker identities and strategy signatures with additional SQL statements inside the same database submission as the row and projection changes; missing, stale, failed, or unmaterialized evidence rolls the whole program back and is diagnosed through the existing typed contribution errors. This keeps projected writes at one mutation exchange even when a Cloudflare Worker constructs a fresh Neon HTTP, D1, or libSQL backend for each request, while retaining the server-side evidence check. - [#568](https://github.com/nicia-ai/typegraph/pull/568) [`cda8bdd`](https://github.com/nicia-ai/typegraph/commit/cda8bdd72b0ea8568bda892f09a666e3b55a790f) Thanks [@pdlug](https://github.com/pdlug)! - Harden deployment-wide base-schema adoption under concurrent upgrades, make managed SQLite and libSQL installations repair pre-0.52 edge storage without bypassing numbered lifecycle steps, and add ratchets that bind ordered physical DDL and supported provisioning paths to executable match-identity storage rather than trusting the version marker alone. - [#589](https://github.com/nicia-ai/typegraph/pull/589) [`09cc34f`](https://github.com/nicia-ai/typegraph/commit/09cc34f957c8114a45b3252d58fb493a61657fb5) Thanks [@pdlug](https://github.com/pdlug)! - Fence constrained writes inside caller-adopted SQLite transactions by taking the writer slot before decision-driving reads, acquire graph-merge locks in the canonical schema-first order, and compile temporal-system `orderBy` fields against their physical columns when the schema does not declare a same-named property. Declared properties retain precedence so filtering and ordering use the same field. Bulk node and edge wrappers now have regression coverage that pins one durable revision advance per public bulk call rather than one advance per member. ## 0.52.0 ### Highlights TypeGraph 0.52 introduces the authoritative command and atomic SQL-program foundations later expanded in 0.53. Eligible schema-managed creates can fold their fences, constraint decisions, projections, and row write into one authoritative statement, while eligible bulk node and edge creates execute as one bounded atomic submission on bundled Neon HTTP, Cloudflare D1, and libSQL roots. Unsupported dimensions continue through the complete interactive-transaction path or a typed refusal; the optimization does not silently omit requested behavior. Edges can now declare a durable graph-local `matchIdentity`. TypeGraph persists and arbitrates the canonical endpoint-and-property key, maintains it through ordinary and import writers, and uses it to converge eligible `getOrCreateByEndpoints()` calls without a read-then-insert race. This release also adds durable schema-only graph templates through `registerGraphTemplate()` and `instantiateGraphTemplate()`. ### Upgrade notes - Custom `GraphBackend` implementations must provide the required `commands: { session, execute }` port and move managed node create, edge create, and edge convergence behavior out of the removed specialized hooks. A command that cannot honor a requested dimension must return its typed `unsupported` result before executing SQL. - `OptionalTransactionExecution.atomic` is replaced by the discriminated `execution.mode: "interactive-transaction" | "sequential"`. Custom `lockSchemaVersionAndGraphWrite` implementations must now return the effective `GraphCommandIsolation` observed by the same pinned session that acquired the lock. - Adding, removing, or changing an edge `matchIdentity` is a breaking schema change. The affected edge kind must be empty during activation; export and hard-delete its rows, migrate the schema, then import them so every row receives the durable key. - Inside `store.transaction()`, issue work through the callback-scoped Store. Calls through an enclosing root Store do not join the callback's transaction and may observe or create a different execution boundary. ### Minor Changes - [#558](https://github.com/nicia-ai/typegraph/pull/558) [`48532f8`](https://github.com/nicia-ai/typegraph/commit/48532f87de7ce5a769c5e090b5afddae325dbdc5) Thanks [@pdlug](https://github.com/pdlug)! - Eligible schema-managed generated-id node creates and `cardinality: "many"` edge creates on bundled root backends can now execute as one authoritative statement, including Neon HTTP and Cloudflare D1, where all required claims, projections, and side effects are either absent or fused into that statement. History/revision and Operational Identity work, plus other managed writes, continue to require the interactive transaction or an explicit typed refusal. Clarify and harden execution boundaries for managed writes. The authoritative command helper now validates command/result correlation once, with typed node, edge, and convergence overloads; first-party Store consumers no longer duplicate that check, while recorded-capture retains its direct transaction-wrapper assertion. `OptionalTransactionExecution` is now a discriminated `{ mode: "interactive-transaction" | "sequential" }` value; migrate custom consumers from `execution.atomic` to `execution.mode`. Document the distinction between interactive Store transactions, static internal adapter batches, and authoritative one-statement commands. Durable edge `matchIdentity` convergence may qualify for the one-statement root command because its canonical key has a database arbiter; claims/cardinality, undeclared dynamic `matchOn`, history/revision sidecars, and Operational Identity remain interactive-transaction contracts. The static native-batch adapter foundation remains internal; no new public Store batching API is implied. - [#556](https://github.com/nicia-ai/typegraph/pull/556) [`d2c2557`](https://github.com/nicia-ai/typegraph/commit/d2c25576fd4fb32b4db5cfe2af33b92e309060ae) Thanks [@pdlug](https://github.com/pdlug)! - Replace the optional managed-create hook with a required semantic command port that carries its root or transaction session and, only after an advisory graph lock is acquired, a graph- and session-bound coordination token. On PostgreSQL that lock statement records the effective transaction isolation in the same token, so convergence never trusts a requested option or assumed server default and adds no isolation-probe round trip. Transparent `deriveBackend` command wrappers retain the underlying session identity; a wrapper for another connection cannot reuse its token. Under that lock, PostgreSQL endpoint get-or-create folds the match-key read, endpoint validation, and insert into one statement, returning either the created edge or the existing winner without another application-level read. Adopted transactions may return an existing match at any isolation, but the create leg refuses repeatable read. Breaking change and migration: custom `GraphBackend` implementations must add a `commands` member with `{ session, execute(command, context) }`. Move managed-create and specialized edge-insert behavior into the `node.create`, `edge.create`, and `edge.converge-create` command cases, and return the typed `unsupported` result for dimensions the backend cannot apply. The optional combined-fence hook `lockSchemaVersionAndGraphWrite` now returns `Promise` instead of `Promise`; custom implementations must return the normalized effective isolation observed by the same pinned-session statement that acquires both locks. Built-in adapters already provide this port. - [#554](https://github.com/nicia-ai/typegraph/pull/554) [`ddf851d`](https://github.com/nicia-ai/typegraph/commit/ddf851d99774b5d7be9a962a7994bf405188d7ee) Thanks [@pdlug](https://github.com/pdlug)! - Introduce authoritative node create commands and reduce first-party PostgreSQL write round trips by folding schema and graph fences, uniqueness and disjointness verdicts, endpoint and cardinality checks, and generated fulltext/vector projections into atomic statements. Managed Store transactions now lease one schema fence across their writes without sacrificing the fused first statement, and endpoint get-or-create decisions are confirmed from transaction-scoped evidence on caching transports. The backend planning API now uses one required semantic command port for node and edge creates. Commands carry an explicit root or transaction session and, when an advisory graph lock was actually acquired, a graph- and port-bound coordination token; custom backends implement that same contract rather than silently falling back to a second decision path. Breaking change and migration: add `commands: { session, execute }` to custom `GraphBackend` objects and route node/edge create plans through it. A backend that cannot honor a requested plan dimension must return its typed `unsupported` result; it must not silently ignore the dimension. See the authoritative command sessions section of the backend setup guide. - [#554](https://github.com/nicia-ai/typegraph/pull/554) [`ddf851d`](https://github.com/nicia-ai/typegraph/commit/ddf851d99774b5d7be9a962a7994bf405188d7ee) Thanks [@pdlug](https://github.com/pdlug)! - Replace the three specialized edge-insert backend hooks with the shared semantic command port. Managed edge creates now compile endpoint validation, an optional schema fence, and an optional cardinality claim into one all-or-nothing `edge.create` command with an explicit result. Custom backends implement the required command contract and must apply or refuse every requested dimension. Breaking change and migration: custom backends should remove their old specialized edge-insert hook wiring and implement `commands.execute` for `edge.create` (and `edge.converge-create` when convergence is supported). Return the typed `unsupported` dimension result when endpoint, schema-fence, cardinality, or convergence behavior is unavailable. - [#561](https://github.com/nicia-ai/typegraph/pull/561) [`cdcb814`](https://github.com/nicia-ai/typegraph/commit/cdcb8144df618b4cd2294ef39969f637df831ad2) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible durable-match and cardinality-constrained `edges.bulkInsert` and `edges.bulkCreate` calls as one schema-fenced atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. The program maintains stale-claim takeover and legacy-incumbent detection, preserves typed endpoint, match-identity, and cardinality refusals, and rolls back every edge row when any constraint sidecar fails. - [#557](https://github.com/nicia-ai/typegraph/pull/557) [`6799b65`](https://github.com/nicia-ai/typegraph/commit/6799b65cff2e667aa858b625b0e0236e1bcfac28) Thanks [@pdlug](https://github.com/pdlug)! - Add graph-local durable edge match identities. An edge registration can declare one named, canonical property-field set; TypeGraph persists its directed endpoint/property key on every edge row, maintains it across normal and trusted import writers, refuses ordinary updates to identity fields, retains it across soft deletion, and releases it on hard deletion. SQLite and PostgreSQL provision and idempotently upgrade the edge relation with a pair-null check and unique arbiter. Schema-managed root `getOrCreateByEndpoints` calls using a declared identity now compile endpoint validation, the schema fence, conflict arbitration, and the created/found result into one database statement on bundled SQLite and PostgreSQL backends. Dynamic call-level `matchOn` remains available through the transaction-fenced compatibility path, and a supplied field list on a declared edge must exactly match the declaration. Bulk endpoint candidate reads use the set-oriented heterogeneous endpoint member instead of one read per endpoint pair on bundled backends. Direct creates use the same durable arbiter at every cardinality. Built-in bulk creates preserve set-oriented insertion through a conflict-arbitrated batch command rather than falling back to one managed write per row. Operation-end hooks now report `outcome: "written" | "unchanged" | "unknown"`. An authoritative get-or-create command that finds an incumbent completes as `"unchanged"` and does not fire `onError`; the same explicit outcome prevents revision/history churn for the no-write leg. Commands without an authoritative physical-write verdict report `"unknown"` instead of guessing from success. Normal import uses the same set-oriented durable command for claimless slices and savepoint-protected batch recovery for exceptional conflicts, including on history-enabled stores. Non-transactional backends refuse ambiguous per-row retry after a failed batch rather than re-inserting a possibly committed prefix. Adding, removing, or changing a match identity is a breaking schema change. The initial migration contract refuses activation while the affected edge kind holds rows; export and hard-delete those rows, migrate the schema, then import them so every row receives the new durable key. - [#539](https://github.com/nicia-ai/typegraph/pull/539) [`2dcae2f`](https://github.com/nicia-ai/typegraph/commit/2dcae2f074c00a46f521f2c45f73e49f12bfee2d) Thanks [@pdlug](https://github.com/pdlug)! - Add durable schema-only graph templates with idempotent v1 instantiation. - [#560](https://github.com/nicia-ai/typegraph/pull/560) [`da56b5f`](https://github.com/nicia-ai/typegraph/commit/da56b5f08d67fb5c867b895b8d6d988efdaecd96) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible unconstrained `edges.bulkInsert` and `edges.bulkCreate` calls as one schema-fenced native atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. The closed program validates live endpoints at the write boundary, preserves input-order results, and rolls back every bind-budget chunk when any statement fails. Cardinality claims, durable match identity, history, revision tracking, caller-owned transactions, derived backends, and unproven custom backends retain the existing transaction or fallback path. Cast closed-program CTE values to their destination column types so PostgreSQL accepts JSON and temporal values in both node and edge native batches. - [#559](https://github.com/nicia-ai/typegraph/pull/559) [`ff42eb4`](https://github.com/nicia-ai/typegraph/commit/ff42eb4c1eced3a66628ec6bd71ebfb273443394) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible schema-managed, generated-ID `nodes.bulkInsert` batches as one schema-fenced native atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. This first closed-program slice excludes claims, Operational Identity, projections, history, revision, and caller-supplied IDs; unsupported shapes retain their existing transaction or fallback path. Skip the guaranteed-empty existence-priming read for generated-ID bulk node operations while preserving caller-ID existence and resurrection checks. ### Patch Changes - [#537](https://github.com/nicia-ai/typegraph/pull/537) [`57ea6ac`](https://github.com/nicia-ai/typegraph/commit/57ea6ac74aef5a8ce22043f4cbf066d8fcac834f) Thanks [@pdlug](https://github.com/pdlug)! - Compute reciprocal-rank-fusion scores with floating-point division. ## 0.51.1 ### Patch Changes - [#526](https://github.com/nicia-ai/typegraph/pull/526) [`dbe8ee6`](https://github.com/nicia-ai/typegraph/commit/dbe8ee6a6d0f41ccb43204c0c3433e5dbe653e17) Thanks [@pdlug](https://github.com/pdlug)! - Read persisted unique constraints that omit `scope` or `collation` by applying the documented `"kind"` and `"binary"` defaults. Schema-management APIs can now inspect databases written with those omitted fields instead of reporting a malformed schema document. ## 0.51.0 ### Highlights TypeGraph 0.51 removes `drizzle-orm` from the dependency graph of its portable entrypoints. Applications using the root, backend, core, schema, indexes, graph-extension, interchange, profiler, graph-merge, or provenance entrypoints can now install and run TypeGraph without Drizzle; managed SQLite and PGlite Stores and explicit `/adapters/drizzle/...` entrypoints still use it. Custom backend behavior is now resolved through explicit capability declarations and shared bundles instead of scattered optional-member checks. The first bundle set covers claims, statement execution, recorded revision origins, batch point reads, unique-sidecar batching, and contribution health. Recursive traversal and write-fence support are also explicit: unsupported engines can refuse recursive operations with a stable reason, while stateful features require a fence plan backed by real locks or engine-serialized writers. ### Upgrade notes - Applications using managed SQLite or PGlite Store entrypoints, or any explicit Drizzle adapter entrypoint, must keep `drizzle-orm` installed. Managed Store factories report `MISSING_PEER_DEPENDENCY` with the installation command when it is absent. - A custom backend hosting Operational Identity, `history: true`, or `revisionTracking: true` must declare truthful `capabilities.pessimisticLocks`. TypeGraph refuses an unfenced declaration rather than assuming safety from the dialect name. - A custom backend without recursive SQL or an equivalent graph-native operation should declare `capabilities.recursiveTraversal: { supported: false, reason }`. Omission retains the pre-0.51 assumption that recursive traversal is supported. - A custom backend that created the timestamp-only recorded-time preview schema must implement `recordedTableDdl(tableNames)` before running `migrateLegacyRecordedTime`; bundled SQLite and PostgreSQL backends already provide it. ### Minor Changes - [#521](https://github.com/nicia-ai/typegraph/pull/521) [`da251ef`](https://github.com/nicia-ai/typegraph/commit/da251ef021c8b2b61717f0c8fa7c07535a28599a) Thanks [@pdlug](https://github.com/pdlug)! - Custom backend capabilities now resolve through six shared bundles instead of scattered `undefined` checks. This makes each operation family consistently choose one of three outcomes: use the declared member, take its documented fallback, or refuse with a typed error. The pilot covers `claims`, `statementExecution`, `recordedRevisionOrigins`, `batchPointRead`, `uniqueSidecarBatch`, and `contributionHealth`. Batch point reads and supported sidecar operations degrade to their existing per-item implementations when the batch member is absent. Operations without a safe fallback keep their existing typed refusal. A backend whose capability declaration disagrees with the members reachable on its execution port now refuses with `CONSTRAINT_CLAIM_SURFACE_MISMATCH` for claims or `BUNDLE_PORT_SURFACE_MISMATCH` for the other bundles. `CAPABILITY_BUNDLES`, the six named definitions, their verdict and binding types, and the `resolveBundle`, `bindCore`, `bindExtra`, and `bindExtraIfReachable` helpers are public for backend conformance tooling. Both bundled backends already implement every required member, so their behavior is unchanged. This pilot covers six of twenty-one member-bearing operation families; the other fifteen continue to work through their existing paths. See [Capability bundles](https://typegraph.dev/backend-setup#capability-bundles) for the complete member and fallback matrix. - [#520](https://github.com/nicia-ai/typegraph/pull/520) [`b6478f6`](https://github.com/nicia-ai/typegraph/commit/b6478f6e67d57767784c59691268049ce8585f75) Thanks [@pdlug](https://github.com/pdlug)! - `drizzle-orm` is now optional when using TypeGraph's portable entrypoints. Applications that import only the root, backend, core, schema, indexes, graph-extension, interchange, profiler, graph-merge, or provenance entrypoints no longer need Drizzle installed. Applications using a managed SQLite or PGlite Store, or an explicit `/adapters/drizzle/...` entrypoint, must still install it with `npm install drizzle-orm` when their package manager does not install optional peers automatically. Managed Store factories report a typed `MISSING_PEER_DEPENDENCY` error with that command; explicit Drizzle adapters retain the runtime's raw module-resolution error. See [Managed Store Entrypoints](https://typegraph.dev/backend-setup#managed-store-entrypoints) and [installation troubleshooting](https://typegraph.dev/troubleshooting#missing-optional-drizzle-orm-peer). - [#520](https://github.com/nicia-ai/typegraph/pull/520) [`b6478f6`](https://github.com/nicia-ai/typegraph/commit/b6478f6e67d57767784c59691268049ce8585f75) Thanks [@pdlug](https://github.com/pdlug)! - Portable entrypoints no longer reach Drizzle through recorded-time migration, claim comparison, or removal-statement builders. This completes the separation that lets consumers use the ten portable entrypoints without installing `drizzle-orm`; source and packaged-output checks now prevent those imports from returning. Custom backends that need to migrate the timestamp-only recorded-time preview schema must implement the new optional `GraphBackend.recordedTableDdl(tableNames)` member. It returns backend-owned table and index DDL for the temporary and final recorded-relation names. A migration that reaches the legacy rewrite without this member throws `UnsupportedBackendCapabilityError` with `details.capability: "recordedTableDdl"` instead of importing Drizzle or crashing. See [Migrating Preview Recorded Time](https://typegraph.dev/schema-management#migrating-preview-recorded-time). - [#521](https://github.com/nicia-ai/typegraph/pull/521) [`da251ef`](https://github.com/nicia-ai/typegraph/commit/da251ef021c8b2b61717f0c8fa7c07535a28599a) Thanks [@pdlug](https://github.com/pdlug)! - Custom backends can now declare whether they support recursive traversal with `capabilities.recursiveTraversal`. Absence means supported for backward compatibility; an engine without recursive SQL or a graph-native equivalent declares `{ supported: false, reason }`. The decision is resolved through a branded `RecursiveTraversalVerdict` that only `resolveRecursiveTraversal` can construct, so a caller cannot forge one by writing `{ supported: true }` inline. Also exported: `assumeRecursiveTraversalSupported` (the one sanctioned way to obtain a verdict without a backend, used by the query compiler's no-backend entry point), `assertRecursiveTraversal`, `recursiveTraversalUnsupportedError`, and the `RecursiveTraversalCapability` type itself. Variable-length queries, `store.subgraph()`, and the three recursion-dependent historical identity reads now refuse an unsupported declaration with `ConfigurationError` code `RECURSIVE_TRAVERSAL_UNSUPPORTED`; `details.operation` and `details.reason` identify the affected path and engine limitation. `weightedShortestPath` keeps working when temporary statements are available, reconstructing the same path through `pathLength + 1` predecessor reads instead of one recursive extraction statement. `CompileQueryOptions` gains an optional `recursiveTraversal`, threaded by `propagateOptions` into every set-operation sub-compile so a `union()`/`intersect()`/`except()` operand carries the same verdict as its parent query. Bundled factories refuse contradictory declarations — unsupported without a reason or supported with a dangling reason — using `CAPABILITY_DECLARATION_CONTRADICTION`. The bundled SQLite and PostgreSQL backends declare support, so their query behavior is unchanged. Custom backend note: the factory-owned clone of `backend.capabilities` is now deep-frozen, so mutating it after construction throws. Objects supplied through factory options remain caller-owned and mutable. See [Recursive traversal capability](https://typegraph.dev/backend-setup#recursive-traversal-capability) and the [recursive query guide](https://typegraph.dev/queries/recursive#backend-support). - [#521](https://github.com/nicia-ai/typegraph/pull/521) [`da251ef`](https://github.com/nicia-ai/typegraph/commit/da251ef021c8b2b61717f0c8fa7c07535a28599a) Thanks [@pdlug](https://github.com/pdlug)! - **Custom backend migration:** a backend not built by `createSqliteBackend` or `createPostgresBackend` must now declare `capabilities.pessimisticLocks` before hosting Operational Identity, `history: true`, or `revisionTracking: true`. Store construction refuses an undeclared backend immediately instead of risking an unfenced concurrent write. PostgreSQL backends normally declare `pessimisticLocks: { advisoryLocks: true, tableLocks: true, serializedWriters: false }`; SQLite backends normally declare `pessimisticLocks: { advisoryLocks: false, tableLocks: false, serializedWriters: true }`. Verify those values against the engine's actual guarantees rather than copying them for a different topology. `BackendCapabilities` gains an optional `pessimisticLocks` field (`{ advisoryLocks, tableLocks, serializedWriters }`) declaring how an engine serializes concurrent writers, if at all. `resolveWriteFencePlan` is the one place that declaration turns into a `WriteFencePlan` (`lock` / `engine-serialized` / `unfenced`) every lock site now consumes instead of re-deriving from `dialect` inline, and `requireWriteFence` is the one place an operation's specific lock requirement (`"advisory-lock"` / `"table-lock"`) is checked against the resolved plan, refusing with `WRITE_FENCE_UNAVAILABLE` when it cannot be met. This consolidates eight call sites that used to spell the same dialect-keyed decision independently. `BackendCapabilities` also gains an optional `recordedTimeOwnership` field (`"typegraph-relations"` | `"engine-native"`) naming who allocates recorded-time revisions. Absent means `"typegraph-relations"` — today's behavior for every existing backend. Declaring `"engine-native"` together with `history`/`revisionTracking` is refused at construction with `ENGINE_NATIVE_RECORDED_TIME_NOT_IMPLEMENTED` as an interim measure, independently of the write-fence plan, because the engine-native read/write path does not exist yet; a later release lifts this refusal with that path. Both bundled backends already declare their write-fence support, so shipped configurations keep their existing behavior. See [Write fence declaration](https://typegraph.dev/backend-setup#write-fence-declaration-pessimisticlocks), [recorded-time ownership](https://typegraph.dev/backend-setup#recorded-time-ownership-recordedtimeownership), and the [stable error codes](https://typegraph.dev/errors#write-fence-declaration-codes). ### Patch Changes - [#521](https://github.com/nicia-ai/typegraph/pull/521) [`da251ef`](https://github.com/nicia-ai/typegraph/commit/da251ef021c8b2b61717f0c8fa7c07535a28599a) Thanks [@pdlug](https://github.com/pdlug)! - **For contributors:** CI now compares the public `etc/*.api.md` snapshots with the last published tag through `test:api-surface`. The check fails when an external consumer would lose a member, see an optional member become required, or need to supply a newly required member through a contravariant API position. This adds no runtime or published API; it makes breaking surface changes visible before release. See the [release verification commands](https://github.com/nicia-ai/typegraph/blob/main/docs/RELEASE.md#pre-release-verification). ## 0.50.0 ### Highlights TypeGraph 0.50 completes database-arbitrated enforcement across the declared constraint families. Hierarchy-wide uniqueness, `disjointWith`, and edge `one`, `unique`, and `oneActive` cardinality now remain fenced under concurrent writers, including interchange import. Claim ownership is the concrete `(kind, id)` pair, so namesake nodes in different kinds cannot take over one another's reservations and lifecycle cleanup releases only the claims a node owns. `store.verifyConstraintFences()` audits uniqueness, disjointness, and cardinality violations that predate these fences without modifying data. Internally, every Store and interchange write now runs through one typed write-plan/session pipeline; this architectural consolidation does not intentionally change public statement order, lock scope, or error behavior beyond the constraint corrections described above. ### Upgrade notes - Databases created before 0.50 need the new `typegraph_edge_claims` relation before their first constrained edge write. Run the normal idempotent bootstrap path or apply the SQL emitted by `generatePostgresMigrationSQL()` / `generateSqliteMigrationSQL()`; a missing relation is reported as `EDGE_CLAIM_RELATION_MISSING`. - The new fences prevent future conflicting writes but do not choose winners among violations already stored. Run `store.verifyConstraintFences()` after upgrading and resolve any reported owners or edge ids deliberately. - Custom backends that implement constraint claims should add the `edgeClaims` table name and the edge-cardinality claim members, then declare matching `capabilities.constraintClaims`. Omission remains a supported opt-out; a declaration/member mismatch is refused. - Custom `deleteUnique` implementations must scope release by `concreteKind` and `nodeId`, and apply the `nodeKind` predicate only when `params.nodeKind` is present. Keeping the old unconditional predicate can leak lifecycle claims permanently. ### Minor Changes - [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Issue a node's claims on the side of the row write their placement names, and refuse the writes a backend with no transactions cannot undo. A uniqueness claim whose axis spans kinds beyond the writer's own is the **only** fence for that axis — the nodes primary key is `(graph_id, kind, id)`, so an `Employee`'s insert does not collide with a `Contractor`'s — and a fence issued after the write it fences is not a fence. Those claims now precede the row insert they gate, with the reservations given back if that insert does not land, so a refusal leaves zero net effect. On the import path this is what turns a violation into a refusal instead of a committed row: `importGraph` recovers per row and takes no per-graph lock, so a claim written after the row it was supposed to refuse let the row commit. A claim whose axis **is** the writer's own kind keeps the position it has today, after the row: the uniques primary key at that axis is already the complete fence for it, moving it would buy nothing, and it would cost a refusal on backends with no transactions. Placement is decided once, per claim, from the one fact both readings turn on — does this claim's axis span kinds beyond the writer's own? — and carried as data through the entry, the claim seam, the refusal and the lock projection. Which writes take the per-graph advisory lock, and the reason each reports, are unchanged. Two new refusals follow, both on `transactions: false` backends (Cloudflare D1, `drizzle-orm/neon-http`, `transactionMode: "none"` SQLite), and both `ConfigurationError` / `CONSTRAINT_WRITE_FENCE_UNSUPPORTED`: - **`importGraph` / `importGraphStream` into a graph any of whose node kinds declares a unique constraint of any scope, or any of whose edge kinds is non-`many`.** Import writes claim rows like every other writer but is not covered by the write-transaction refusal, so it would have written reservations with nothing to roll them back. The refusal is computed up front, before the first chunk, so a streamed import cannot commit _k-1_ chunks and then fail. Disjointness owes no claim yet — that fence lands in a later batch — so a disjoint-only graph is not refused here. - **A node UPDATE or RESURRECT whose kind declares only `scope: "kind"` unique constraints**, reason `nodeUniquenessClaim`. This closes an existing hole rather than paying for a new one: the transition seam already claims before its gated row write for every scope, so that path already wrote a reservation with nothing to undo it. The matching **create** is not refused — its claim stays after the row — which is the pair that makes the rule legible: same kind, same constraint, opposite verdicts, decided only by placement. `ConstraintFenceReason` gains `nodeUniquenessClaim` for that refusal. It is never returned by the lock projection, so it cannot widen the set of writes that take the per-graph lock. Claim statements within each placement group are issued in one canonical order — code-point on `(relation, graph, axis, constraint, key)` — and the pre-insert group is always issued first, so two writers touching the same claim rows for one row acquire them in the same order instead of deadlocking. For a node create this is observable as statement order: a kind owing only own-axis claims emits exactly what it emits today, a kind owing a cross-kind claim emits it ahead of the row insert, and a kind owing both emits two claim statements, one on each side. - [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Tolerate the concurrent `CREATE EXTENSION` race when materializing trigram indexes, and give every extension install one owner. `method: "trigram"` needs `pg_trgm`, the extension is database-global, and the claim that serializes an index build is keyed per index — so two materializers building different trigram indexes both reach `CREATE EXTENSION IF NOT EXISTS pg_trgm`. That statement is not a concurrency primitive on PostgreSQL: its existence check cannot see another session's uncommitted `pg_extension` row, so the loser waited for the winner and was handed SQLSTATE 23505 instead of a notice, reporting `failed` for an extension the winner had already installed. `GraphBackend` gains an optional `ensureExtension(name)` member — the single owner of "install a database-global extension idempotently" — which the bundled PostgreSQL backend implements with both fences: a transaction advisory lock keyed on the extension, so same-key installers never raise at all, and the concurrent-DDL retry its table and column creates already use, which clears the 23505 an installer that did NOT take that lock can still hand it (a peer on an older version, or a `capabilities.transactions: false` backend with no transaction to hang the lock on). The name is validated against the exported `DATABASE_EXTENSION_NAMES` allowlist rather than interpolated freely. `GraphBackend.ensureTrigramExtension`, the `pg_trgm`-only member added in 0.47, is deprecated in favour of `ensureExtension` and now says exactly the same thing: the bundled PostgreSQL backend implements it by delegating, and index materialization consults it only after `ensureExtension`, so a backend written against 0.47 keeps its fence unchanged. A backend implementing neither keeps issuing the bare statement with materialization's own one-shot retry, so a third-party trigram index is still materialized. The advisory-lock key changed from `typegraph:pg-trgm-ddl` to `typegraph:extension-ddl:` when the fence generalized to any allowlisted extension — a 0.47 peer therefore takes a different key, which is exactly why the retry is retained on the locked path too. Closes [#446](https://github.com/nicia-ai/typegraph/issues/446). - [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Fence `disjointWith` with a claim on the declared pair, and enforce it under `importGraph`. Disjointness was probed and never fenced. The nodes primary key is `(graph_id, kind, id)`, so `Person "X"` and `Company "X"` are two different rows by construction — the exact collision the axiom forbids is the one the database cannot refuse — and the probe was only as good as the serialization around it. `importGraph` takes no per-graph lock and, until now, ran no disjointness probe at all: an import could commit both halves of a violating pair. A create of a kind with a `disjointWith` partner now reserves one claim row per partner, in the same relation uniqueness claims use, at the pair's own axis with the node's id as the key. Both kinds of a pair fold to one axis through the registry's own canonical pair label, so their claims contend for one row and its primary key refuses the second writer. The claim precedes the row it gates and is given back if that row does not land, so a refusal leaves zero net effect. Because the two families arrive through one list of claim sites, every path that already maintained uniqueness reservations — create, batch create, delete, import — maintains disjointness reservations too. A **resurrect** — a soft-deleted node revived by `.create()` on its tombstoned id, `upsertById`, `upsertByIdFromRecord`, `bulkUpsertById`, or `getOrCreateByConstraint` — reserves the same claim, since reviving a tombstone re-introduces a live id under a kind exactly as a create does; the resurrect leg no longer has a window where it could revive a node under an id a disjoint partner already holds live. `importGraph` gains the per-row disjointness probe both node paths were missing, and the per-row recovery it sits in is widened from `UniquenessError` to every declared-constraint refusal. Behavior deltas: - **Import now enforces disjointness.** A payload containing a `Person` and a `Company` with the same id, in one batch or in sequence, refuses the second **row** and reports it in `errors` while the import continues — the accepted rows commit. Previously both committed silently. A _concurrent_ violation, taken by another writer between this row's probe and the batch's claim, still surfaces from the claim and aborts the import; that asymmetry is what import already does for uniqueness. - **`importGraph` / `importGraphStream` is refused on a `transactions: false` backend when any node kind has a disjoint partner**, joining the unique-constraint and non-`many`-cardinality cases (`ConfigurationError` / `CONSTRAINT_WRITE_FENCE_UNSUPPORTED`). A disjoint create's claim precedes its row, and without a transaction a failure between the two would leave a reservation with no repair path. - **A kind or unique constraint name containing `U+001E` is refused at `defineNode` / `defineGraph`.** That code point builds the axes that are not kinds, so a name carrying it could spell the reserved disjointness axis. New refusal on an input no real schema carries. The refusal a caller sees is the family's own `DisjointError`, with the same payload whichever layer produced it: the probe reads the partner's node row, the claim reads its reservation's owner, and both name the holder's concrete kind. - [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Fence edge cardinality with a claim relation, and enforce it under `importGraph`. Declared edge cardinality was probed and never fenced. `one` and `oneActive` are predicates over `(kind, from)` and `unique` is one over `(kind, from, to)`, while the edges relation's only uniqueness is its `(graph_id, id)` primary key — so two writers could both count zero sibling edges and both commit, and nothing in the schema re-decided at write time. `importGraph` made it worse by running no cardinality probe at all: a payload could commit any number of edges a `cardinality: "one"` declaration forbids. A new relation, `typegraph_edge_claims`, keyed `(graph_id, axis, key)`, is the fence. Each constrained edge write reserves the axis its declaration spans (`:`) against the endpoint identity the declaration covers, in two statements: a decision-free create-or-lock that reports the committed holder, then — only when the holder is a different edge — a conditional takeover that succeeds exactly when that holder is no longer an edge the axis and key describe. Deciding inside one upsert would read the pre-lock snapshot of the edges relation under READ COMMITTED and accept both writers; the split is what makes the second one lose. The claim needs no release path. A holder that is soft-deleted, hard-deleted, or (for `oneActive`) ended fails the takeover's liveness predicate and is replaced in place, so no delete, end, cascade or kind-removal path participates in the fence. The holder is identified by its kind and source endpoints as well as its id, because edge ids are caller-suppliable: a reused id would otherwise read as a live holder and block its axis forever. `EDGE_CARDINALITY_SPECS` is the one table both the TypeScript probe and the takeover's SQL read for which endpoints an axis covers, whether an edge born already ended claims at all, and what a holder must still be — so the probe and the fence cannot drift apart. Behavior deltas: - **Import now enforces edge cardinality** (`one` / `unique` / `oneActive`). Two edges from one source in one payload refuse the second **row** and report it in `errors` while the import continues; the accepted rows commit. Import's edge slice reuses the store's own in-batch cardinality accounting to make that per-row rather than a whole-slice abort. A _concurrent_ violation, taken by another writer between a row's probe and the slice's claim, still surfaces from the claim and aborts the import — the same asymmetry import already has for uniqueness. - **`PostgresTableNames` / `SqliteTableNames` / `SqlTableNames` gain `edgeClaims`**, and `BackendCapabilities` gains an **optional** `constraintClaims`. Absent means `false`: a backend that predates the claim relations keeps every fence it has today and is never refused for the absence. Both bundled dialects declare `true` and implement every claim member. A backend whose declaration and surface disagree in either direction is refused with `ConfigurationError` / `CONSTRAINT_CLAIM_SURFACE_MISMATCH` rather than silently unfenced. - **`GraphBackend` gains optional `claimEdgeCardinality`, `claimEdgeCardinalityBatch` and `purgeEdgeClaims`.** Additive; a custom backend that omits them declares `constraintClaims: false` and keeps working. - **A database bootstrapped before this release needs the new table.** It is emitted by the existing idempotent boot path and by `generatePostgresMigrationSQL` / `generateSqliteMigrationSQL`. A store reaching a missing relation on its first constrained edge write is refused with a typed `ConfigurationError` (`EDGE_CLAIM_RELATION_MISSING`) naming the relation and the migration to run, instead of an opaque driver failure. The refusal a caller sees is the family's own `CardinalityError`, built by the same functions the probe calls, so it is indistinguishable from the serial refusal it replaces. - [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Scope uniqueness-claim releases to the node that owns the claim. A `typegraph_node_uniques` row records both the axis it fences on (`node_kind`) and the node that owns it (`concrete_kind`, `node_id`). Releases keyed on the axis alone could not tell those apart: a soft delete gave up whatever row sat at the node's own kind, and a kind removal deleted every claim whose _axis_ was the removed kind. Both readings are wrong the moment an axis and a concrete kind differ — they leak a claim that blocks its key forever, and they delete a surviving sibling's claim. Release now has three explicitly different shapes, each with one owner: a **lifecycle** release gives up every claim the node holds for a constraint and key at whatever axis it sits on (soft delete, an update's key-change release, the resurrect diff); a **compensating** release undoes exactly the row a failed write claimed, at the axis it claimed on; and **kind reaping** removes every claim the removed kind's nodes own, through the new `buildHardDeleteUniquesByConcreteKind` builder that `materializeRemovals` and the new optional `hardDeleteUniquesByConcreteKind` backend member both compile. `DeleteUniqueParams` gains the owner pair `concreteKind` / `nodeId`, and its `nodeKind` becomes optional — present selects the compensating shape, absent the lifecycle one. A third-party `GraphBackend` that implements `deleteUnique` must make BOTH changes: add `concrete_kind` and `node_id` to its predicate, and make the `node_kind` term _conditional_ on `params.nodeKind` being present. Doing neither does not leave the old behavior in place: the lifecycle release now passes no `nodeKind`, so a predicate that still spells `node_kind = :nodeKind` unconditionally compares against NULL, matches zero rows, and releases nothing — every soft delete and every key-change release leaks its claim, and the key stays blocked forever. TypeScript cannot catch it either, since an `undefined` bound into a SQL template is accepted silently. - [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Fence `scope: "kindWithSubClasses"` uniqueness on a shared claim axis, and decide claim ownership by `(concrete_kind, node_id)`. A shared-scope unique constraint used to reserve its key under the writer's OWN kind, while the probe walked the whole hierarchy. Sibling kinds therefore reserved rows that could never collide — the `typegraph_node_uniques` primary key was structurally incapable of refusing the second writer, and only the per-graph lock stood between two concurrent creates and a duplicate ([#436](https://github.com/nicia-ai/typegraph/issues/436)). The claim is now written at the scope's **axis**: the code-point minimum of the connected `subClassOf` component, which every kind in that component computes identically, so two writers of two kinds contend for one row and the primary key is the fence. Under multiple inheritance or multiple roots this is stricter than before — the component is what the old "walk one root's descendants" reading was documented to mean — and the probe still visits every kind in scope, so rows written before the upgrade are still read and no data migration is required. Ownership of a claim is the pair `(concrete_kind, node_id)`, not the id alone. Ids are unique only per kind, so `Employee "X"` and `Contractor "X"` are two different nodes; comparing ids let the second one match the "this row is already mine" arm of the upsert, rewrite the incumbent's `concrete_kind`, and read its own id back as proof it had won. Both upsert builders now compare the pair and return it, the accept/refuse test compares the pair, and the batch-validation cache remembers the pair — so a live claim held by a namesake under another kind is a refusal where it was previously a silent takeover. In one import batch, that refusal is now reported per row (with the earlier rows committed) instead of aborting the whole batch at the flush. Two payload/behavior corrections come with it. `UniquenessError.kind` now names the **holder's own kind** rather than the `node_kind` the row was found under, which is the same value on a single-kind scope and the meaningful one on a shared scope. And the cross-kind lookups behind `findByConstraint` and `getOrCreateByConstraint` state their preference explicitly — axis first, then the remaining kinds in code-point order, live rows preferred over tombstoned ones — so a database carrying both a pre-upgrade and a post-upgrade row for one key resolves deterministically instead of by iteration order. The per-graph write lock is unchanged: which writes take it, and the reason each one reports, are byte-identical, now derived from the same claim-site classification the claim itself is written from. - [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.verifyConstraintFences()`, the read-only audit of constraint violations that predate the fence. The claim relations refuse the second live claimant of an axis from the first write after upgrade onward, but they repair nothing that is already there. A database that carried two live siblings sharing a `scope: "kindWithSubClasses"` key, an id live under both kinds of a `disjointWith` pair, or two live `cardinality: "one"` edges from one source keeps carrying them: the next write that touches such an axis is refused with the ordinary typed error naming the incumbent, and until then nothing says so. This is the diagnostic that says so. It reads the relation each constraint is **declared over**, never a claim relation's primary key. A claim key admits one row per axis by construction, and a database written before the claim tables existed holds no edge claims at all, so a claim scan would report zero violations on precisely the data the audit exists to find. Uniqueness is read from the live `uniques` rows and folded onto the axis each row's `node_kind` belongs to — which is how a pre-upgrade duplicate sitting at two different `node_kind`s is found at all — restricted to constraint names the graph declares, so disjointness claims (whose `node_kind` is a pair label, not a kind) are audited from the nodes relation instead. Contention is counted in distinct **owner pairs** (`concrete_kind`, `node_id`), not rows, so one node legitimately holding its key at a legacy axis and at the current one is not reported. Each entry names the claim row two claimants contend for — built by the same functions the fence writes with — plus the conflicting `owners` (uniqueness, disjointness) or `edgeIds` (cardinality). It writes nothing and repairs nothing: choosing which claimant keeps an axis is a data-loss decision that belongs to the operator. - **`GraphBackend` gains an optional `readConstraintFenceViolations`.** Additive; both bundled dialects implement it through one shared statement per family. `store.verifyConstraintFences()` refuses with `ConfigurationError` / `CONSTRAINT_FENCE_AUDIT_UNSUPPORTED` on a backend without it, rather than returning an empty report a caller would read as "clean". - **`KindRegistry` gains `disjointKindPairs()`**, the declared pairs as kind pairs — the inverse of the internal pair label, so an enumerating caller never spells the label's form itself. The parity matrix in `backend-setup.md` gains the three rows this mechanism owes a reader: the `constraintClaims` capability, PostgreSQL's `40001` in place of the typed error above READ COMMITTED, and the claim row's lock being held to end-of-transaction on both dialects — refusal included, so a caller that catches a constraint error and continues blocks other writers of that axis for the rest of its transaction. ### Patch Changes - [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Restore a merged node's `disjointWith` reservations after a resolved merge write set commits. A resolved node write set (the graph-merge apply path, and the set update it shares its preflight with) validates the whole after-image, then clears the affected nodes' sidecar rows so its upserts can take the approved keys in any order, then rebuilds them once at the end — the rebuild is what keeps a coalesced, otherwise side-effect-free upsert from leaving its key unreserved. The clear is keyed on the claim's OWNER, so it takes every reservation the affected nodes hold, and since 0.48 that includes their `disjointWith` claims as well as their uniqueness claims. The rebuild now goes through the same claim writer an ordinary create uses, which restores whatever the row's kind owes rather than the uniqueness slice alone; previously a merged node came out of every merge with its disjointness axis unreserved, leaving it unfenced against a disjoint namesake for the rest of the graph's life. - [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Close the write-pipeline seam: the row-work read projection is now the type every write path actually uses. The preparation helpers, constraint and uniqueness probes, batch validation caches and identity hooks the migrated modules reach are re-typed off `GraphBackend | TransactionBackend` onto the narrow handle a write frame hands out, so the counted `unfencedTarget` widening falls from seventeen call sites to one. That one is structural rather than migration debt — the bulk `getOrCreateByEndpoints` legs re-enter the executor against their enclosing frame's target, and re-entry mints a session — and the ratchet now records it as a reasoned floor with a second escape failing the build. Three seams are stated instead of implied along the way. `IdentityTarget` is an explicit facet composition of what an identity statement needs (reads plus the optional raw-statement port) rather than the whole backend union, with the service context's `backend` named for what it is: the handle the service opens its own write frames on. `ConstraintContext` and the uniqueness probe carry read facets, which states in the type that no check in either module writes. The executor's overlaid-session mint takes the READS to answer rather than a backend to write through, so row work can no longer hand the session an arbitrary backend. `src/store/operations/index.ts` publishes the seam (`runWritePlan`, `WritePlan`, `WriteSession`) and still re-exports no step or sidecar module, and the write-pipeline files come off knip's ignore list. No public API, behavior, error type, statement or lock scope changes. - [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Route every edge write through the write pipeline: the raw `insertEdge` / `updateEdge` / `deleteEdge` / `hardDeleteEdge` calls move into a new `edge-write-pipeline.ts` step module and the insert dispatch, reached through the session's seven edge methods under an edge write plan. All nine edge entry points — including the bulk `getOrCreateByEndpoints` batch — now declare their constraint probe as plan data instead of spelling it at the transaction call, and an edge update states its asserted identity and validity bound as a fence record whose keys are required, so a partially stated fence is a type error rather than a silently unfenced write. The zero-row diagnosis (`withUnmatchedEdgeUpdateRefusal`) moves to `edge-write-fences.ts` and stays caller-applied, because the store's converge-or-refuse reading and interchange import's report-and-continue reading are genuinely different recovery policies. No public API, behavior, error type, statement, or lock scope changes. - [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Route interchange import through the write pipeline: its hand-built write context and its own `runInWriteTransaction` call are gone, and all six write legs — the batched and per-row node creates, the node update, the batched and per-row edge creates, and the edge update — now run as session calls under one write plan whose identity participation the executor acquires. Import's hand-rolled `insertNodesBatch === undefined` / `insertEdgesBatch === undefined` probes converge on the insert dispatch that already owns that decision, and its edge update states the five immutable identity components and the window guard's stored lower bound as a fence record with required keys instead of a spread convention. The write-pipeline exemption list has no migration debt left: every remaining entry is a step, sidecar or reasoned carve-out. No public API, behavior, error type, statement or lock scope changes. - [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Route every node write except the set update through the write pipeline: the eight managed entry points in `node-operations.ts` now compose a write plan and run through the executor, and their row and sidecar writes are the session's fused units rather than hand-paired calls. The identity-participation decision moves from eight inline conditions to one declaration per plan, and the update path's validity lower-bound fence becomes a required argument instead of a spread convention. No public API, behavior, error type, or lock scope changes. - [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Add the internal write-pipeline seam: a total, disjoint classification of every `GraphBackend` member, a typed write plan, per-kind write fences with total applier maps, the fused write session, and the executor that is the single sanctioned caller of the write transaction. An ESLint rule now bans direct backend mutation calls outside the step and sidecar modules that own them, with a declared exemption list a ratchet holds equal to the tree. No public API, behavior, statement order, or lock scope changes. - [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Route the set-based node update through the write pipeline: `updateWhere`'s transaction body is now `applyNodeSetUpdate`, a node write step, reached through `session.reviseNodeSet` under a write plan. The uniqueness drop it performs moves to the uniqueness sidecar module, and the fence the set UPDATE has no field to carry is now refused by name instead of being absent from the call. With this, no module outside the declared step and sidecar modules calls a backend mutation member for a node write. No public API, behavior, error type, or lock scope changes. ## 0.49.0 ### Minor Changes - [#488](https://github.com/nicia-ai/typegraph/pull/488) [`0db8f9c`](https://github.com/nicia-ai/typegraph/commit/0db8f9c68ae0a13c9363f393c889d56c48114f4c) Thanks [@pdlug](https://github.com/pdlug)! - Allow `importGraph` and `importGraphStream` to stage interchange data directly into an opaque ingestion branch while preserving its deferred-uniqueness boundary. - [#486](https://github.com/nicia-ai/typegraph/pull/486) [`b82d436`](https://github.com/nicia-ai/typegraph/commit/b82d436a5c151ee6fa230591861785f82aaee302) Thanks [@pdlug](https://github.com/pdlug)! - Let constraint-aware ingestion branches stage Operational Identity same and different assertions through a conditional assertion-only facade, so duplicate unique aliases and their identity evidence can reach merge planning together. ### Patch Changes - [#485](https://github.com/nicia-ai/typegraph/pull/485) [`184fc96`](https://github.com/nicia-ai/typegraph/commit/184fc965e20973b6196fb1349e91e4139c3e6b27) Thanks [@pdlug](https://github.com/pdlug)! - Stop publishing a private workspace ESLint config in package metadata, preventing lockfile-refreshing pnpm installs from resolving an unpublished package. - [#489](https://github.com/nicia-ai/typegraph/pull/489) [`f4e31d9`](https://github.com/nicia-ai/typegraph/commit/f4e31d91d04faf93033b124bc351c39458ec16d3) Thanks [@pdlug](https://github.com/pdlug)! - Translate missing fulltext storage failures from Cloudflare Durable Objects SQLite into `ContributionUnavailableError` while preserving the underlying database error and transactional rollback. ## 0.48.0 ### Highlights TypeGraph 0.48 adds constraint-aware ingestion branches. Applications can stage overlapping records without prematurely enforcing node uniqueness, then resolve duplicates and validate the complete merge write set atomically against the target graph. Writes with an implicit start and a historical end now preserve an unknown lower validity bound, so an already-ended record remains readable before its end. The new `repairInvertedValidityWindows()` utility reports or repairs older rows whose stored start is later than their end. PostgreSQL connection detection also recognizes additional single-connection configurations, preventing streaming interchange from waiting indefinitely on a connection it already holds. ### Upgrade notes - Existing inverted validity windows are not repaired automatically. Run `repairInvertedValidityWindows()` in report mode first; apply repairs with writers stopped, preferably across both live and recorded relations, then re-baseline outstanding merge branches. - Rows created with an implicit start and a historical `validTo` can now return `meta.validFrom: undefined`. Custom backends should use `resolveStampedValidityLowerBound` for insert and node-resurrection stamping. - String-valued PostgreSQL single-connection settings now trigger the same interchange guards as numeric `max: 1`. Working-copy clones on those connections use a materialized in-memory export; use the `serializedResource` declaration when connection topology cannot be inferred correctly. ### Minor Changes - [#476](https://github.com/nicia-ai/typegraph/pull/476) [`a58ba03`](https://github.com/nicia-ai/typegraph/commit/a58ba031b439f3befbcc073e016f896b178af00e) Thanks [@pdlug](https://github.com/pdlug)! - Store no validity lower bound for a write that would otherwise be born already ended. A write that stamps a `valid_from` the caller did not state now stores the write instant only when doing so leaves a window some coordinate can read: if a stated `validTo` falls at or before that instant, the row is stored with no lower bound at all — "ended at T, start unknown" — and reads back at every `asOf` before its end instead of at none ([#407](https://github.com/nicia-ai/typegraph/issues/407)). The decision lives in the SQL builders, so it holds for every `GraphBackend` caller, including interchange import and trusted import, not only for the store paths. Three consequences worth naming: - `meta.validFrom` is `undefined` for such a row, where it used to be an instant no query could match. - A resurrecting `upsertById` / `bulkUpsertById` that names a lone historical `validTo` on a tombstoned node **no longer refuses**: it reaches the same stored shape a `create` on the same id reaches. One stated window, one outcome, whichever entry point resets the window. - A `validTo` in the future is unchanged — it still stamps the write instant, so a scheduled-end row stays invisible before it existed. Rows already stored with an inverted window are not rewritten; they stay invisible at every coordinate until repaired. Custom `GraphBackend` implementations should route node/edge insert stamping and node resurrection stamping through `resolveStampedValidityLowerBound`, now exported from `@nicia-ai/typegraph/backend`, so adapter-specific builders cannot drift from the shared validity-window contract. Edge resurrection retains its stored lower bound and therefore does not stamp one. - [#477](https://github.com/nicia-ai/typegraph/pull/477) [`09318d6`](https://github.com/nicia-ai/typegraph/commit/09318d672e05f50bc3b1c61a4c9037b5f6f0c1be) Thanks [@pdlug](https://github.com/pdlug)! - Recognize every spelling of a one-connection Postgres cap that its driver actually honors, so interchange refuses the pairs that would otherwise hang. Serialized-connection detection previously required a numeric `max: 1`, which missed three configurations that really do run every statement on one connection: 1. `new Pool({ max: "1" })` and the legacy `new Pool({ poolSize: "1" })` — the shape `max: process.env.PG_MAX` produces. pg-pool never coerces the value, so the cap stays a string, and its own `_clients.length >= options.max` check then coerces it: the pool really is capped at one. 2. `postgres(url + "?max=1")` — postgres-js resolves `max` from the URL and does not coerce it either. 3. `PGMAX=1` with postgres-js — the same cap through the environment. On those three backends, a streaming export/import pair now throws `INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` where it previously hung, and two concurrent streaming imports now throw `INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS` where they previously interleaved and succeeded slowly — the lease is exclusive across all four pairings, so this one conservative refusal is inherited verbatim from the existing numeric `max: 1` behavior. `store.withWorkingCopy` and branch cloning also switch from a streamed clone to a fully materialized in-memory export on those backends, a memory-profile change on large graphs. Deliberately still unmarked: a postgres-js client given a non-numeric string cap other than one (`?max=5`), which opens exactly one connection today only because postgres-js does not coerce it — marking that would be marking on an upstream bug, and would refuse legitimate concurrent work the day the driver fixes it. A `pg` pool given `max: "5"` genuinely opens five connections and is likewise unmarked. Existing correctly-detected backends see no change: same marks, same refusal codes, same messages. Shipping in the same release, so nobody surprised by a new refusal is stuck: `createSqliteBackend` and `createPostgresBackend` gain an optional `serializedResource` declaration. `{ mode: "shared", resource: client }` marks a connection TypeGraph cannot detect — the `?max=5` shape above, Bun `SQL`, `expo-sqlite`, `op-sqlite`, `sqlite-proxy`, `pg-proxy` — and two backends naming the same object are one serialized resource. `{ mode: "independent" }` escapes a detection that is wrong for your topology. A `"shared"` declaration naming a different object than the one detected is refused with a `ConfigurationError` carrying `details.reason: "serialized-resource-conflict"` and a constructor-name description of each side (`details.declaredKind` / `details.detectedKind`, never the handles themselves — `details` is what `toLogString()` serializes, and a driver handle there would log the credentials that driver stores) rather than silently preferred. `"independent"` lifts the shared-resource refusal between two distinct backends — one SQLite backend exporting into itself still reports `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`, which is a fact about one handle rather than a claim about connection topology. That surviving refusal is SQLite-only, so on PostgreSQL a backend declared independent may export into itself. - [#479](https://github.com/nicia-ai/typegraph/pull/479) [`e285a71`](https://github.com/nicia-ai/typegraph/commit/e285a717a5d1df20868c3bc7a2c0f3a7914f5d40) Thanks [@pdlug](https://github.com/pdlug)! - Add constraint-aware ingestion branches that defer node uniqueness during staging and validate the resolved merge write set atomically. - [#476](https://github.com/nicia-ai/typegraph/pull/476) [`a58ba03`](https://github.com/nicia-ai/typegraph/commit/a58ba031b439f3befbcc073e016f896b178af00e) Thanks [@pdlug](https://github.com/pdlug)! - Add `repairInvertedValidityWindows`, the explicit operator action that makes rows an older version stored with a backwards window (`valid_from > valid_to`) observable again. Such a row is readable at no coordinate at all; upgrading deliberately rewrites nothing, so repairing is a decision an operator takes rather than a side effect of a deploy. `mode: "report"` counts and writes nothing — it reads through `execute`, a required backend member, so detection works on every backend including a history-capturing one and one with no statement-execution support. `mode: "apply"` normalizes the rows it counted to no lower bound ("ended at T, start unknown"), the shape today's write paths store, and is idempotent and convergent. `relations` is required and takes `"live"` or `"live-and-recorded"`; `"live-and-recorded"` is recommended, because repairing only the live axis leaves the recorded twin inverted and re-materializes the invisible row at any `asOfRecorded` coordinate. The repair mints no revision, bumps no `version` and does not move `updated_at`: it normalizes storage for rows that were never observable, so it is not a logical write. Run it with writers stopped, and re-baseline outstanding merge branches afterwards — `valid_from` is part of the `base@V` content fingerprint. `apply` refuses rather than guessing on the states it cannot honor: a backend without statement execution, a recorded-capture backend, and (on SQLite, where bounds compare as text) a relation holding non-canonical bounds it cannot classify. ## 0.47.0 ### Highlights TypeGraph 0.47 separates merge planning from execution with JSON-serializable `MergePlanArtifact` values. Applications can inspect the resolved write set and entity-resolution evidence, then apply the reviewed artifact against its target, schema, and revision fence without rerunning candidate generation or conflict callbacks. Existing one-call merge APIs retain their behavior. Operational Identity assertions gain explicit half-open validity windows, temporal contradiction checks, endpoint coverage validation, and window-aware interchange and merge behavior. Node and edge update/upsert APIs gain `clearValidTo: true` to reopen ended records, while dynamic collection results can participate directly in identity operations. Streaming exports also gain an optional consumer-idle timeout that releases their snapshot transaction and connection lease. ### Upgrade notes - Custom similarity scorers must return finite numbers; `NaN` and infinity now raise `MatchEvidenceError`. Candidate-source failures use `CandidateSourceError`, and deterministic merge constraint failures use `MergeConstraintConflictError` with the original store error as their cause. - With `coalesceUnchangedUpserts` enabled, an unchanged endpoint edge get-or-create returns `"found"`, not `"updated"`. The coalescing path requires the endpoint convergence fence and refuses with `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` on backends that cannot provide it. - Compiled-query metadata timestamps now use canonical fixed-width UTC ISO 8601 across drivers. Consumers comparing serialized timestamps should expect the same rendering as collection reads. - Reopening a validity window rechecks applicable constraints, including `oneActive` cardinality. Custom backends that cannot honor `clearValidTo` refuse it explicitly. ### Minor Changes - [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Add an opt-in `idleTimeoutMs` safety bound to `exportGraphStream`. The timeout measures how long a delivered chunk remains unacknowledged, then rolls back the snapshot transaction, releases the serialized stream lease, and reports the new typed `ExportStreamIdleTimeoutError`. Time spent waiting for the backend does not count as consumer idleness, and existing `AbortSignal` and cooperative `break`/`return` cancellation behavior is unchanged. - [#468](https://github.com/nicia-ai/typegraph/pull/468) [`c53c006`](https://github.com/nicia-ai/typegraph/commit/c53c0067f8a349e307f4f6d05e316172bcd726c2) Thanks [@pdlug](https://github.com/pdlug)! - Add public snapshot and incremental merge planning APIs that return stable, JSON-serializable `MergePlanArtifact` values. Plans bind the reviewed write set to the target graph, active schema, durable revision origin and revision, carry a content digest, and can be applied exactly once with `applyMergePlan`. Applying a plan validates the artifact and checks its fence atomically without re-running candidate generation, scoring, embeddings, canonical selection, or conflict callbacks. Existing `merge` and `mergeIncremental` entry points retain their one-call compatibility behavior while sharing the same resolution, validation, and mechanical write owners. Explain entity resolution with deterministic decisive edges, complete built-in candidate-source attribution, and scored strategy/score/threshold evidence while keeping definitional matches distinct from similarity scores. Add opt-in, deterministically bounded accepted/rejected candidate diagnostics. Default evidence excludes raw compared values. Custom similarity scorers that return `NaN` or infinity now fail with `MatchEvidenceError`; non-finite values cannot be represented faithfully in a serialized evidence artifact. Candidate-source failures now use the more specific `CandidateSourceError`; legacy `details.source` remains available alongside `details.sourceId` for base-source configuration failures. - [#473](https://github.com/nicia-ai/typegraph/pull/473) [`44d1486`](https://github.com/nicia-ai/typegraph/commit/44d1486938661bbde6d171bd0c99a0b269b597c3) Thanks [@pdlug](https://github.com/pdlug)! - Allow nodes returned by runtime string-keyed collections to participate directly in Operational Identity reads, assertions, bulk operations, and pair retractions. Identity result types now honestly include runtime-evolved members alongside compile-time graph references. - [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Add explicit half-open validity windows to scalar and bulk Operational Identity assertions, with bounded temporal contradiction checks, endpoint coverage validation, archival interchange support, node-window integrity guards, scalable branch-merge validation, and report-visible window reconciliation. - [#469](https://github.com/nicia-ai/typegraph/pull/469) [`8e50bdb`](https://github.com/nicia-ai/typegraph/commit/8e50bdbda614942b2848c6355b9cfb11c6468d2f) Thanks [@pdlug](https://github.com/pdlug)! - Add `clearValidTo: true` across node and edge update/upsert APIs so applications can reopen an ended valid-time window without changing entity identity. Built-in SQLite and PostgreSQL backends apply the clear, unchanged replays coalesce, `oneActive` relationships are rechecked when reopening, unsupported custom backends refuse explicitly, and graph merge carries branch-authored reopenings while rejecting delete-and-resurrect window artifacts. - [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Return `MergeConstraintConflictError` when a resolved graph merge would violate a deterministic store constraint, preserving the typed store error as its cause and exposing actionable constraint details. ### Patch Changes - [#470](https://github.com/nicia-ai/typegraph/pull/470) [`0b0022c`](https://github.com/nicia-ai/typegraph/commit/0b0022cf2b3c13483b468536404be4f35a8dfe39) Thanks [@pdlug](https://github.com/pdlug)! - Coalesce unchanged endpoint edge get-or-create updates when `coalesceUnchangedUpserts` is enabled. A coalesced replay now returns action `"found"`; `"updated"` means an update actually ran. The coalescing check needs the endpoint match-key convergence fence. On a backend without top-level transactions, such as Cloudflare D1 or `neon-http`, an otherwise unchanged endpoint replay now refuses with `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` instead of running unfenced. The bulk endpoint form and the create leg already required this fence. This option does not coalesce node `getOrCreateByConstraint` updates; use `upsertById` for replay projectors that need unchanged node writes to avoid history churn. - [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Canonicalize node and edge metadata timestamps returned by compiled queries. All supported database drivers now expose the same fixed-width UTC ISO 8601 rendering through compiled-query projections and store collection reads. - [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Serialize PostgreSQL `pg_trgm` extension installation across concurrent trigram index materializers. - [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Preserve PostgreSQL vector index build failures throughout serial-fallback preparation and durable `parallel_workers` cleanup, report the exact manual repair when cleanup fails, and reset built-in pgvector tables before every materialization attempt so recovery survives backend recreation without mutating custom strategy storage. ## 0.46.2 ### Patch Changes - [#459](https://github.com/nicia-ai/typegraph/pull/459) [`0e2afe2`](https://github.com/nicia-ai/typegraph/commit/0e2afe2d76381b5bc8309485d3a7e84c7bda092e) Thanks [@pdlug](https://github.com/pdlug)! - Extend `onImmutableLowerBound: "preserve"` to endpoint-matched edge writes. `getOrCreateByEndpoints` accepts the policy in its options, and `bulkGetOrCreateByEndpoints` accepts it per item alongside `validFrom` and `validTo`. The policy applies a stated `validFrom` on create or resurrection, while a live `ifExists: "update"` preserves the stored lower bound and still applies properties and `validTo`. Strict refusal remains the default. - [#462](https://github.com/nicia-ai/typegraph/pull/462) [`44f60cf`](https://github.com/nicia-ai/typegraph/commit/44f60cfbfaa6290323b02fb33fad2ca3541b1127) Thanks [@pdlug](https://github.com/pdlug)! - Surface lost fulltext contribution storage on gated operations as a typed `ContributionUnavailableError` with `state: "physical-storage-missing"` and rebuild guidance. Healthy operations retain the cached marker fast path; the error path translates only a missing-relation failure whose same driver error names the declared fulltext table. ## 0.46.1 ### Patch Changes - [#456](https://github.com/nicia-ai/typegraph/pull/456) [`a091902`](https://github.com/nicia-ai/typegraph/commit/a091902264fcbcd8336179c893a1e0a16eab528c) Thanks [@pdlug](https://github.com/pdlug)! - Add an explicit event-materializer policy for node upserts. Passing `onImmutableLowerBound: "preserve"` applies `validFrom` when the upsert creates or resurrects a row, but preserves a live row's stored lower bound while still applying props and `validTo`. The strict `IMMUTABLE_VALIDITY_LOWER_BOUND` refusal remains the default. The policy is available on `upsertById`, `upsertByIdFromRecord`, and each `bulkUpsertById` item, including unchanged coalescing replays. Widen the optional `better-sqlite3` peer range through 13.x and exercise 13.0.3 in this repository. Correct the event-log projector guidance to update existing endpoint-matched edges, document historical replay window requirements, and clarify that `MergeReport.validityEnds` only reports inherited-row claims. ## 0.46.0 ### Highlights TypeGraph 0.46 introduces Operational Identity: an opt-in profile for asserting that graph nodes represent the same or different entities. Typed Store, transaction, and temporal-view APIs expose identity membership, representatives, assertions, retractions, and history. Applications can choose same-ID folding across kinds or assertion-only identity, and use identity-expanded traversal without physically merging the underlying nodes. Identity truth travels through interchange and graph merge, with transactional contradiction checks and derived closure storage on SQLite and PostgreSQL. Merge handling also preserves inherited validity-window changes, retains parallel edges unless repointing causes a collision, and gives committed rows precedence when a collapse selects a survivor. This release completes the contribution maintenance sequence with read-only `probeContributions()`, non-destructive `repairContributions()`, and explicit `rebuildContribution()`. Validity-window validation is shared across write paths, and streaming exports support cancellation while holding a consistent snapshot. ### Upgrade notes - Audit graph definitions and persisted extension schemas before upgrading. Duplicate ontology relations, hierarchical self-loops, disjointness contradictions, incompatible inverse endpoints, multiple inverse partners, and unresolved extension edge names now fail validation on both construction and reload. - Operational Identity requires a supported transactional backend. Use ordinary interchange import for identity-bearing data; trusted import refuses identity-enabled targets and identity-bearing streams. Exports now use format `2.0`, while imports still accept `1.0`. - Custom backend table-name resolutions and `SqlSchema` subclasses must provide the identity assertion, recorded identity assertion, and closure relations. Restore missing assertion ledgers from backup; a missing derived closure can be recreated and rebuilt before traffic resumes. - `StoreView` and `RecordedStoreView` retain construction and `instanceof` support but can no longer be subclassed. Exhaustive merge-report and import-error consumers must handle the new `"identity"` entity kind, and tooling that assumes `FORMAT_VERSION` is the literal `"1.0"` must be updated. - Writes now refuse inverted validity windows and live-row updates that state a different immutable `validFrom`. Trusted imports require canonical UTC validity timestamps. Node creation or upsert on a soft-deleted same-kind ID now resurrects the row instead of exposing a storage constraint error. - Changing `identity.sameIdAcrossKinds` requires an explicit schema migration, which rebuilds closure with the schema commit. The type-level `sameAs` and `differentFrom` factories are deprecated in favor of Operational Identity. ### Minor Changes - [#427](https://github.com/nicia-ai/typegraph/pull/427) [`facef56`](https://github.com/nicia-ai/typegraph/commit/facef560d14c38607d6414818c65e43dc65a88d2) Thanks [@pdlug](https://github.com/pdlug)! - Add `signal` to `exportGraphStream`, and keep a streaming import's statistics refresh inside its connection lease. An export stream holds one `repeatable_read` / `read_only` transaction for its whole life, and on a single-connection backend it holds that connection's exclusive interchange-stream lease with it. Every cooperative exit already settled both, because each runs the generator's `finally`: `break` or `throw` out of a `for await`, and an explicit `iterator.return()`. A consumer that pulls `next()` and then simply DROPS the iterator — the `Promise.race([iterator.next(), timeout])` pattern — has no cooperative exit, because async-generator `finally` blocks do not run on garbage collection. That leaked the snapshot transaction for the life of the process, and with it the lease, so every later export and every later import on that connection was refused for a stream nobody was reading. `ExportOptionsSchema` now accepts `signal?: AbortSignal`, so both `exportGraph` and `exportGraphStream` take it, on both capability arms — a non-transactional export holds no snapshot and no lease, but it still owes its consumer an answer rather than a silent stall, and the same cancellation path gives it one. On a transactional backend, aborting rolls the snapshot transaction back and releases the lease whether or not anyone is waiting on `next()`; the pull that is in flight when the abort lands — and a pull from a consumer that walked away and came back — rejects with the new `ExportStreamCancelledError` (`code: "INTERCHANGE_EXPORT_STREAM_ABORTED"`), carrying the signal's own `reason` as `cause`. A signal that is already aborted refuses the export before any transaction is opened or any lease claimed. The listener is subscribed before anything is claimed or opened and `signal.aborted` is re-checked immediately after subscribing, so an abort at any instant is either seen by that re-check or delivered to the listener — including one raised synchronously by a driver inside `backend.transaction(...)`, which an `AbortSignal` never replays to a listener that arrives later. Everything else is unchanged: an export without a signal behaves exactly as before, and a cooperative exit still reports a clean end rather than a cancellation. There is deliberately no garbage-collection fallback. A `FinalizationRegistry` on the iterable cannot work here — not merely unreliably, but never: the producer is interruptible only where it is parked waiting for the consumer, so any cleanup state able to settle an abandoned stream must reach the stream's internal channel, and a registry holds its held value strongly, so holding anything that reaches that channel keeps the abandoned stream permanently reachable and the entry can never fire. The signal is the mechanism, and it is a contract rather than a hint. Separately, `importGraphStream` now holds its target connection's stream lease across the trailing planner-statistics refresh instead of releasing it when the chunk loop ends. That `ANALYZE` is a write like the chunks were, and running it outside the lease left it to be stranded by an export snapshot opening in that window — swallowed as a warning, because the refresh is best-effort. `importGraph` never had the hole (`withImportStreamLease` spans its whole call), so this also removes a divergence between the two import surfaces. The lease is still released on every exit, including the error paths. Also fixes a silent cross-kind edge overwrite in import. Edge ids are unique per graph but the import's existence probe (`getEdge` / `getEdges`) is keyed on `(graph_id, id)` with no kind comparison, so a document edge of kind A whose id was already held by a kind-B row matched that row: `onConflict: "update"` wrote A's properties onto the kind-B row with nothing in `result.errors`, and `onConflict: "skip"` counted the document's edge as already present when no edge of its kind existed. Both are now reported as a per-row `ImportError` prefixed `INTERCHANGE_EDGE_KIND_CONFLICT`, naming the stored kind and the stated one, with the stored row left untouched — the check runs before the conflict strategy, so all three strategies answer alike. `backend.updateEdge` is additionally called with `kind`, which `UpdateEdgeParams` documents as MUST-apply, so the predicate lives in the UPDATE's own `WHERE` and the check cannot be raced by a concurrent hard-delete-and-recreate; a write that consequently matches no row is reported as the same per-row error rather than aborting the import. Nodes were never affected — their probe is kind-scoped. The export's snapshot guarantee is now stated as the capability-scoped fact it is, in the API docs, the option docs, the error class, and the abort message: a backend reporting `capabilities.transactions` reads the whole export inside one repeatable-read transaction, while one without (SQLite `transactionMode: "none"`, session-less HTTP Postgres drivers) paginates statement by statement and can show a mid-stream write in later pages. `ExportStreamCancelledError`'s message now says which of the two it is describing, so a cancelled non-transactional export no longer claims to have rolled back a snapshot it never opened. - [#417](https://github.com/nicia-ai/typegraph/pull/417) [`9d3014c`](https://github.com/nicia-ai/typegraph/commit/9d3014c05f2936fcdadb3fa50950445a0d8e2652) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: judge the edge fold's property union against base, and report a target-precedence window discard Two adjacent gaps in the edge repoint/window path. One is a bug fix, the other adds an optional field to a report type, so this ships at the higher `minor` bump and covers both. The repoint fold's property union had no base to compare against, unlike the node path's three-way merge, so a staged copy of an INHERITED edge contributed its whole fork property bag as first-class `(branch, value)` claims — including the values it never touched. Under any rank-based `onPropertyConflict` an untouched base value could therefore outvote a value a branch actually authored, decided by whichever branch label happened to ride on the untouched copy. The window-only carrier made it observable: an inherited row whose only change is its end-of-validity is staged solely to give that ending somewhere to ride, its properties ARE the base's, and its branch is merely whichever sorted first in staging. The union now filters every contributor to the properties it CHANGED from its own base — a branch-created edge has no base, so everything it carries stays a full claim — which means a carrier contributes no claim and raises no conflict at any rank. Genuine disagreements are unaffected: two members that changed one property differently still conflict, over their real values alone. Filtering claims does not erase content: the folded row commits the same property set as before, and a key only a non-survivor carries keeps the value held by the member with the minimum edge ID — the row, never the branch label riding on it, since for these keys no branch claimed anything and an arbitrary label deciding the committed value is the very thing being fixed. `MergeReport.validityEnds` now also reports the window claims that target precedence discards. When the incremental target had already moved an inherited row's end, the reconciler took the row out of the resolution and the branch claims vanished from the report entirely — less visible than a claim that merely lost the least-claim rule, which stays named in `claimedBy`. Such a row now gets a resolution naming the target's own committed instant, its discarded claimants, and the new optional `ValidityEndResolution.precedence` field set to the exported `VALIDITY_END_TARGET_PRECEDENCE`. The field is absent on every entry the merge itself decided, so existing consumers read what they always read; no write is staged and no provenance credit is minted for such a row, and a row no branch claimed still produces no entry at all. - [#364](https://github.com/nicia-ai/typegraph/pull/364) [`fb29816`](https://github.com/nicia-ai/typegraph/commit/fb2981664c77d704b7f78933b8f887222c796091) Thanks [@pdlug](https://github.com/pdlug)! - Add optional `topK` to `pageRank()` and `personalizedPageRank()`, and optional `minComponentSize` to `weaklyConnectedComponents()`. Both bound only result extraction: the limit and the inclusive component-size filter are applied in extraction SQL after the existing deterministic ordering, so bounded rows never reach the driver. Default results and ordering are unchanged, and the graph computation itself still runs over the whole visible induced subgraph. - [#415](https://github.com/nicia-ai/typegraph/pull/415) [`b68e643`](https://github.com/nicia-ai/typegraph/commit/b68e6437a337cdcd2e3c166754deb87008e25152) Thanks [@pdlug](https://github.com/pdlug)! - Refuse non-canonical validity-window timestamps in trusted import. `trustedImportGraph` / `trustedImportGraphStream` accept a pre-typed stream and never re-parse it, so a `validFrom` / `validTo` that TypeScript types as `string` but is not canonical fixed-width UTC ISO 8601 used to flow straight to SQL. Every temporal filter compares those values AS TEXT against an `asOf` coordinate, so a stored `"2021-01-01"`, `"...T00:00:00Z"`, `"...:00.1Z"` or `"...+01:00"` mis-sorts and silently includes or excludes the wrong rows — and it mis-decided the negative-width window check that the same path performs on the way in. Both window fields of every streamed node and edge are now format-checked with the same `isCanonicalIsoDate` decision the untrusted import schema and the store's own writes make. A violation refuses the WHOLE stream with a `TrustedImportError` carrying the existing reason `invalid_stream`, naming the offending field, row and value; the session's transaction rolls back, so chunks already streamed are not left behind. This is a behavior change: a stream that previously imported and stored an unsortable timestamp now fails loudly. Convert such values with `new Date(value).toISOString()`. The check is format-only — trusted import still skips property, reference and conflict validation — and it leaves an absent field and an explicitly `null` (confirmed open-left) `validFrom` untouched. Also documents a pre-existing bulk-API limitation, with no behavior change: `bulkUpsertById` groups every create ahead of every update, so one batch cannot hand a constrained value from one row to another (releasing a `unique` value or a `oneActive` edge slot and claiming it in the same batch throws `UniquenessError` / `CardinalityError`, where the equivalent sequential upserts succeed). The workaround is two batches, or sequential upserts. - [#374](https://github.com/nicia-ai/typegraph/pull/374) [`fadf932`](https://github.com/nicia-ai/typegraph/commit/fadf93297df40bd619a1ca45b165edc04ef6ebfe) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: stage cascade retractions with their cause instead of inferring intent A node soft-delete ends every open identity assertion touching the node, so a branch that deletes a node stages retractions it never asked for. The merge previously separated those cascade endings from a branch's own retraction with a conservative branch-level heuristic, which deliberately over-dropped the same-branch case: a branch that retracted an assertion and LATER deleted one of its endpoints looked exactly like a pure cascade, so its retraction was dropped whenever the deletion was overruled — silently keeping truth the branch had explicitly ended. The soft-delete cascade now ends assertions at the deleted node's own `deleted_at`, which makes the cause derivable: the state-diff compares each retracted assertion's end instant to the deletion instants of its endpoints and stages the retraction as either a cascade naming the deleted node or the branch's own act. The merge planner drops a retraction only when EVERY branch staged it as the cascade of a deletion that delete/modify resolution then overruled, so an explicit retraction survives even when it comes from the deleting branch. Two cases stay conservative because nothing distinguishes them at the stored resolution — a hard delete (which removes the assertion rows) and a retraction issued in the same millisecond as the delete that followed it. - [#387](https://github.com/nicia-ai/typegraph/pull/387) [`71361d7`](https://github.com/nicia-ai/typegraph/commit/71361d70b1fda83ad8539228c8e133abd0ce57f9) Thanks [@pdlug](https://github.com/pdlug)! - Complete the contribution health lifecycle with a read-only readiness probe and an explicit destructive rebuild, so the three maintenance operations form one escalation ladder: `probeContributions()` (writes nothing) → `repairContributions()` (non-destructive, already shipped) → `rebuildContribution()` (drops and recreates storage). `store.probeContributions()` answers "is search coherent with the graph right now" without mutating anything — safe on a read path, on a replica, and under a least-privilege role. It returns one `ready` / `degraded` entry per search projection plus the durable `graphRevision` the assessment was taken at on a revision-tracked Store. It shares the detection logic of `verifyContributions()` rather than reimplementing it, so a health check can never disagree with the gate the hot path actually consults. A projection with no declared contributions is omitted rather than reported `ready`, and a backend that provisions contributions but cannot probe its catalog refuses instead of answering — "assessed and healthy" and "never looked" never share a return value. `store.rebuildContribution("fulltext")` is the repair that was missing for a `stale` contribution, whose table exists at a shape the current `createDdl` no longer produces: the ensure path's `CREATE ... IF NOT EXISTS` no-ops against it, so re-stamping the marker would leave it blessing storage of the wrong shape. The rebuild drops the storage, recreates it, reconstructs the content from the node rows, and stamps the marker inside one transaction under the schema-write fence, so an interrupted rebuild rolls back rather than leaving storage attested but empty. It is reachable only by name — never from `repairContributions()`, which continues to report these findings as `requires-rebuild`. Vector contributions are not rebuildable, and the call refuses with `ContributionRebuildUnsupportedError` rather than dropping them: TypeGraph stores the vectors callers supply and never the inputs that produced them, so the embeddings exist only in the storage a rebuild would destroy. `reembedVectorField(kind, fieldPath, { embed })` remains the sanctioned destructive path, because it takes the callback that can regenerate them. The same typed error covers a fulltext strategy that declares no `dropDdl` and a backend with no transactional schema fence; all three refuse before anything is dropped, and all three are declared ahead of time on the new `backend.capabilities.contributions` capability. Fixes the drift guard so the ladder can actually be climbed: when the guard refused a shape change it recorded the failed attempt at the _new_ signature, overwriting the only evidence of the shape the table really had. The verdict then read as `missing-marker` rather than `stale`, so `repairContributions()` reported it repaired — re-stamping the marker over the unchanged old-shape table — and the next boot skipped the guard entirely. The guard now preserves the recorded signature, so a `stale` contribution stays `stale` across restarts, `repairContributions()` keeps reporting `requires-rebuild`, and the refusal persists until `rebuildContribution("fulltext")` fixes the shape. Reach that call from a `createStore()` / `createVerifiedStore()` Store: the managed factory's boot step is what the guard refuses. Also adds optional `dropDdl` to `TableContribution` — declared by both bundled fulltext strategies — which is what opts a strategy into the rebuild. - [#376](https://github.com/nicia-ai/typegraph/pull/376) [`8c3a8e6`](https://github.com/nicia-ai/typegraph/commit/8c3a8e6af5ec2813a26d0aa13bf58da2c50fbaa3) Thanks [@pdlug](https://github.com/pdlug)! - Add a database-level contradiction backstop for Operational Identity. A `different` assertion and a `same` assertion that would place both of its endpoints in one identity class are a contradiction, and until now only application code stood between such a write and a committed graph: the plan-time simulation and the identity applier's validation both decide by reading state and comparing, so a bug in either commits the contradiction silently. Identity now also maintains a derived **separation relation** — one row per pair of identity classes a current `different` assertion holds apart, keyed by the two class keys under a `CHECK (class_key_low < class_key_high)` constraint. Every transaction that fuses two classes relabels the affected separation rows in the same statement batch, so fusing two separated classes relabels both sides of their shared row to one key and the database aborts the transaction. A write that reached the ledger through a path that skipped identity validation can no longer commit a contradictory graph; it fails with the new typed `IdentitySeparationViolationError`. The relation is derived and requires no application changes: it is maintained wherever the identity closure is (assert, retract, fold, delete, merge, import, rebuild), `rebuildIdentityClosure(store)` recomputes it from the assertion ledger, and store-open identity validation checks it against that recomputation. Upgrading an existing identity-enabled database needs no manual step. A store opened with `createStoreWithSchema` / `createAdapterStoreWithSchema` creates the new `typegraph_identity_separation` relation through the same idempotent identity DDL path as the other identity relations and recomputes it from the ledger once, before anything reads it. A missing assertion ledger or closure relation is still refused as data loss. Custom backend authors: the resolved table-name types (`ResolvedSqlTableNames`, `SqliteTableNames`, `PostgresTableNames`) and the `ensureIdentityTables` parameter each gained a required `identitySeparation` entry, so an implementation that builds one of those objects needs the new name added. Code that only reads `backend.tableNames`, or that passes a partial name override to `createSqliteTables` / `createPostgresTables` / `createSqlSchema`, is unaffected — an omitted name still resolves to the default `typegraph_identity_separation`. - [#397](https://github.com/nicia-ai/typegraph/pull/397) [`d2b935e`](https://github.com/nicia-ai/typegraph/commit/d2b935ed071fc2eb8e4310cfb1bdeecca64072b0) Thanks [@pdlug](https://github.com/pdlug)! - Graph Merge now keeps the **inherited** edge when a repoint-induced collapse folds a committed row together with a branch-created one, instead of keeping the lexicographically-minimal edge id. Previously the survivor of such a collapse was whichever edge id sorted lowest. A collapse rewrites the row it keeps and ends none of the rows folded into it, so when a branch-created id sorted below the committed one, the merge wrote the branch's row as a new edge and left the committed edge live beside it at its pre-merge properties — two live rows for one folded relationship, the edit staged for the committed row never written, and `merged.edges` counting one of them. Which of the two you got depended on an id sort, so it was not behavior a caller could depend on. The surviving edge id reported in `PropertyConflict.entityId`, window resolutions and provenance is consequently the inherited row's id whenever the collapse involved one. That is the id of the row that actually persists, and it no longer moves with branch-created id lexicographics. Collapses among branch-created edges alone are unchanged, as is the property/window reconciliation applied to the survivor. - [#427](https://github.com/nicia-ai/typegraph/pull/427) [`facef56`](https://github.com/nicia-ai/typegraph/commit/facef560d14c38607d6414818c65e43dc65a88d2) Thanks [@pdlug](https://github.com/pdlug)! - `store.rebuildContribution("fulltext")` is now scoped to the graph it is called on. The fulltext projection is one physical table holding every graph's rows keyed by `graph_id`, while the rebuild runs under the per-graph schema fence — so the old unconditional `DROP TABLE` destroyed every other graph's search index on the same database, with no concurrency required: a neighbouring graph's `fulltext` marker survived the drop (markers are keyed by `graph_id`), so `probeContributions()` and `verifyContributions()` went on reporting it `ready` while every search it served returned nothing. The rebuild now removes only the calling graph's rows, through the same `DELETE ... WHERE graph_id` statement `clear()` uses — one exported builder both call — and escalates to dropping and recreating the shared table only when that table holds no other graph's rows. That drop remains the one repair for storage provisioned at a shape the current DDL no longer produces, and the lock scope now matches the decision it authorizes. Two locks, protecting different resources: a constant-keyed advisory lock (`typegraph:contribution-ddl`) serializes the contribution's DDL across graphs — it survives the drop and exists even when the table does not, which is what a relation lock cannot do — and, on the path that may drop, `LOCK TABLE ... IN ACCESS EXCLUSIVE MODE` excludes ordinary writers, which take no advisory lock at all and could otherwise commit a row between the probe and the `DROP TABLE` that the probe had already decided was safe. The verdict is re-established under that lock before any drop; the cheap unlocked probe ahead of it exists only to keep the graph-scoped path off the relation lock, and can only err toward keeping the table. Both are no-ops on SQLite, whose `BEGIN IMMEDIATE` fence already holds the database's single writer slot from probe through commit. When the recorded shape is `stale` — the state only a recreate repairs — and the storage that would have to be recreated holds another graph's rows, the rebuild refuses with `ContributionRebuildUnsupportedError` and the new reason `shared-storage-in-use` instead of either destroying content it cannot reconstruct or re-stamping this graph's marker over a physical shape nothing verified. The refusal names the other graph ids and the sanctioned maintenance-window sequence. Vector contributions are unaffected: their storage is per-`(graph, kind, field)`, so no other graph's data is ever in reach. - [#427](https://github.com/nicia-ai/typegraph/pull/427) [`facef56`](https://github.com/nicia-ai/typegraph/commit/facef560d14c38607d6414818c65e43dc65a88d2) Thanks [@pdlug](https://github.com/pdlug)! - Harden the Operational Identity release and adjacent write paths found during its adversarial review. Identity interchange now exports one repeatable-read snapshot, uses target-bound keyset pagination pinned to code-point order via the dialect adapter's `binaryText` seam so a `base@V` content token minted on 0.45 still matches its recomputation on PostgreSQL, cancels cleanly, and refuses streams that would deadlock a serialized connection — a PGlite connection, a bare `pg`/neon `Client` (including a checked-out `PoolClient`), a `Pool` explicitly configured with `max: 1`, a postgres-js client built with `{ max: 1 }`, a better-sqlite3 handle, a `bun:sqlite` database, a sql.js database, a local (`file:`/`:memory:`) libSQL client, or Cloudflare Durable Object storage, whose transaction frame is ambient on the storage object. The refusal is one EXCLUSIVE long-lived-stream lease per serialized resource, not a one-time observation and not a cross-kind-only exclusion: at most one interchange stream of any kind holds a given connection, so all four pairings are refused rather than only the two that mix kinds — an import behind an export snapshot (including through a user-wrapped stream that no longer identifies its source backend), an export snapshot behind a streaming import, and now export-behind-export and import-behind-import too, which previously reached the driver as a nested `BEGIN` after chunks had already committed. Whichever long-lived stream starts second gets a typed `ConfigurationError` instead of both hanging: `details.code` names the condition that holds the connection (`INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` behind an export snapshot, `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT` when the object-identity detector is what answered, or the new `INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS` behind another import), while the new `details.requested` and `details.heldBy` name which pairing was actually refused, so a same-kind refusal is never reported as something it is not. `"import-stream"` is the kind of EVERY long-lived import, not only the chunk-streaming one: `importGraph` takes the lease for the whole call and `trustedImportGraphStream` / `trustedImportGraph` hold it for the whole trusted session, so both APIs can now throw these serialized-connection `ConfigurationError` codes — a new error TYPE on a trusted-import surface that previously threw only `TrustedImportError`. Every exit releases the lease, including a mid-stream producer failure and a synchronous throw out of `backend.transaction(...)` (a closed handle, a refused pool checkout), which would otherwise strand the connection for the life of the process. Relatedly, SQLite's manually framed transactions no longer let a failing `ROLLBACK` mask the failure that caused it: SQLite auto-rolls-back on `SQLITE_FULL` / `SQLITE_IOERR` / `SQLITE_NOMEM`, so the unwinding `ROLLBACK` can itself fail with "cannot rollback - no transaction is active" — the caller now receives the ORIGINAL error and the rollback failure is warned instead of thrown. Graph merge uses injective composite keys, preflights provenance sidecar collisions, and refuses merge options it cannot honor instead of ignoring them. Edge identity checks now include kind and endpoint kind on every create, delete, and get-or-create path, including tombstoned rows, and that check is carried by the write statement itself rather than re-derived beside it: `UpdateEdgeParams`, `DeleteEdgeParams`, and `HardDeleteEdgeParams` each gain an optional `kind`, and a backend that receives it MUST scope the statement to that kind. Both bundled Drizzle backends satisfy that contract through one shared predicate, but a hand-written `GraphBackend` has to honor it or it will silently widen a write it was told to narrow. Because a kind-scoped statement that also requires `deleted_at IS NULL` is its own recheck, the redundant in-transaction re-read that used to precede each edge delete and hard delete is gone — the statement either matches the edge the caller named or affects nothing. Node `bulkDelete` remains one atomic, hookless bulk operation exactly as in 0.45. Edge `bulkDelete` changes behavior in 0.46: 0.45.x looped single deletes and fired per-item `onOperationStart`/`onOperationEnd` hooks for each one, and 0.46 makes it one atomic single-transaction batch that emits NO hook events at all — neither per-item nor bulk, since `onBulkOperationStart`/`onBulkOperationEnd` fire only for node `updateWhere` and no bulk-hook coverage for deletes exists yet — so a consumer that relied on those per-item events for audit or metrics must either keep deleting individually (single `delete` still fires per-item hooks) or capture the deletions another way, such as from the ids it passes in and the rows it reads back; an id in the batch that belongs to another edge kind is refused with `ValidationError` carrying `EDGE_IDENTITY_MISMATCH_CODE`, rolling back every delete already applied earlier in the same batch. Constrained writes no longer take their decision from a read the write cannot vouch for. Every write whose correctness rests on a check-then-write — edge cardinality `one`, `unique`, and `oneActive` (including the create and resurrect legs of `getOrCreateByEndpoints`, single and bulk), node-kind disjointness on create, and a `kindWithSubClasses` uniqueness constraint that actually expands to more than one kind (a scope covering a single kind probes exactly the row the uniques table's own primary key then reserves, so that key IS its fence) — now runs its probe and its write under the same per-graph mutual exclusion, whether or not the store enables `history` or `revisionTracking`. That exclusion previously arrived only as a SIDE EFFECT of recorded capture's advisory lock, so the DEFAULT PostgreSQL store — no history, no revision tracking — raced: two writers each probed a graph that satisfied the constraint and each committed, producing exactly the duplicate the constraint exists to prevent, with no error on either side. On PostgreSQL the fence is that same transaction-scoped advisory lock, now taken for the constraint's sake rather than the clock's; on SQLite it is the `BEGIN IMMEDIATE` writer slot the backend already holds. Writes with nothing to check — an unconstrained create, any delete, a cardinality-`many` edge — take no lock at all, so the cost tracks the constraints a graph actually declares rather than becoming a blanket serialization. A backend running WITHOUT transactions has no fence to take, and a constrained write there is now REFUSED rather than run unfenced. Both halves of the fence are transaction-scoped constructs — SQLite's `BEGIN IMMEDIATE`, PostgreSQL's `pg_advisory_xact_lock`, which outside a transaction is acquired and dropped inside its own implicit single-statement one and excludes nothing — so "can this backend fence" and "does this write run inside a transaction" are the same question, and the refusal is keyed on exactly that. It is a `ConfigurationError` with `details.code` `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` and `details.constraint` naming WHICH declared constraint needed the fence — `edgeCardinality`, `edgeMatchKeyConvergence`, `nodeDisjointness`, or `nodeUniquenessScope` — because "this backend cannot fence constrained writes" is unusable advice while "your `cardinality: 'one'` edge cannot be enforced here" is actionable; the `suggestion` carries the per-class way forward. The blast radius is Cloudflare D1 (auto-detected as `transactionMode: "none"`), `drizzle-orm/neon-http`, and any SQLite backend explicitly built with `transactionMode: "none"` — Durable Objects are NOT affected, since `do-sqlite` reports `capabilities.transactions: true` and fences normally. On those three, a declared-constraint write that previously raced silently now throws: a `cardinality` other than `many` cannot be created or resurrected, a `disjointWith` kind cannot be created, a `kindWithSubClasses` unique that actually expands past one kind refuses on create AND update, and `getOrCreateByEndpoints` can no longer take its CREATE leg (a call that FINDS an existing edge still returns it, and a `many` resurrection is an id-keyed UPDATE re-deriving no verdict, so both keep working; the BULK form fences its whole batch and therefore refuses whatever the outcome would have been). Everything unconstrained is untouched — a `many` edge created, updated and deleted, any node delete including one whose kind participates in a disjointness axiom, and a `scope: "kind"` unique whose uniques primary key IS its fence — so this is not a blanket loss of write access on those engines. Refused rather than degraded, per the accepted-or-refused rule: a constraint enforced only when nothing races is the exact defect the fence exists to close, and reporting it as enforced would make the invariant above false precisely where it matters. A store created with `coalesceUnchangedUpserts` likewise stopped letting an optimization change an answer. A single-node `upsertById` decided to skip its write from an autocommit read, so a writer committing between that read and the skip left the caller told its props were stored while the store in fact held the other writer's — a DIFFERENT outcome from the same call with the flag off, where the update's own in-transaction re-read merges the caller's props over whatever it finds. The skip is now taken only on evidence re-read inside a transaction after that first observation, and a losing verdict falls through to the ordinary write path; only an upsert that is about to coalesce pays the second read, so a store without the flag, or one whose props differ, keeps the single read and single write it always had. `getOrCreateByEndpoints` now converges rather than retrying once and hoping. Its single-shot retry became a bounded loop of three attempts whose ordinary case is cheaper than before — under the fence a losing writer learns of the winner from its own in-transaction lookup instead of from a `CardinalityError` — and whose exhaustion is a typed refusal rather than a livelock or a stray constraint error leaking out of a lost race: a competitor that repeatedly creates and removes the same match key ends in a `DatabaseOperationError` naming that pattern and telling the caller to serialize or retry. The bulk edge `getOrCreateByEndpoints` and bulk node `getOrCreate` paths, which had no retry at all and surfaced a concurrent winner as a raw constraint violation, gained one. `efSearch` is now refused everywhere it cannot be applied, not only on PostgreSQL. The guarantee that vector search never silently drops an accepted option held only for the PostgreSQL path: every SQLite backend accepted `efSearch` and ignored it, on the vector path and the hybrid path alike (the hybrid path dropped it in a second place, while rebuilding the vector parameters), so a caller tuning recall got the default frontier and no indication. All of them now ask one owner, and an engine with no per-search ANN frontier refuses with `UnsupportedBackendCapabilityError` — `details.capability` `vector.searchFrontierTuning`, `details.reason` naming the limitation (`sqlite-vec`'s `vec0` KNN takes only `k`, the page size; libSQL's DiskANN `vector_top_k` fixes `search_l` at index-creation time). PostgreSQL is unchanged, including its existing refusals for a non-HNSW slot and a driver that cannot scope `SET LOCAL`, which now come from that same owner instead of from a second spelling of the same decision. Separately, sql.js backends could not execute a compiled query at all. The compiled-execution adapter recognized any client exposing `prepare()` and then called `all()` on the resulting statement, but sql.js's `Statement` has no `all()` — it is a `bind`/`step`/`getAsObject`/`free` cursor — so the first prepared statement threw. Client detection is now shape-specific, sql.js is excluded from the compiled path, and it runs through Drizzle's own session, which drives that cursor correctly; `bun:sqlite`, whose statement DOES expose `all()`, keeps the compiled path. The merge-provenance sidecar now claims its graph id MARKER-FIRST instead of inferring ownership from circumstantial evidence: the durable `ProvenanceOwner` marker is the sidecar's FIRST write of any kind, committed inside the schema fence (`schemaWriteTransaction` — the same per-graph fence every schema commit and schema-managed write already respects) and BEFORE the sidecar schema is registered, which is possible because the marker is a plain node row needing no per-graph DDL. A competing writer therefore either commits first and is seen, or waits until the claim has committed; a refused open leaves the occupant byte-identical, writing no schema row, no marker, and no provenance row. Because the marker precedes the schema, the resumable interrupted state is marker-WITHOUT-schema (or marker beside a pre-marker legacy schema), which resumes by registering or migrating the schema — while a graph carrying the exact current sidecar schema with NO marker is a state this module cannot produce and is refused UNCONDITIONALLY as `unowned-exact-schema-graph`, contents never consulted: empty and provenance-shaped occupants are refused exactly like any other, since contents an application could have written are not evidence of authorship. Freedom is judged by occupancy across EVERY per-graph row table the backend names through its `tableNames` port — nodes and edges, but equally recorded-time history, the revision clock and origins, identity assertions with their recorded ledger, closure and separation, fulltext, and unique keys — so a graph id holding only, say, identity or fulltext rows is occupied and refused, and an unregistered schema is never taken as evidence of a free namespace. Only the exact validated live marker counts as ownership: a tombstoned, malformed, wrong-target, or non-canonical `ProvenanceOwner` row refuses with the reason `corrupt-ownership-marker` and is never overwritten or resurrected. Refusals report one of five typed reasons under `GRAPH_MERGE_PROVENANCE_ID_COLLISION` — `application-graph`, `empty-legacy-sidecar`, `unupgradeable-legacy-sidecar`, `unowned-exact-schema-graph`, or `corrupt-ownership-marker` — each carrying remediation specific to the state actually found, and a backend that exposes no schema fence refuses an unclaimed sidecar with `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED` rather than claiming without atomicity (an already-owned sidecar needs no claim and still opens there). One writer class takes neither the per-graph advisory lock nor the active schema row — a schema-LESS raw `createStore` writer, or a direct `backend.insertNode` / `insertEdge` — and at PostgreSQL's READ COMMITTED its insert could commit between the claim's fenced re-inspection and the claim's commit, leaving the marker on a graph id an application had just made its own. That window is closed rather than accepted: on PostgreSQL the claim issues `LOCK TABLE , IN SHARE ROW EXCLUSIVE MODE` inside the fence and before the re-inspection, draining in-flight row writers and holding new ones off until the marker commits. `SHARE ROW EXCLUSIVE` and not `SHARE` because the mode must be SELF-exclusive: two concurrent claims on different sidecar ids hold different advisory locks, so under `SHARE` both would acquire it and then both request `ROW EXCLUSIVE` for their own marker INSERT — a lock-upgrade deadlock PostgreSQL resolves by aborting one. The cost is real and bounded: while a claim runs, every node and edge write on the whole DATABASE waits, for the duration of a few probes and one INSERT with no caller code inside — and the lock is taken only when a sidecar is created, upgraded from the pre-marker schema, or resumed after a crash, never on the common path where an already-owned sidecar opens with no fence at all. SQLite takes no such lock; `BEGIN IMMEDIATE` already owns the single writer slot. `persistProvenance: true` is now honored or refused, never dropped: the sidecar is opened and claimed PRE-COMMIT, so an occupied sidecar graph id or a backend that cannot fence the claim refuses the whole merge as `InvalidMergeOptionsError` (`details.option` `"persistProvenance"`, with the originating `ConfigurationError` as `cause` and its code echoed as `details.provenanceErrorCode`) and leaves the target unmodified — where 0.45 would have committed the merge and reported the same configuration verdict as a `warnings` entry. Those verdicts are as true before the merge as after it, so reporting them post-commit left the caller with a committed graph and a stated option TypeGraph had silently ignored. The post-commit best-effort warning path survives only for what is genuinely transient — a row write failing against a sidecar this library already owns. One visible consequence of claiming early: after a `persistProvenance` merge the sidecar (marker and schema) exists even if the merge itself later fails, holding an owned, empty sidecar and no target change. `mergeIncremental`'s refusal of a non-`"flag"` `onBasePropertyConflict` is now a typed `InvalidMergeOptionsError` (`MERGE_ERROR_CODES.invalidOptions`, category `user`) instead of a plain `MergeError`, so it is catchable the same way every other refused merge option is. `mergeIncremental` additionally refuses a fork point that moved under it. Its plan is a set of diffs against one `base@V`, and only the TARGET was ever allowed to advance while the merge ran — but nothing checked, so a write landing on the fork-point store mid-call left the commit applying diffs against an ancestor that no longer existed. The fork point is now frozen for the duration of the call: the version read before planning is carried into the commit and re-compared as the first act of the commit transaction, and a mismatch raises `BaseVersionMismatchError` (`GRAPH_MERGE_BASE_VERSION_MISMATCH`) naming the expected and live fork-point bases rather than committing. And `branch()` no longer leaks the working copy's backend when the post-transfer schema-anchor read fails — ownership of that engine transfers to `branch()` on the strategy's success path, and a `branch()` that reports failure as `err(...)` hands the caller no handle to close. Operational Identity's derived separation relation is never published in a state that under-reports separations. The relation is created INSIDE the transaction that fills it — the fenced path issues its DDL under `schemaWriteTransaction`, the schema-commit path returns the DDL as data for the commit transaction to issue — so a commit refused by the `IDENTITY_PROFILE_MIGRATION_PENDING` gate, a stale CAS, or a contradiction now creates nothing at all, where previously it stranded a readable, empty relation that the next open skipped because "present" was what suppressed the rebuild. What that cannot undo, a per-graph predicate heals: the fill decision is "does THIS graph hold live `different` assertions and no separation rows", not "does the table exist", so a relation left empty by an older version or by another graph's provisioning is rebuilt at the next open of the graph that owns the assertions. The predicate is exact in both directions — it shares the fill's registry kind filter and additionally requires the assertion's two endpoints to resolve to DIFFERENT identity classes, since a contradicted ledger projects to a degenerate pair the relation's CHECK refuses, which is its own fault with its own error rather than an unfilled relation. Identity DDL is serialized database-wide by a constant-keyed advisory lock (`typegraph:identity-ddl`; a no-op on SQLite, whose writer slot already serializes the database), taken inside the per-graph schema fence and outside the per-graph identity locks, because the identity relations are shared by every graph while the fence is not. A backend that cannot publish that upgrade atomically — missing `schemaWriteTransaction` or `identityTableDdl` on the fenced path, or `executeSchemaDdl` on the commit path — is refused with the new `ConfigurationError` code `IDENTITY_UPGRADE_REQUIRES_ATOMIC_DDL` naming the missing ports, but only when a fill is actually owed; both bundled Drizzle backends implement all three when transactions are enabled. Finally, `isSeparated` no longer trusts an empty read: a graph with zero separation rows whose ledger holds a live, kind-filtered `different` assertion across two distinct classes raises `IDENTITY_STORAGE_MISSING` with the new `details.reason` `"unfilled"` rather than answering "not separated". That is the state a Store handle opened while the relation did not exist would otherwise slide into the moment another graph's upgrade created the shared relation mid-session; the remedy is to reopen the Store, which runs the fill, and the error says so. That proof is taken once per Store handle rather than once per read: it settles a property of the graph, not of the pair, and the assertion ledger has no index that answers it cheaply — proving it per read cost a 200-assertion same-only import 32% and eight concurrent `assertSame` 56% on SQLite, on the workload class whose separation relation is legitimately empty forever. A handle opened while the relation was absent still refuses, because a handle that cannot read the relation never records a proof; what a kept proof no longer re-detects is a relation truncated out of band midway through one handle's life, which `validateIdentity()` reports and the CHECK constraint still refuses at the next fusing write. Three smaller guards join that one. The recorded revision clock can no longer move backward: its upsert advances the stored row only WHERE the stored revision is strictly less than the one being written, so a late allocation cannot rewind a clock other readers have already passed, and a caller that supplies an explicit stale `previousRevision` now gets a `ConfigurationError` stating that the write would have moved the graph's revision clock backward — carrying the graph id and both revisions — instead of quietly winning. Operational Identity's SQLite writes now name the one state they cannot recover from: an identity mutation inside a transaction the CALLER began and TypeGraph adopted (`store.withTransaction(externalTx)` / `store.withRecordedTransaction(externalTx)`) can find its read snapshot invalidated before it ever takes the writer slot if that transaction was opened `BEGIN DEFERRED` and another connection committed first, and SQLite cannot upgrade a stale snapshot in place. That surfaces as a `ConfigurationError` with `details.code` `IDENTITY_TRANSACTION_NOT_WRITE_FENCED` and `details.sqliteCode` `SQLITE_BUSY_SNAPSHOT`, telling the caller to roll back and reopen with `BEGIN IMMEDIATE`; TypeGraph's own transactions already open that way, so the state is unreachable without an adopted frame. And the three PostgreSQL error shapes a concurrent `CREATE ... IF NOT EXISTS` race can take — SQLSTATE `23505`, `42701`, and `XX000` carrying "tuple concurrently updated" — are now classified by one shared predicate rather than by each call site's own partial spelling of the set, so a race one site tolerated is no longer a hard failure at the next. Edge `matchOn` composite-key construction and its per-field match comparison, embedding/fulltext field extraction, and the uniqueness path — both the unique-key computation and the `where` predicate's evaluation — now read a props bag by declared own key rather than plain property access, so a field named after an `Object.prototype` member (`toString`, `constructor`, `valueOf`) can no longer resolve to the inherited prototype member instead of the field's actual (absent) value: a unique constraint over an absent field named `toString` keyed on the inherited function, producing the empty key under `binary` collation and throwing `TypeError` under `caseInsensitive`, where it must key as absent like every other missing value. The remaining NUL-joined cache and bucket keys in edge and node operations, and legacy provenance record ids, are now built with the same injective tuple encoding already used elsewhere, closing the last collision-prone key constructions. That own-key discipline now extends past reads of a props bag. A prototype-named field the schema DECLARES is projected and returned as its stored value instead of being short-circuited as prototype noise, so selecting a field called `toString` or `valueOf` answers with what was written rather than with the inherited function; the field tracker asks the schema introspector whether the name is declared before deciding, and an UNDECLARED prototype name still resolves exactly as it always did. Schema canonicalization builds its sorted form on a null-prototype bag, so a schema carrying a `__proto__` property is no longer canonicalized — and therefore hashed and diffed — identically to one where the property is absent; the output is byte-identical for every schema without such a key, so no existing schema's hash moves. Graph merge's property bags got the same treatment, which is what lets a fork-side DELETION of a `__proto__` property record as a deletion: the deletion marker is an assignment, and on an ordinary object literal `Object.prototype`'s setter swallowed it, silently reverting the delete to the base value. A graph-extension document that declares a PROPERTY named `__proto__` is refused outright with `RESERVED_PROPERTY_NAME`, because schema validation cannot carry it — at any depth, so a NESTED object field named `__proto__` is refused on the same grounds rather than only a top-level one. `defineNode` / `defineEdge` refuse the identical declaration at definition time (`ConfigurationError`, `details.conflicts`) — at ANY nesting depth, walking nested object schemas and every wrapper (optional, nullable, default, arrays, records, unions, lazy) structurally through Zod's public `def`, with a dotted path in the error — so the two authoring paths no longer disagree about the same unstorable field: it was a typed refusal on the document path and silent data loss on the typed one. It is reachable only through a computed key — `z.object({ __proto__: … })` written literally sets the shape object's prototype instead of creating an entry, while `z.object({ ["__proto__"]: z.string() })` yields a shape whose `Object.keys` really does contain it — and it is UNSTORABLE rather than merely reserved, because Zod drops the key from every parse result and reports success even when the field is required. The same misreading has a WRITE side, and it is now closed as a class rather than case by case. `bag[key] = value` on a `{}` literal does not create an entry when `key` is `__proto__`: it invokes `Object.prototype`'s `__proto__` setter, which reparents the bag for an object value and does nothing at all for a primitive, so the value is dropped and every later own-key read agrees the writer never wrote it. Kind names (`isValidKindName` admits `__proto__` exactly as it admits `toString`), schema property names, JSON-Schema keywords, query aliases and `JSON.parse`d document keys are all data, and all of them admit it. `normalizeEdges` in `defineGraph` and every other data-keyed accumulator in the tree now build through one owner, `createDataKeyedBag`, so an EDGE kind named `__proto__` survives `defineGraph`, schema serialization, and a live store round trip instead of vanishing between the config and the registration. Because the class had already recurred twice from an incomplete enumeration, it is made self-enforcing: a ratchet test scans `src/**` for statement-position `{}` initializations and fails on any that is not allowlisted with a stated reason. Behavior note for callers: none of this is observable on returned values — every record TypeGraph hands back (serialized schema maps, aggregate rows, select contexts, migration counters, extension documents) accumulates on a null-prototype bag internally and is spread into an ordinary object at the public boundary, which preserves an own `__proto__` alias as data while restoring `Object.prototype`, so `row.toString()` and `record instanceof Object` behave exactly as before. Relatedly, the field tracker and the selective projection now track a DECLARED field named `toJSON` as the stored data it is, instead of exempting the name unconditionally; the exemption survives, unchanged, for kinds that do NOT declare it, where it exists only to keep an incidental `JSON.stringify` of the tracking context from being recorded as a field access. Two remaining leaks of that internal null prototype are closed, and one of them was a behavior that depended on which query plan ran. A smart-selected alias object and its `meta` are guarded PROXIES, and a proxy's target is caller-observable — `instanceof`, `Object.getPrototypeOf` and every other internal method resolve against it, and no `get` trap can disguise it — so `ctx.p instanceof Object` answered `false` under a selective projection and `true` under the full mapper, for the same query. The boundary spread now happens inside the guard, so both mappers hand back objects rooted at `Object.prototype` while a projected field named `__proto__` survives as an own key; the tracking context handed to the `select` callback on the field-tracking pass got the same treatment, so the probe and the engine agree. `TransactionReceipt`'s `writes.nodes` and `writes.edges` are likewise ordinary objects now, matching `writes.identity`, which always was — a `__proto__` kind still reads back as an own key with its count. And the builder handed to an index `where:` callback is no longer built on a null prototype either. `defineNode` / `defineEdge` now REFUSE a schema containing a `z.lazy()` whose getter cannot run yet, with a `ConfigurationError` naming the kind and the dotted path. This is reachable from one shape: a mutually recursive pair declared AROUND the definition call, so the second `z.object()` const is still in its temporal dead zone when the first one's getter fires. Previously that branch was skipped, and skipping it was a fail-open — a definition is validated exactly once, so a `__proto__` nested under the unreadable subtree was accepted at definition time and then silently dropped by every parse, which is the precise outcome the unstorable-name refusal exists to prevent. Recursion itself is not refused: declaring both consts before the definition — which the error message asks for — resolves every getter, and the walk then reports the real conflict at its full nested path. Note that a `z.lazy` property field is typed `unknown` by the query introspector regardless, so predicates over it degrade; recursive property schemas are not a supported shape, and this makes the one silently-wrong case loud. An import UPDATE now asserts every component its verdict read, closing the temporal half of the class rounds 6 and 7 closed for edge kind and endpoints. `UpdateNodeParams` and `UpdateEdgeParams` each gain an optional `expectedValidFrom`, with the same MUST-apply contract as `UpdateEdgeParams.kind` — a backend that receives it has to put it in the statement's own `WHERE`, and the three states are distinct: omitted asserts nothing, `null` asserts `IS NULL` (an open-left window), a string asserts equality. Both bundled Drizzle backends satisfy it through one shared NULL-safe predicate builder; a hand-written `GraphBackend` must honor it or it will silently widen a write it was told to narrow. All four import update legs — node and edge, batched and per-row — now state the bound they validated the document's window against, so a concurrent hard-delete-and-recreate between the probe and the write matches no row instead of ignoring a `validFrom` the document stated or persisting a `validTo` below the new row's `validFrom`. They state it on exactly the terms the store paths do, because it is the same verdict object: a document naming neither `validFrom` nor `validTo` made the verdict read no bound, so its properties update is fenced on identity and liveness alone and a concurrent recreate that only moved the bound no longer refuses it. A node write that consequently matches nothing is reported per row with the new message prefix `INTERCHANGE_NODE_UPDATE_TARGET_CHANGED`; the edge equivalent keeps the published `INTERCHANGE_EDGE_KIND_CONFLICT` prefix, with the validity bound added to its message text. `ImportError` still carries no `code` field, so the prefix is the branchable token. Relatedly, a node update's row write and its uniqueness transition now commit or fail as ONE unit: the new keys are claimed before the row write (the claim upsert reports the key's final owner, so it IS the conflict gate, and a transaction holding the key cannot lose it to a peer), the row write follows, and the old keys are released only once it lands — with the claims compensated away if it does not. Import is the reason this has to hold on its own terms rather than on the transaction's: `onConflict: "update"` catches a per-row `UniquenessError`, records it, and commits everything else, so an ordering where the claim can fail AFTER the row changed reported `updated: 0` for a row whose props HAD changed, whose old reservation was released, and whose new reservation belonged to another node. The fulltext and embedding syncs still run after the row write, since a write that lands on nothing must not re-derive them. The same fence now covers the STORE update paths, which read the probed row's `valid_from` for exactly the same verdict. `store.nodes.*.update` / `upsertById` and `store.edges.*.update` / `upsertById` carry `expectedValidFrom` into the statement's own `WHERE` — but only when the window verdict actually consulted the row's bound, which is when the caller stated a `validFrom` to compare against it or a lone `validTo` to invert against it. A plain `update({ props })` names no window, reads no bound, and is fenced by nothing extra; that conditionality is not an optimization but the same "only what it asserted" rule the edge identity components already followed, since predicating a write on a component the caller never claimed refuses writes that are legitimate. The decision has one owner, and it hands over the predicate rather than a flag: `assertWritableValidityWindow` now RETURNS the `expectedValidFrom` fence its verdict obliges the write to carry — empty when the verdict read no stored bound — so the answer comes from the branches the guard actually took and there is nothing left for a caller to re-derive. Interchange import had re-derived it, asserting the probed `valid_from` unconditionally and over-fencing exactly the props-only updates the store paths left alone; that second spelling is gone rather than corrected. When the assertion does catch a replaced row the update CONVERGES rather than failing: it re-reads, re-merges the caller's partial props over the current props, and re-judges the window against the current bound, so a stated window that no longer fits is refused with the same typed `ValidationError` it would have raised on the first attempt, and one that still fits is applied to the row that really exists. Convergence is bounded at one retry; a peer that keeps replacing the row ends in a `DatabaseOperationError` naming the contention instead of a livelock or a false "not found". Two adjacent defects in the same paragraph of code are fixed with it: `applyNodeResurrect` reserved its uniqueness keys before the gating `deleted_at IS NOT NULL` update and kept them when the gate refused, so a resurrection that lost its race left reservations behind for a revival it never performed (it now runs through the same claim/gate/release transition as `applyNodeUpdate`, which gives the reservations back when the gate matches nothing); and the bulk `getOrCreateByConstraint` decided whether to resurrect from the uniques row its batch probe captured, while the single-item path decided from the node row it was about to write — one decision with two owners, now read from the node row on both. Relatedly, the builder a uniqueness `where` clause names fields on now answers for every declared field rather than only the fields the props bag happens to carry — which is what its type has always promised (`-?` makes every schema field required on the builder, precisely so a partial constraint can ask whether an OPTIONAL field is present). Naming an absent field previously hit the builder object's prototype instead: an everyday partial constraint over an absent optional field threw `TypeError: Expected a defined value` for every node written without it, and a field named `toString` found `Object.prototype.toString` and threw `isNull is not a function`. Such a field now evaluates as null, which is what a partial constraint means by absent — the same builder shape schema serialization has always captured a `where` clause with. `defineGraph` also refuses a `where` clause it can already see is broken, at definition time rather than on the first write it distorts: a callback that returns something other than a predicate, or a predicate naming a field the kind's schema does not declare, throws a `ConfigurationError` naming the kind, the constraint, and — for the undeclared field — the fields the kind actually declares. A statically typed caller could express neither mistake, so this bites generated or untyped definitions, where the old behavior was a constraint that quietly matched every row or partitioned on a field that was absent forever. A third state joins those two: a constraint carrying a `where` on a kind whose schema exposes no `.shape` — not an object schema, so there is no declared-field set to check the clause against — is REFUSED rather than left unvalidated, because skipping the check silently would disable the guard for exactly the untyped callers it was written for, who are also the only callers able to put a non-`ZodObject` there. Narrowly so: a plain `unique: [{ fields }]` on such a schema needs no shape to be meaningful and still works. And the same malformed clause is refused at EVALUATION too, not only at definition — `checkWherePredicate` throws the equivalent `ConfigurationError` for a callback that returns a non-predicate, so the third reader of a `where` clause now agrees with the other two (definition-time validation and persistence-time capture, all three reading the clause through one owner) instead of quietly treating a broken constraint as one that applies to every row. That matters for constraints built outside `defineGraph`, which never passed the definition-time gate. Because the check evaluates the clause, a `where` callback now runs one extra time when the graph is defined, so it must be pure — which it already had to be, since the uniqueness path evaluates it per write. This validates node kinds whose schema exposes an object shape; edge `unique` constraints are not covered. This adds public API surface — `InvalidMergeOptionsError`, `ExportStreamCancelledError`, `EDGE_IDENTITY_MISMATCH_CODE`, `MERGE_ERROR_CODES.invalidOptions`, the `GraphBackend.identityTableDdl` port with its `IdentityTableNames` type, the `VectorSearchFrontierTuning` type, `ExportOptionsSchema`'s `signal?: AbortSignal` (accepted by both `exportGraph` and `exportGraphStream`), the optional `kind` on `UpdateEdgeParams` / `DeleteEdgeParams` / `HardDeleteEdgeParams` from the `@nicia-ai/typegraph/backend` entry point plus the four optional endpoint assertions `fromKind` / `fromId` / `toKind` / `toId` on `UpdateEdgeParams` alone (one assertion that moves together or not at all, under the same MUST-apply contract as `kind`), the `"shared-storage-in-use"` member of `ContributionRebuildRefusal`, `ValidationErrorDetails.operation` widened with `"delete"` and `"hardDelete"`, the `details.requested` / `details.heldBy` pairing on every serialized-connection interchange refusal, and the new error codes: `INTERCHANGE_EXPORT_STREAM_ABORTED` on `ExportStreamCancelledError`, the `vector.searchFrontierTuning` value of `UnsupportedBackendCapabilityError`'s `details.capability`, and the interchange, identity, and provenance `ConfigurationError` codes (`INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS`, `INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT`, `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`, `IDENTITY_UPGRADE_REQUIRES_ATOMIC_DDL`, `IDENTITY_TRANSACTION_NOT_WRITE_FENCED`, `IDENTITY_STORAGE_MISSING`'s new `details.reason` `"unfilled"`, `GRAPH_MERGE_PROVENANCE_ID_COLLISION` with its five refusal reasons, `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED`, and `CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, whose `details.constraint` carries one of `edgeCardinality` / `edgeMatchKeyConvergence` / `nodeDisjointness` / `nodeUniquenessScope` — the four members of the internal `ConstraintFenceReason` union, which is not itself exported; branch on the string values). Three refusals join the surface without a stable `details.code`, so match them by class plus `details` rather than by code. `assertApproximateMetricSupported` throws a `ConfigurationError` when `similarTo(..., { approximate: true })` is combined with a `metric` override that differs from the slot's declared metric, carrying `details` `{ nodeKind, fieldPath, requestedMetric, declaredMetric, indexType }`. `defineNode` / `defineEdge` throw a `ConfigurationError` for a schema property named `__proto__`, carrying `details.conflicts` and a `nodeType` / `edgeType` key. And a per-row import failure whose message is prefixed `INTERCHANGE_EDGE_KIND_CONFLICT` appears in `result.errors` when an interchange edge's id belongs to another kind — `ImportError` has no `code` field, so the message prefix is the branchable token, following the existing `windowErrorOf` idiom used by the validity-window import errors. Two of those are BREAKING for callers who reach past the bundled implementations. `VectorCapabilities.searchFrontierTuning` is REQUIRED, not optional: a hand-written vector strategy must now state whether its engine has a per-search ANN frontier knob — `{ tunable: true, parameter, indexType, requiresTransactionScope }` or `{ tunable: false, reason }` — rather than inheriting silence, which is the exact defect the field closes, and a strategy that omits it no longer compiles. And a hand-written `GraphBackend` must apply the new `kind` on the three edge params when it is present, and the four endpoint assertions on `UpdateEdgeParams` alongside it; a backend that accepts and ignores either turns a write the caller narrowed into an unscoped one, and nothing above it re-reads to catch that any more. Kind alone is not enough for the update: an edge's endpoints are immutable for a given row but its id is not, so a concurrent hard-delete-and-recreate under the SAME kind with DIFFERENT endpoints satisfies a kind-only predicate, and an upsert that resolved the id BY endpoints would write to an edge pointing somewhere it never looked. The endpoint fields are stated only by a write that actually checked them — a plain `update` on a kind-scoped collection resolved the edge by id and kind and states none of them, because predicating on endpoints it never checked would refuse legitimate writes. `tests/edge-write-self-verification.test.ts` asserts the contract against the bundled backends. It also changes behavior callers can observe: edge `bulkDelete`'s hooks; `importGraph` / `trustedImportGraph` / `trustedImportGraphStream` newly throwing the serialized-connection `ConfigurationError` codes (a new error type on the trusted-import surface); `persistProvenance`'s new pre-commit refusal turning a merge that previously committed-and-warned into one that refuses without touching the target; `mergeIncremental` refusing a commit whose fork point moved mid-call, where it previously committed diffs against a vanished ancestor; a stale explicit `previousRevision` now refused instead of rewinding the recorded clock; a SQLite vector or hybrid search that supplies `efSearch` now throwing `UnsupportedBackendCapabilityError` where it previously searched at the default frontier and said nothing; `defineGraph` throwing on a unique `where` clause that names an undeclared field, returns a non-predicate, or sits on a kind whose schema is not an object schema, and evaluating every such callback once at definition time; a graph-extension property named `__proto__` refused with `RESERVED_PROPERTY_NAME` at any nesting depth, and the same name in a `defineNode` / `defineEdge` schema refused with a `ConfigurationError`; a constrained write on D1, `neon-http`, or a `transactionMode: "none"` SQLite backend now refused with `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` where it previously committed unfenced; `similarTo` with `approximate: true` and a mismatched `metric` override now refused where it previously served the exact scan and dropped one of the two options silently (`store.search.vector` / `hybrid` already refused every mismatched override on their own broader rule, and the builder's EXACT path stays deliberately wider — only the silent half is closed); an interchange edge whose id belongs to a row with a different kind OR different endpoints now reported as a per-row `INTERCHANGE_EDGE_KIND_CONFLICT` error naming the mismatched components, under `onConflict: "update"` (which previously overwrote the other row's properties while silently keeping its endpoints) and `"skip"` (which previously counted it present and silently lost it) — and the import's `updateEdge` statements carry the full identity assertion in their own `WHERE`, closing the concurrent-recreate window for endpoints exactly as it was closed for kind; and `MergeIncrementalArgs.options` narrowed to `Omit` — a compile error for code that passed `options.target` to `mergeIncremental`, and a runtime `InvalidMergeOptionsError` for untyped callers that still do, where the named `target` argument previously won and `options.target` was silently ignored. So this release ships as a `minor`, not a `patch`. - [#426](https://github.com/nicia-ai/typegraph/pull/426) [`eb4b9c1`](https://github.com/nicia-ai/typegraph/commit/eb4b9c1076eb67d54cf3a7692ebe8f9a3b203453) Thanks [@pdlug](https://github.com/pdlug)! - Refuse a `validFrom` that a live row's update cannot store, instead of accepting it and writing without it. An in-place update never rewrites `valid_from`; only a resurrection does. Stating a bound that named a different instant used to block coalescing, so the upsert wrote — bumping the version and capturing a history row — while the bound itself was dropped at the SQL builder and the row's window never moved. It now raises a `ValidationError` whose issue carries the new exported code `IMMUTABLE_VALIDITY_LOWER_BOUND`, naming both the stated instant and the one the row holds so the caller can restate it without a second read. This reaches every path that accepts `validFrom` against a live row: `upsertById` and `bulkUpsertById` (nodes and edges, including a repeated id in one batch, which is judged against the row the batch just queued), `getOrCreateByEndpoints` and `bulkGetOrCreateByEndpoints` with `ifExists: "update"` — which previously dropped the option before it reached any guard — and interchange import's `onConflict: "update"` legs, where it is recorded as a per-row error prefixed with the code rather than aborting the import. What stays legal: restating the bound a row already holds (nothing to apply, so nothing is ignored); a create or a resurrection, both of which store a stated bound and are the way to give a row a different one; zero-width windows; and `getOrCreateByEndpoints` returning an existing edge, which performs no write at all. Previously-accepted writes now refuse, so this is a MINOR bump — the same precedent as the window refusals in the two releases before it. Note for temporal imports: replaying an `includeTemporal: true` export over rows that were created separately now reports those rows instead of updating their props under a lower bound it ignored. Omit `validFrom` from the update document, export with `includeTemporal: false`, or import into a fresh graph. - [#383](https://github.com/nicia-ai/typegraph/pull/383) [`bab0752`](https://github.com/nicia-ai/typegraph/commit/bab075213b480d61c91be06a2c833e510ab61a18) Thanks [@pdlug](https://github.com/pdlug)! - Merge an inherited row's end-of-validity instead of discarding it `update(id, {}, { validTo })` on a branch is an ordinary write, but the merge silently dropped it: modification detection compared properties only, so a branch that ended an inherited node's or edge's validity merged as a no-op. There was no workaround preserving row identity and history — deleting the row was the only statement the merge honored, and it is a strictly stronger one. An end-of-validity is now treated as a **sibling of deletion**: - one branch ends a row → that end is committed, including a _later_ end that extends the window; - several branches end it differently → no conflict; the **earliest** end wins (a fixed, commutative rule, so the merge stays order-independent); - `mergeIncremental()`'s target already ended it → the target's end stands, the same committed-target precedence identity survivors already get; - one branch ends it and another deletes it → deleted, with **no** `DeleteModifyConflict` — the stronger statement absorbs the weaker one; - a branch re-states the end the target holds → nothing is staged at all: no write, no version bump, no history row. `MergeReport` gains `validityEnds`, listing every row whose end the merge changed and the branches that claimed it — the arbitration is silent by design, so this is how a caller sees it happened. Window deltas the commit cannot apply to a live row (a fork `validFrom` divergence, or a `validTo` cleared back to open — both reachable only by soft-delete + resurrect inside a fork) are now reported in `dropped` with reason `"window-not-applicable"` instead of being ignored. **Behavior change.** Merges where a branch ended an inherited row now write that end, so new version bumps, history rows, and recorded-time entries appear where a no-op used to be. There is no opt-out flag: a permanent knob for "does the merge lose data" is worse than this note. Nothing that previously succeeded now fails. Also hardens `coalesceUnchangedUpserts`: the requested and stored valid-time bounds are compared as instants rather than as raw text, so the decision cannot come to depend on a dialect's timestamp rendering. A bound that is not a representable instant still counts as a change, so it reaches the write path that rejects it rather than being coalesced away. - [#426](https://github.com/nicia-ai/typegraph/pull/426) [`eb4b9c1`](https://github.com/nicia-ai/typegraph/commit/eb4b9c1076eb67d54cf3a7692ebe8f9a3b203453) Thanks [@pdlug](https://github.com/pdlug)! - Report a non-canonical validity bound whether or not `coalesceUnchangedUpserts` is on. A parseable-but-non-canonical bound equal to the stored instant — `"2100-06-01T00:00:00Z"` against a stored `"2100-06-01T00:00:00.000Z"` — compared as "unchanged" with coalescing on, so the write was skipped and the `ValidationError` the same call raises with coalescing off was swallowed. An unrelated performance flag decided whether malformed input was reported. A non-canonical REQUESTED bound now counts as a window change, so it reaches the write path and raises identically either way. Re-stating a window in canonical form still coalesces, including against a driver that renders the stored value as an equivalent zoned string: only the stored side needs canonicalizing, because the requested side is held to canonical form by the write validation this no longer hides. - [#394](https://github.com/nicia-ai/typegraph/pull/394) [`6dd43c4`](https://github.com/nicia-ai/typegraph/commit/6dd43c40bbb988fc8d112f177f03caab86c18359) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: scope the edge repoint/dedupe fold to collisions repointing caused A TypeGraph store is a multigraph: nothing enforces uniqueness on `(from, kind, to)`, `create()` makes a parallel edge, and `getOrCreateByEndpoints()` is the opt-in set-semantics accessor. The merge's edge fold nevertheless grouped **every** staged edge by `(from, kind, to)` and collapsed each group onto its lowest-sorting edge id, so a branch that created a parallel edge lost one of the two rows — and _which_ one it lost depended on how the branch-created id happened to sort against the existing one. The fold is now restricted to what it was designed for. It groups staged edges by the endpoint pair they named **before** repointing, and collapses one row per pair — so a collision the canonicalization itself induced (`x → a` and `x → b` both becoming `x → c*`) still folds to a single edge, keeping the existing min-id-survivor, property-reconciliation, and end-of-validity behavior. Edges that already shared their endpoints are no longer folded together: each distinct edge id commits as its own parallel row, and a valid-time end lands on the row whose author claimed it rather than migrating to an unrelated survivor. A group that mixes the two folds only **across** the pairs. A repointed `x → b` joining two parallel `x → a` rows merges into one of them, and the other row still commits with the edit its author made — repointing said nothing about the rows that were already there. Previously the whole group collapsed, which dropped that edit silently: a folded-away row is never rewritten. What makes two staged edges "the same row" is their **edge id**, not equal properties. One inherited edge staged by several branches still folds into a single write with its property disagreements reconciled; a branch-created edge is a new parallel row even when its properties coincide with an existing one's. Merges that previously collapsed parallel edges will now commit both, and the spurious `PropertyConflict` those collapses reported between two rows that were never the same row is gone. - [#406](https://github.com/nicia-ai/typegraph/pull/406) [`248f56a`](https://github.com/nicia-ai/typegraph/commit/248f56ab2c8bc849fe9385157f02344c7ceb612d) Thanks [@pdlug](https://github.com/pdlug)! - Refuse valid-time windows of negative width. A write whose `validTo` precedes the row's effective `validFrom` describes a row that stopped being true before it started — observable at no `asOf` coordinate, and unrepairable by any later write — and it used to be accepted silently on every path except a node update. It now raises a `ValidationError` whose issue carries the new exported code `INVERTED_VALIDITY_WINDOW`. This is a behavior change: writes that previously succeeded now fail. Two shapes refuse where they did not before. - A stated `validFrom` / `validTo` PAIR must be ordered, on node and edge `create`, `upsertById`, `bulkUpsertById`, `getOrCreateByEndpoints` and its bulk form, and on an imported document. `getOrCreateByEndpoints` judges the pair before its existence probe, so whether a call is valid no longer depends on whether the edge happens to exist yet. - An UPDATE's lone `validTo` must not precede the lower bound the row carries. Nodes already enforced this; edges did not, which is how a graph merge could hand a committed edge an end predating its start and still report success. This covers a resurrecting write too: an edge RETAINS its `valid_from` across resurrection, so reviving one into a window that closed before it began now means restating the start — pass `validFrom` alongside `validTo`. Landing a revived edge in the ENDED state is otherwise unchanged. `getOrCreateByEndpoints` and its bulk form now honor `validFrom` on the `"resurrected"` branch, where they previously accepted it and silently dropped it — which is what left the refusal above with no way to satisfy it. As the backend has always documented for a resurrecting write, naming `validFrom` asserts the COMPLETE window, so an accompanying `validTo` is applied and an omitted one REOPENS the revived row rather than leaving the tombstoned incarnation's end in place. A `"found"` or `"updated"` live edge is unaffected: its stored lower bound is history and still stays put. Interchange import records the refusal as a per-row error prefixed with `INVERTED_VALIDITY_WINDOW`, so one bad row does not abort the import; its `onConflict: "update"` legs are held to the existing row's `valid_from` exactly as a direct `update` is. Trusted import refuses the whole stream with reason `invalid_stream`. Two shapes stay legal, deliberately. A ZERO-width window (`validTo === validFrom`) is what a same-instant retraction produces at millisecond precision, so the store's own output still round-trips. An INSERT carrying a lone historical `validTo` still means "born already ended": the write instant stamped as `valid_from` is a storage convention rather than a caller assertion, and such a row is read back through `includeEnded`. - [#380](https://github.com/nicia-ai/typegraph/pull/380) [`2c5dd29`](https://github.com/nicia-ai/typegraph/commit/2c5dd29d9c23024119589d5360b265a0c3ab49da) Thanks [@pdlug](https://github.com/pdlug)! - Store the cause of an identity assertion's ending instead of deriving it. The identity assertion relation and its recorded mirror gain nullable `ended_by_kind` / `ended_by_id` columns. A node soft-delete cascade stamps the deleted node's `(kind, id)` onto every assertion it ends, in the same statement that closes the row; `NULL` means the row was retracted explicitly. Graph merge's `RetractionCause` now reads that column instead of comparing an assertion's `valid_to` against a deleted endpoint's `deleted_at`. This removes the derivation's same-millisecond residue: a retraction issued in the same millisecond as the delete that followed it is now classified as `explicit` and survives a merge whose deletion is overruled, where the timestamp comparison could only read the tie as a cascade and drop the branch's intent. The hard-delete residue remains by design — a hard delete removes the assertion rows outright, so no evidence survives to read. Archival interchange carries the cause as an optional `endedBy` on each assertion, so an export/import round-trip preserves why an assertion ended. Import rejects an `endedBy` on an open assertion (`IDENTITY_IMPORT_ENDED_BY_WITHOUT_END`) or one naming a node that is not an endpoint of the assertion (`IDENTITY_IMPORT_ENDED_BY_NOT_ENDPOINT`); a CHECK constraint on the relation backs both rules at the database. Operational Identity has not shipped in a release, so the relation changes shape with no migration path. - [#268](https://github.com/nicia-ai/typegraph/pull/268) [`9721ba2`](https://github.com/nicia-ai/typegraph/commit/9721ba2854f6e5504018ee8c3b7a0eaaf87314bb) Thanks [@pdlug](https://github.com/pdlug)! - Add the opt-in TypeGraph Identity Profile with typed store, transaction, and temporal-view APIs; configurable same-ID folding or assertion-only identity; kind-branded and hydrated member reads; idempotent assertion receipts; ended assertion retraction results; assertion history; interchange and graph merge propagation; identity-expanded traversal; cross-backend closure storage; and fail-fast capability errors for non-transactional D1 and neon-http drivers. Harden ontology construction and reload validation: propagate disjointness through interleaved subclass and equivalence closure, validate inverse endpoint compatibility and partner uniqueness, reject unresolved extension edge names in `inverseOf` and `implies` while retaining absolute external IRIs, recompute serialized closures, and deprecate the type-level `sameAs` and `differentFrom` factories in favor of Operational Identity. **Behavior changes.** Ontology and registry validation is now stricter and runs both at graph construction and when a persisted schema is loaded, so a few patterns earlier versions silently accepted now throw a `ConfigurationError`: duplicate ontology relations, hierarchical self-loops, disjointness contradictions (a kind disjoint with itself, with a subclass ancestor, a common subclass of two disjoint parents, or a kind declared both `equivalentTo` and `disjointWith`), multiple distinct `inverseOf` partners for one edge, inverse endpoint incompatibility, and unresolved extension edge names in `inverseOf` or `implies`. To recover, fix the graph definition; for a persisted extension document, correct the stored document before upgrading (or rewrite it through the previous minor, which still accepts it). Interchange documents remain readable across versions — `1.0` documents are still accepted on import, and exports write `2.0`. Trusted import rejects identity-enabled target stores (`identity_unsupported`) and identity-bearing input (`invalid_stream`) rather than silently dropping assertions or leaving the derived closure empty; use `importGraphStream` for an export that carries identity truth. Bundled SQLite and PostgreSQL backends provision the three identity relations, including effective custom `SqlSchema` names, before first-enable preflight. An already-enabled graph with missing identity storage instead fails with `IDENTITY_STORAGE_MISSING`; restore missing ledgers from backup, or recreate a missing derived closure and rebuild it before serving traffic. `create()`/`upsertById()` of a soft-deleted same-`(kind, id)` row now resurrects that row on every graph (properties replaced, validity window reset so `validFrom` becomes the resurrection instant) rather than leaking a storage constraint error. These are additive-strictness and semantics-pinning changes on top of the new opt-in profile, hence the minor bump. **Type-level breaking notes for backend and tooling authors.** 1. `ResolvedSqlTableNames` gained three required fields (`identityAssertions`, `recordedIdentityAssertions`, `identityClosure`). Out-of-tree `GraphBackend` implementations must supply them; the `SqlTableNames` input type keeps these optional, so only the resolved type is total. 2. `SqlSchema` (the abstract class) gained three abstract members (`identityAssertionsTable`, `identityClosureTable`, `recordedIdentityAssertionsTable`). External subclasses must add them; the `createSqlSchema` factory path is unaffected. 3. `FORMAT_VERSION`'s literal type changed from `"1.0"` to `"2.0"`. Comparisons like `FORMAT_VERSION === "1.0"` are now type errors; both versions remain accepted on import. **Behavioral note.** `revisionNow()` now returns `Promise` (a branded string, assignable to `string`; use `asRecordedInstant` to round-trip). **Review-hardening pass (same release).** - `store.identity` and the read-only view `identity` surfaces now use the same conditional presence as `tx.identity`: the property does not exist on identity-disabled graph types, so misuse is a compile error. The `IdentityFacadeFor` / `IdentityReadFacadeFor` helper aliases and the duplicate `IdentityNodeRef` type are gone (use `IdentityFacade`, `IdentityReadFacade`, and `GraphNodeReference`); the loose input type formerly named `GraphNodeRef` is now `IdentityNodeRefInput`. - `StoreView` and `RecordedStoreView` are now type aliases over an implementation class plus `ViewIdentityAccess`, exported alongside a construction-compatible `const`. `new StoreView(...)` and `instanceof StoreView` keep working; subclassing them does not. - `MergeReport.merged` gained an `identity: { asserted, retracted }` section (`MergedCounts`), and `DroppedItem` is now a discriminated union (`kind: "node" | "edge" | "identity"`) so dropped identity assertions are enumerable in the report. - Identity merge conflicts — including transitive `same`/`different` contradictions, retract/reassert races, and assertions over merge-deleted nodes — are detected at plan time and surface as `IdentityMergeConflictError` (`GRAPH_MERGE_IDENTITY_CONFLICT`) through `merge()`'s returned `Result`. Convergent edits (the re-asserting branch itself also retracted the pair) merge cleanly. - `ImportError.entityType` widened to `"node" | "edge" | "identity"`; identity import failures are recorded in `result.errors` instead of throwing. Archival identity imports now bound validity windows (`validTo` must not be in the future for ended rows, `validFrom` must not be for open rows) with `IDENTITY_IMPORT_FUTURE_VALID_TO` / `IDENTITY_IMPORT_FUTURE_VALID_FROM`. - Changing `identity.sameIdAcrossKinds` is now classified a breaking schema change requiring explicit migration; explicit `migrateSchema()` rebuilds the identity closure atomically with the schema commit, and an unapplied identity-only breaking change surfaces `IDENTITY_PROFILE_MIGRATION_PENDING` rather than a generic `MigrationError`. **Performance.** Current-coordinate identity reads (`membersOf`, `areSame`, `areDifferent`, `representativeOf`, `nodesOf`) were O(total graph size) on SQLite — the class-members lookup defeated the closure's class index and the planner scanned every live node per read. The rewritten statement is O(class size): ~40x faster on a populated graph (0.013 ms vs 0.56 ms per read at ~6,000 nodes), with a smaller improvement on PostgreSQL. **Follow-up hardening (same release).** `ValidationIssue` gained an optional `assertionId` field carrying the offending identity assertion structurally; identity import failures (self-assertions included) attribute their `result.errors` entries by that id, never by message parsing. The identity enablement preflight is derived inside `initializeSchema()` itself, so every public first-commit path — bare `ensureSchema`/`initializeSchema` included — builds and validates the closure atomically with version 1. Identity reads on `includeTombstones` views hydrate soft-deleted rows the coordinate makes visible instead of silently dropping them. **Import error attribution.** The import coordinator tags rethrown errors with the id of the assertion it was applying, so `ImportResult.errors` attribution for contradictions and missing endpoints identifies the failing assertion rather than the first assertion sharing its endpoints. **The identity preflight is not substitutable.** `initializeSchema()` and `SchemaManagerOptions` no longer accept a schema-commit preflight callback — a no-op callback could suppress the mandatory closure build at version 1. Both (and `MigrateSchemaOptions`) instead accept the effective `SqlSchema` (`schema`), and every identity-enabled schema commit derives the closure preflight internally from it. **One schema source in the batteries-included constructors.** The nested `schemaManagement` option no longer accepts `schema` (typed out and stripped at runtime): the effective `SqlSchema` has exactly one source, `store.schema`, which also drives physical table provisioning — a second schema could name tables that were never created. The manager brand-validates the `schema` option with `requireSqlSchema()` before any DDL or version commit, so a schema-shaped plain object is rejected (`INVALID_SQL_SCHEMA`) instead of committing a closure into tables the Store never reads. **Historical bridges must exist; plan-time simulation knows the profile.** Archival identity imports now require every ended assertion's endpoints to exist structurally (soft-deleted rows qualify; the store's own exports already satisfy this), so a hand-built document can no longer conduct historical identity through a node that never existed. The graph-merge plan-time contradiction check now simulates the target's identity semantics — implicit same-id folds under `sameIdAcrossKinds: "fold"` and ontology `disjointWith` between class member kinds — so those contradictions surface as `GRAPH_MERGE_IDENTITY_CONFLICT` at plan time instead of a generic commit failure. Counterfeit schema objects are rejected before any identity DDL runs, on fresh and already-enabled graphs alike. **Assertion-free nodes join the plan-time simulation.** The merge planner's contradiction check now seeds its universe with every post-merge canonical node and the live target peers sharing their ids (one kind-free indexed probe, only under `sameIdAcrossKinds: "fold"`), so a node no assertion names — newly created, retyped, or an existing same-id peer — can no longer fold into a disjoint-kind class undetected and fail at commit as a generic merge error. **Universe seeding, precisely.** The plan-time simulation seeds retyped canonical nodes under the kind the commit writes (not their pre-retype kind), the live same-id peer probe reads the merge TARGET when it differs from the diff source (`mergeAgainstBase`, `mergeIncremental`), and the incremental commit revalidates the probed peer set inside its transaction — a same-id peer landing in the plan→commit window is refused as the same typed replan error the other window guards raise. **The window guard ranges over the committed plan.** The incremental fold-peer revalidation compares only ids the final plan folds on — commit-ready canonical nodes and remapped assertion endpoints — so a window row at an id canonicalization dropped is tolerated as an ordinary target advance instead of raising a spurious replan error. **The window guard is class-transitive.** The incremental fold-peer guard also snapshots each final seed's structural identity class at plan time and revalidates the fingerprints inside the commit transaction — a window row or assertion that joins a seed's class through another member (leaving the seed's direct same-id peers untouched) is refused as the typed replan error, and a rerun surfaces the contradiction as a plan-time `GRAPH_MERGE_IDENTITY_CONFLICT`. **A validated baseline, exactly.** The incremental identity guard now re-probes and snapshots the final seeds' classes AFTER planning and re-runs the identity simulation against that exact snapshot — its members join the simulation universe unlinked, with connectivity rebuilt from the deletion-filtered fresh ledger and fold unions — so drift landing between planning and the snapshot fails as a typed plan-time conflict instead of becoming the guard's baseline. Fingerprints are structurally encoded (injective for ids containing any character) and carry a liveness bit, so a planned assertion endpoint deleted in the commit window is refused as the typed replan error rather than failing generically. **Negative truth in the baseline.** The post-plan identity recheck consumes the target's FRESH assertion ledger (not the pre-planning staging capture), and the transaction guard carries a deterministic fingerprint of the `different` assertions touching the guarded universe — a `different` committed in either window is refused typed instead of surfacing as a generic commit failure. **The identity guard covers both profiles.** The incremental identity baseline, class/liveness fingerprints, and negative-ledger guard run for every identity-enabled merge — under `sameIdAcrossKinds: "ignore"` too, where explicit assertions still change plan legality. Only the same-id fold expansion stays profile-gated; the plan-time simulation additionally models the profile-independent create-time constraint that one id cannot be shared by ontology-disjoint kinds, and the direct-peer window check refuses a disjoint same-id arrival under `"ignore"` while tolerating a benign one. **Replacement is legal.** Planned node deletions are excluded from both sides of the incremental identity guard (peers, liveness, class members, and the ledger slice), and `applyMergePlan` soft-deletes nodes BEFORE the node writes — so a plan replacing a node with a disjoint same-id one (the order the create-time constraint permits, and the order the same operations run directly on a store) commits instead of being falsely rejected or failing at apply. **Deleting a bridge splits the class.** The incremental recheck derives connectivity from the deletion-filtered fresh ledger and the checker's fold unions — never by pre-linking the old closure's filtered member lists — so a plan that deletes an identity bridge and asserts its former ends `different` commits instead of being falsely rejected. Snapshot class members still join the simulation universe (unlinked) so fold links at unprobed ids keep participating. **The transaction re-derives legality.** The incremental commit guard's final step re-runs the full identity simulation on transaction reads — fresh deletion-filtered ledger, snapshot members, fold unions — so drift that leaves every fingerprint unchanged (a redundant `same(a, b)` that becomes the surviving link once the plan removes the pair's bridge) is refused as the typed replan error instead of failing generically at apply. **One assertion id, one truth — validated where it can be typed.** The planner refuses one id staged for two different complete truths and any staged id already identifying different truth among the target's stored rows (ended included, exactly the set the import coordinator compares); the commit transaction revalidates every planned id against transaction reads (both commit modes), so a window row reusing a planned id — even with endpoints entirely outside the guarded universe — refuses as the typed replan error instead of a generic id-conflict at apply. **Retractions carry their complete truth.** A merge plan's identity retractions are full expected rows, never bare ids: the planner validates each one against the row its id identifies on the target and SKIPS — reported as `identity:retraction-target-mismatch` in `dropped` — a retraction whose id the target reuses for different truth, instead of ending a row the branch never saw. The commit transaction revalidates the surviving retractions (and every planned assertion id) by id in BOTH commit modes; snapshot commits need this explicitly because the legacy base@V token fingerprints only CURRENT assertions, so an ended window row claiming a planned id would otherwise slip through to a generic apply failure. The raw staged assertions are also checked one-id-one-truth BEFORE the semantic survivor dedupe, closing the validity-only collision (same id, same pair, different `validFrom`) that dedupe used to collapse silently while the report listed the id as both applied and dropped. **The applier is the completeness backstop, typed.** Any identity refusal that still escapes the commit — an invariant the plan-time simulation does not (yet) mirror — is translated into the typed `IdentityMergeConflictError` with the applier's error as its cause, instead of surfacing as the generic merge wrapper. Identity-typed environment errors (missing profile, non-atomic backend) pass through unchanged. A property-based law suite additionally quantifies the merge contract over randomized identity histories on both backends: refusals are always typed, a committed ledger is internally consistent, pre-merge truth survives unless a branch retracted it or deleted an endpoint, and the report never lists an id as both dropped-as-duplicate and newly current. **Truth replacement is visible to the diff.** The identity diff compares ids present on both sides by COMPLETE truth, not presence: a branch that hard-deletes an assertion's endpoint (physically removing the row), recreates it, and imports the same id for different truth used to diff as empty — the merge silently kept the base truth the branch had replaced. The replacement now stages as a retraction plus a new assertion, and because the applier never reuses an ended row's id, the merge refuses typed instead of silently preserving either side. **Identity semantics extracted; translation at the applier boundary.** The plan-time identity derivation, contradiction simulation, and commit guards now live in `graph-merge/merge-identity.ts` with a one-directional dependency from the merge orchestrator (functions take a structural `IdentityPlanSlice`, never the full plan type). The typed-conflict translation wraps exactly the identity-apply call inside the commit, so it also classifies refusals whose identity code lives in nested validation issues (`details.issues[].code`) and — because only identity rows are applied at that boundary — a missing-node error there can only mean a vanished assertion endpoint, which now translates too instead of surfacing as the generic wrapper. Exact-duplicate staging (two branches importing one identical row) no longer reports the id as dropped while applying it. **Five laws, three lanes.** The property suite now also holds every successful merge to BRANCH-EFFECT accounting — every truth a branch holds is applied with equal complete truth, enumerated as dropped, retracted, or invalidated by an endpoint deletion; silent loss is a law violation — and runs the whole law set in three lanes: snapshot `merge()` under both identity profiles (with hard-delete/recreate and same-id fold peers in the operation alphabet) and `mergeIncremental()` against a target that ADVANCED after the fork, where branch truth meets independently-moved target truth. Truth-preservation and branch-effect exclusions are truth-aware: a retraction excuses a row's death only when the retracted COMPLETE truth matches, and a hard-delete/recreate excuses exactly the rows it physically killed, not everything ever touching the node. A dropped-as-duplicate id must never be current post-merge. The generator skips only expected semantic refusals (contradiction, missing node); any other error fails the run rather than silently emptying the histories. Independent-target merge semantics are now documented in the identity guide. **The survivor pick respects committed truth.** The law suite caught its first live defect within a day: a branch-minted assertion id could win the semantic-pair dedupe against the target's own committed row — the applier (idempotent per pair) then skipped the write, so the report claimed an id as applied that never landed while listing the target's committed row as dropped. Ids already committed on the target with the exact staged truth now always win the survivor pick, pinned by a deterministic incremental test alongside the law. **The simulation uses the plan's REAL canonical map.** Both closure re-runs (post-plan and in-transaction) previously reconstructed the member→survivor map from the report-shaped resolutions, which drops pure ontology-retype clusters and mis-keys mixed-kind members — degrading the decisive in-transaction backstop into judging endpoints at pre-merge identities (a false negative) and enabling an unresolvable replan loop (a false refusal). The plan now carries the exact `canonicalOf` map the commit repoints edges with, and the reconstruction is deleted. The simulated base ledger is also deletion-filtered inside the checker itself, so all three call sites share one post-deletion rule. **An overruled deletion no longer ends identity truth.** A node soft-delete cascades — it ends every open assertion touching the node — so the deleting branch's diff stages those endings as retractions indistinguishable from intent. When the delete/modify resolution keeps the modification (the default `"flag"` and `"modifyWins"` policies), the node survives, and the cascaded retraction is now dropped with it — reported as `identity:deletion-overruled` — instead of ending the resurrected node's assertions anyway. **Identity-only merges advance the revision clock.** The interchange import records capture touches through its own recorded binding, so a merge whose only effect was creating assertions never marked the mutation as written: the durable revision clock stayed unmoved and every base@V token went stale, letting a later commit's target-unchanged guard pass against a target that DID move. The apply now marks the write from the import summary, with a regression test on a revision-tracking store. **Guard structure hardened.** The by-id freshness check is invoked directly by BOTH commit paths (never through the peer-probe guard's early return), the environment-code passthrough covers the identity environment/corruption codes that must never be translated into replan advice, and one id staged as both a new assertion and a retraction — an applier-refusing shape currently unreachable through any supported staging path — refuses typed defensively at plan time. **External-review hardening (cross-model pass).** An independent review with a different model produced six verified fixes: (1) the deletion-overruled retraction filter is provenance-aware — a retraction is dropped only when EVERY contributing branch is explained by an overruled endpoint deletion, so a branch that retracted independently keeps its effect (the earlier filter silently suppressed it); (2) committed-row precedence in the survivor dedupe is RE-DERIVED after endpoint canonicalization, closing the collision the first fix missed when reconciliation collapses a branch pair onto a committed target pair; (3) merges refuse, typed, any branch whose store ran a schema operation after forking (its committed schema hash no longer matches the fork source's) — schema side effects can no longer be smuggled into a data merge as bare identity changes; (4) a kind-dropping `migrateSchema()` now cascades the assertion ledger exactly as `Store.removeKinds()` does, instead of stranding current assertions on unregistered kinds where a later "no-op" merge would end them; (5) a staged survivor's valid-time window travels with the commit write, so a branch-authored — possibly already ended — window survives resurrection instead of being reset to merge time and silently joining a live fold class; (6) `merged.identity` reports rows the applier actually created and ended (idempotent skips excluded) instead of planned intents, and the replan-vs-conflict error suggestions are path-specific. Temporal windows on MODIFIED inherited nodes remain outside merge state — a documented boundary. **Second cross-model pass: the fixes' own compositions.** A follow-up external review of the previous round's fixes produced seven more verified corrections. The schema-drift guard now anchors on the branch's AT-FORK `(version, hash)` row — a round-trip migration that restores the document hash still advances the monotonic version and is refused, and unmanaged fork sources are no longer falsely rejected; revision-anchored `base@V` tokens bake in the active schema version, fencing the same round-trip on the target side (the legacy content fingerprint already covered it). Kind-dropping schema operations cascade the assertion ledger even when identity is DISABLED at drop time (the ledger, not the schema profile, is the signal), and first enablement purges assertions naming unregistered kinds, so historical orphans cannot be adopted into a fresh closure. Node resurrection carries `validFrom` through the internal update path (a branch-authored ended window no longer inverts into merge-time-start), merged edges carry their staged windows exactly as nodes do, and when the live incremental target itself contributed the surviving member, the TARGET's committed window wins over a branch re-window. Canonicalization that would move a COMMITTED assertion's own endpoints refuses with a specific typed conflict (committed rows cannot be rewritten), and window-identical upserts coalesce again instead of rewriting version and history state. **Final pre-merge pass.** A last scoped external review of the previous hardening commit returned three refinements, all applied: the disabled-identity cascade's outside-transaction emptiness probe is skipped when THIS commit is the one disabling identity (writers on the still-enabled prior schema could otherwise slip an assertion in between probe and lock — the locked cascade always runs for that shape); node writes validate the EFFECTIVE validity lower bound, so a lone historical `validTo` on a resurrecting upsert refuses typed instead of persisting a born-inverted, permanently invisible window (edge resurrection keeps its sanctioned resurrect-as-ended contract — edges retain their stored lower bound, so the node-side corruption cannot arise there); and bulk edge coalescing compares explicit windows against the stored window, so no-op incremental merges stop rewriting byte-identical target edges. The property law lanes carry explicit five-minute test budgets sized for coverage-instrumented CI shards. - [#375](https://github.com/nicia-ai/typegraph/pull/375) [`fc6075d`](https://github.com/nicia-ai/typegraph/commit/fc6075ddca02de4c0fa9d15167c124868ab947b5) Thanks [@pdlug](https://github.com/pdlug)! - **The merge commit proves its own identity result.** After a merge commit's identity DML, and inside the same transaction, the applier now re-derives the identity classes the merge TOUCHED and refuses a contradiction there — a class whose member kinds the ontology declares disjoint, or a current `different` assertion whose endpoints share a class. Both commit modes run it, seeded from the planned assertion and retraction endpoints plus (under `sameIdAcrossKinds: "fold"`) the node identities the commit writes, so the cost is proportional to the affected classes rather than the graph. This makes a committed identity ledger correct independently of the plan-time simulation, which reasons about state read before any write. The simulation and the commit-window fingerprints remain as the diagnosability layer: they refuse early, before anything is written, naming exactly what drifted. Because the scans resolve classes through the materialized closure — the same authority every current identity read uses — a closure that lags its ledger can hide a contradiction as easily as invent one. On any inconsistency the closure is rebuilt from the base relations inside the commit transaction and the scans re-run against it: a clean second pass means the closure was stale and is now repaired atomically with the merge, while a repeated contradiction aborts the whole merge. There is no partial commit either way. `IdentityContradictionErrorDetails.operation` gained a `"merge"` member for this refusal, which reaches callers as the existing `IdentityMergeConflictError` (`GRAPH_MERGE_IDENTITY_CONFLICT`) with the contradiction as its cause. ### Patch Changes - [#418](https://github.com/nicia-ai/typegraph/pull/418) [`5795127`](https://github.com/nicia-ai/typegraph/commit/57951275ee6310bcbfebadb1e3bdc46204769052) Thanks [@pdlug](https://github.com/pdlug)! - store: coalesce a bulk upsert that re-states a row's own validity window `bulkUpsertById` now decides whether a requested `validFrom` / `validTo` is a change the same way `upsertById` does, through one shared comparison, so a batch and the same items applied one at a time write the same rows. Two defects met in that comparison. The node bulk path refused to coalesce whenever an item named a bound AT ALL, so any caller that re-stated a row together with the window it already holds — a merge commit, or any read-modify-write loop that round-trips `meta.validFrom` — bumped the row's version and wrote a history and revision entry for a row that did not change. The edge bulk path did compare, but compared the bounds as DRIVER TEXT: a Postgres driver that renders `timestamptz` as a zoned string rather than a `Date` yields text that is equivalent to the caller's canonical ISO bound without being identical to it, so the same batch could coalesce on one backend and write on another. Both paths now compare INSTANTS, and an unrepresentable bound still counts as a change so the write path raises the `ValidationError` the caller is owed rather than coalescing it away. The bulk paths also track the window each queued write leaves behind, so a repeated id in one batch is compared against the batch's own pending state rather than the once-prefetched row. Previously an edge item that re-stated the window the row held BEFORE the batch was read as unchanged and skipped, dropping a write the sequential path performs. A later copy that re-states the window a queued write established coalesces; one that names a bound the backend was left to stamp (an omitted `validFrom` on a create) writes, since that instant is not knowable batch-locally. - [#363](https://github.com/nicia-ai/typegraph/pull/363) [`cdc904b`](https://github.com/nicia-ai/typegraph/commit/cdc904b574d4aa5a4ef10f3378b1ca4079373209) Thanks [@pdlug](https://github.com/pdlug)! - Chunk iterative graph-algorithm node-kind initialization within each backend's bind-parameter budget. - [#396](https://github.com/nicia-ai/typegraph/pull/396) [`994c7da`](https://github.com/nicia-ai/typegraph/commit/994c7da9aff06d995680381e3b43c4a05bece5a3) Thanks [@pdlug](https://github.com/pdlug)! - Reach the candidate edge of a current-coordinate identity-expanded traversal by an equi-join instead of a correlated membership scan An identity-expanded hop at the current coordinate read the materialized closure from inside a correlated `EXISTS`, so nothing in the join condition linked the frontier row to the edge row. Both engines were free to enumerate _frontier rows × edges of the matching kind_ and probe the closure per pair, which cost quadratically in graph size. Each traversal step now widens its frontier onto the closure's class members with an outer join, so the candidate edge is reached by the same ordinary indexed equality a traversal without identity expansion uses. One compiler path serves both coordinates and both emitters. On SQLite a hop over 100,000 matching edges from a 500-row frontier drops from 51.6 s to 77 ms, and `EXPLAIN QUERY PLAN` seeks `typegraph_edges_from_idx` where it used to scan every matching edge per source row; PostgreSQL drops from 9.2 s to 61 ms. A traversal at a **historical** coordinate reaches its candidate edge through the same step, so it gains the same join order: on SQLite an `asOf` hop over 100,000 matching edges drops from 31.3 s to 241 ms. That coordinate's own remaining cost is the ledger reconstruction, still tracked in typegraph#310. Results are unchanged at every coordinate: physical edges stay deduplicated, and member visibility, the `sameIdAcrossKinds` profile and the read instant are all resolved exactly where they were. The class members a current-coordinate step joins are reached by seeking the closure from the frontier row, so the cost of the widening tracks the frontier and its classes rather than the identity population — see the follow-up changeset, which replaced the graph-wide relation this change first shipped with that seek. - [#400](https://github.com/nicia-ai/typegraph/pull/400) [`05af68d`](https://github.com/nicia-ai/typegraph/commit/05af68d122371ddeab03269536f2c7484ccf1a74) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: record each provenance contribution once, so the persisted count is the rows actually written Several planning phases legitimately observe the same `(role, canonical, branch, source)` contribution. An inherited edge is credited once when its modification survives delete/modify and again when the repoint fold reads it as a source, and a fold set's `mergedIds` carries one entry per staged copy — so a row staged by several branches re-offered each of its branches once per copy. The tuple is exactly the sidecar row's identity, so those re-observations were never new information: they inflated `provenancePersisted.count`, and because a single `bulkUpsertById` batch cannot create the same id twice, the over-count was the milder half: with `persistProvenance: true`, a merge in which a single branch modified one inherited edge failed the whole best-effort persist, so `provenancePersisted` came back absent, a `provenance persistence failed …` warning was reported, and NO provenance rows were written at all. Contributions are now collapsed at the single recording funnel, so the record list, the in-memory `provenance.byBranch` index and the reported count all speak about distinct contributions. `persistProvenanceRecords` additionally collapses records that hash to one id before the batch, which makes its documented "row count written" true for any caller's record list. Every genuinely distinct contributing branch is still credited. - [#384](https://github.com/nicia-ai/typegraph/pull/384) [`cc7af7b`](https://github.com/nicia-ai/typegraph/commit/cc7af7b1ef53458296f27b73484cd799667df13c) Thanks [@pdlug](https://github.com/pdlug)! - Report the real active schema version in `StaleVersionError.details.actual` when a PostgreSQL schema-managed write loses to a concurrent schema commit. The write fence takes a `FOR SHARE` lock on the active schema row. At `read committed`, a locking read that blocks behind an in-flight schema commit rechecks only the row versions its own statement snapshot saw — so once the winner marked the old row inactive, the fence saw no active row at all and reported `actual: 0`, misrepresenting the database as having no active schema. The fence now settles an empty locked read with a non-locking read, which observes the committed winner, and reports `0` only when a graph genuinely has no active version. The write itself was always correctly rejected; only the error metadata was wrong. - [#427](https://github.com/nicia-ai/typegraph/pull/427) [`facef56`](https://github.com/nicia-ai/typegraph/commit/facef560d14c38607d6414818c65e43dc65a88d2) Thanks [@pdlug](https://github.com/pdlug)! - Bound current-coordinate identity expansion by the frontier instead of the identity population An identity-expanded traversal at the current coordinate built its class relation by self-joining the whole identity closure into a materialized CTE, before any frontier predicate applied. The relation's size is the sum of the squares of every class in the graph, so a hop from a single start row paid for identity classes it never touched: nine unrelated classes of 501 members materialize 2,259,009 seed/member pairs, and the hop measured 564 ms on SQLite and 568 ms on PostgreSQL where the equivalent traversal without expansion costs microseconds. Doubling an unrelated class quadrupled the cost. Each step now seeks the closure from its own frontier rows — the frontier row's class through the closure primary key, that class's members through the class index, each member's node for its visibility — so the peer relation is never built for classes the query does not touch. The same hop measures 0.5 ms on SQLite and 2.4 ms on PostgreSQL, and PostgreSQL's `EXPLAIN (ANALYZE)` reports 18 rows visited against 4,522,557. A single-start-row hop over 50,000 folded triples drops from 387 ms to 1 ms on SQLite. Wide-frontier hops are unchanged: 500 source rows over 100,000 matching edges measures 325 ms against 331 ms, because that shape was already paying for a population it used. The **historical** coordinate keeps its hoisted, materialized relation. Its rows come from a recursive fixed point over the assertion ledger that no frontier row narrows, so evaluating it once per statement is still the win, and the two coordinates are now deliberately different strategies behind one interface rather than one relation with two sources. Both remain a single compilation path across dialects. Results are unchanged at both coordinates: physical edges stay deduplicated, member visibility is still resolved against the read instant, and a frontier row in no class still expands to itself. - [#391](https://github.com/nicia-ai/typegraph/pull/391) [`8eafebd`](https://github.com/nicia-ai/typegraph/commit/8eafebdae8d28fc48343fcd5d498f47b1684311f) Thanks [@pdlug](https://github.com/pdlug)! - Evaluate the historical identity-class reconstruction once per query instead of once per candidate edge An identity-expanded traversal hop under a historical coordinate (`asOf`, `asOfRecorded`, or a non-current `view()`) has no materialized closure to read, so it rebuilds classes from the assertion ledger. That rebuild used to sit inside the correlated edge predicate, where SQLite re-materialized it for every candidate _(source row, edge)_ pair, and under `sameIdAcrossKinds: "fold"` each rebuild also scanned the structural same-id relation across the graph — a quadratic term that quadrupled per doubling of graph size. The reconstruction is now a single materialized query-level relation of `(seed_kind, seed_id, kind, id)` rows, seeded by the nodes that have identity peers rather than by the frontier, so it depends on nothing a traversal step carries and is built once for the whole statement. Each step widens its frontier onto that relation with an outer join, which turns the candidate-edge lookup into the same ordinary indexed equality a traversal without identity expansion uses. On the narrow-edge fixture (SQLite, all _n_ nodes acting as source rows) the hop drops from 122/486/1984/8261 ms at _n_ = 250/500/1000/2000 to 7/7/14/28 ms, and grows linearly rather than quadratically. Results are unchanged at every coordinate. Current-coordinate traversal still reads the materialized closure through its existing correlated predicate. - [#425](https://github.com/nicia-ai/typegraph/pull/425) [`92354bf`](https://github.com/nicia-ai/typegraph/commit/92354bfb559c255a03b4b5e91741b87b98de5777) Thanks [@pdlug](https://github.com/pdlug)! - Ask props bags whether they carry a key with `Object.hasOwn` rather than `in`. A props bag is data: its keys come from a JSON column, so a schema may declare a field named after an `Object.prototype` member — `toString`, `constructor`, `valueOf` — and such a field is ordinary data that survives Zod validation and the JSON round-trip untouched. `in` cannot answer "does this row carry this property" for such a bag, because `"toString" in {}` is `true`: a row that does not carry the key reads as though it does, and the read that follows yields the inherited prototype member instead of stored data. This is a lost-write fix, not only hardening. In a graph merge, a fork's bag is its full intended state, so a base property absent from it was deleted by that fork. Under `in` that deletion was never detected for a prototype-named field: no deletion tombstone was written and the base value survived the merge, silently discarding the fork's write. The same misclassification credited a branch that does not carry such a property with the inherited prototype member as if it were a stored value, letting an invented claim compete in conflict resolution and be reported to the caller as that branch's value. A schema diff also reported a removed prototype-named property as an incompatible schema change rather than a removal, because the absent field resolved to a function that was then compared as though it were the field's new schema. The edge fold and the node cluster union were affected in the same way, and their worst outcome was a committed function. For a property the SURVIVING row does not carry, both ask that row's bag for the value to keep, so under `in` they took the inherited `Object.prototype` member and wrote that function into the merged row instead of the value a member actually carried. The cluster union additionally routed such a property through the separate base-property-conflict policy on the strength of a base member that does not carry it, so the wrong policy decided the committed value. The edge fold's claim filter separately counted a member that says NOTHING about such a field as having AUTHORED it; the shared value collector discarded that phantom claim, so the two agreed only by one absorbing the other's mistake. Two guards were quietly weakened rather than corrupted. Graph-extension validation accepted a unique constraint on an undeclared field named after a prototype member — it answered "declared" against the prototype — and went on to index a field that does not exist. The evolve guard that refuses re-adding a kind whose data cleanup is still pending never counted such a kind as added, so it skipped the refusal. The convention now has one owner, `hasOwnKey`, applied across graph-merge node and edge property resolution, schema-diff property classification, schema-removal reconciliation, interchange unknown-property stripping, graph-extension document validation, query and index schema-field validation, the evolve pending-removal guard, edge `matchOn` composite-key and match-comparison reads, and embedding/fulltext field extraction. `in` remains correct, and still in use, when both the key and membership question are internal: a discriminated union's tag, a capability probe, a brand check, and the deliberate `Object.prototype` lookup in selective projection. A user-supplied field name is always checked as an own key, even when the schema shape itself is statically known, so names such as `__proto__` and `constructor` cannot masquerade as declared fields through `Object.prototype`. A plain `bag[field]` walks the prototype chain exactly as `field in bag` does, so the same misreading reached two more read paths that never used `in` at all. An edge's `matchOn` composite key and its per-field match comparison (`getOrCreateByEndpoints`, edge upsert dedup) read a caller's stored and input props by a schema-declared field name; a field named after a prototype member that neither bag carries as an own key now reads as `undefined` on both sides instead of the same inherited function, so a match or non-match decision is never made on a phantom shared value. `syncEmbeddings` and `computeFulltextContent` read a declared embedding or searchable field the same way, so an undeclared row no longer surfaces a prototype function as if it were the field's stored value there either. `then` and `toJSON` complete the same class from the other end. They are the two names JavaScript itself probes — the thenable check and the `JSON.stringify` hook — so every proxy standing in for a row resolved them to `undefined` up front to stay safe to await and to serialize. They are also legal schema field names, and answering them by NAME before consulting the data made the read side lie: a declared field called `toJSON` came back `undefined` through smart selection while the full mapper returned the stored string, so the same query answered differently depending on whether the optimizer engaged — exactly the equivalence selective projection exists to preserve. The predicate builders (query, traversal, collection, and index WHERE) made such a field unaddressable outright, and field tracking dropped a declared `then`, so the projection could not have carried it. **A declared `then` or `toJSON` field is now tracked, projected, readable through smart selection, and usable in a predicate.** The rule is the one the surrounding fixes already follow: ask the data question first — `hasOwnKey` for a materialized row, `hasDeclaredField` for a proxy whose key set is the schema — and fall back to the probe exemption only once the answer is "not data", which is what keeps `await` and `JSON.stringify` working on a partially projected row. Returning an own `then` is safe as well as correct: props decode from a JSON column, so the value can never be callable, and the thenable check ignores a non-callable `then` exactly as it does on the plain objects the full path returns. `isInteropProbeKey` owns which names those are, and an ESLint rule bans the bare name comparison that used to stand in for the decision. `__proto__`, the case originally reported, is the NARROW variant. Every VALIDATED write path blocks it: Zod drops an own `__proto__` key, and `bag["__proto__"] = value` assigns a prototype rather than creating a key, so an assignment-built bag cannot carry one either. It is still reachable through `trustedImportGraph`, which by contract does not validate properties and writes a caller's bag verbatim — the stored JSON parses back with `__proto__` as an own key on both dialects. Recorded here so the two are not confused: a prototype-named field needs nothing unusual at all, while `__proto__` needs the trusted path. - [#381](https://github.com/nicia-ai/typegraph/pull/381) [`e6fb356`](https://github.com/nicia-ai/typegraph/commit/e6fb35669ed0bfeab1cfce64aafc72afaf5698a2) Thanks [@pdlug](https://github.com/pdlug)! - identity: answer current different-ness with one probe on the separation relation `identity.areDifferent()` and the `assertSame` contradiction precheck resolved both identity classes and then loaded every current `different` assertion touching one of them, scanning in JS for one that spanned the pair. That scan grew with class size and, past the backend's bind budget, took more than one statement. Both now probe the derived separation relation on its primary key `(graph_id, class_key_low, class_key_high)` instead: `areDifferent` reads the assertion ledger not at all, and the precheck reads it only to name the conflicting assertion in the typed error it is already about to throw. Results and typed errors are unchanged. Reads at a valid-time `asOf` or a recorded coordinate still reconstruct from the ledger, since the separation relation projects current assertions onto current classes. A probe never answers "not separated" when it could not read: a missing relation refuses with `IDENTITY_STORAGE_MISSING`, and any other driver failure propagates unchanged so transient conflicts stay classifiable. - [#404](https://github.com/nicia-ai/typegraph/pull/404) [`479ca78`](https://github.com/nicia-ai/typegraph/commit/479ca783781d9449a7b20422446c68fd702f516b) Thanks [@pdlug](https://github.com/pdlug)! - Fix `bulkUpsertById` throwing on a repeated id whose row does not exist yet. `bulkUpsertById` applies items in order, so a repeated id in one batch is last-write-wins — but that only held for an id that already existed. The create branch queued its create without registering the id in the batch-local pending map, so a second copy of a **new** id queued a second create and the batch failed with `Node already exists` / `Edge already exists` (a unique-constraint violation on some paths). Callers feeding a batch straight from a stream or a changeset, where a key can legitimately appear twice, hit this on first delivery of a key. A queued create is now registered like a queued update: a later copy of the id takes the update path over the queued create, which runs after the batch's creates, so the final row is exactly what the equivalent sequence of `upsertById` calls produces — the later copy's props merged over the created row, one version bump per real write, and the created row's validity lower bound. With `coalesceUnchangedUpserts` enabled, a value-identical second copy of a new id now coalesces against the queued create instead of writing a second time. Nodes and edges are both fixed; for edges, as for an id that already existed, a later copy's `from` / `to` are ignored because an update never repoints an edge. Two smaller consequences of routing every queued write through the same state: a repeated id whose dirty check rejected an earlier item's props no longer reports the wrong error, and no later copy can coalesce against a stale prefetched row after an earlier item queued a write. - [#362](https://github.com/nicia-ai/typegraph/pull/362) [`9982960`](https://github.com/nicia-ai/typegraph/commit/9982960e66343c6980b8d6e87e0cb159a981e72a) Thanks [@pdlug](https://github.com/pdlug)! - Translate PostgreSQL read-only and missing-`TEMP` failures during graph analytics into `UnsupportedBackendCapabilityError`, preserving the driver error as the cause. Both refusal points are covered: a standby that rejects the read-write working-table transaction, and a role that cannot create the temporary table inside it. - [#420](https://github.com/nicia-ai/typegraph/pull/420) [`d82fdaf`](https://github.com/nicia-ai/typegraph/commit/d82fdaf369109fe791be00ffe330351fb8aa4d00) Thanks [@pdlug](https://github.com/pdlug)! - Harden two failures at the operations/backend boundary: a create the engine refuses now reports the condition it actually hit, and the last UPDATE path that could store an inverted valid-time window no longer can. **A create refused by the engine reports "already exists", not a raw driver error.** A create learns an id is taken either from its own existence probe or from the engine refusing the INSERT, and the second used to escape as a `DrizzleQueryError` whose `.message` is the raw INSERT text. One condition therefore surfaced as a typed user error down one path and an opaque system error down the other, and callers could not branch on it at all. The engine's report is now classified structurally and both routes raise the same `ValidationError`, on the single and batch create paths for nodes and edges alike. Two things reach the engine's path. A NODE create probes first, but the probe and the INSERT are two statements and PostgreSQL's default READ COMMITTED does not serialize the two write transactions, so a concurrent create of the same new id can commit in between — the issue's reproduction. An EDGE create has no existence probe at all, so the engine's refusal is its only report of a taken id, on every backend and with no race involved. Classification is structural, never message text: SQLSTATE 23505 plus the PostgreSQL protocol's own constraint and relation fields, and SQLite's extended result code, which distinguishes a primary-key duplicate (1555) from any other unique-index duplicate (2067) in the code itself. Every such refusal, from either route, now carries the new exported issue code `ENTITY_ALREADY_EXISTS`, so a caller can recognize it without matching on the message. `details.entityType` and `details.kind` say what was refused; `details.id` names the taken id, and is absent only when the refused statement inserted more than one row, because the engine reports that the statement collided without saying which row did. No race is needed to reach that: a bulk create of edges, whose ids the caller supplied and which nothing probes, is refused this way on every backend. The classification is scoped to the primary key on purpose. A `unique: true` index declaration materializes a UNIQUE INDEX on the same relation, and violating that is a declared-uniqueness failure about the row's VALUES rather than a duplicate identity — PostgreSQL reports it under the index's own name and SQLite under a different extended code, so it never matches and is unaffected. Neither is a declared `unique` constraint conflict, which still raises `UniquenessError`. SQLite never reached the node race: `BEGIN IMMEDIATE` gives the writer slot to one transaction at a time, so a second create cannot sit between its probe and its INSERT while the first commits. Its probe is authoritative there, and the refusal was already the typed error — it now carries the code too. A duplicate EDGE id on SQLite did surface as a raw `SqliteError`, and now raises the same error as it does on PostgreSQL. **A node resurrection stores the bound its window guard measured against.** A resurrection rewrites `valid_from` rather than retaining it, so the guard that refuses inverted windows has no stored bound to check and used the write instant instead — sampled in the operations layer, while the backend went on to stamp its own, strictly later, sample. A `validTo` at the guard's instant passed as zero-width and committed as NEGATIVE width a millisecond later, the exact shape the previous release exists to refuse. The operations layer now passes the instant it validated against explicitly, so the bound that is checked is the bound that is stored. Stating both endpoints is unaffected; the only change to a successful write is that a resurrection's `valid_from` is the operations layer's instant rather than the backend's — sampled a moment earlier inside the same locked write, before the uniqueness entries it re-checks and re-inserts. Edge resurrection was never exposed: an edge RETAINS its stored `valid_from` unless the write names a new one, so its guard measures against a value already on disk and predicts nothing. - [#403](https://github.com/nicia-ai/typegraph/pull/403) [`c0279fc`](https://github.com/nicia-ai/typegraph/commit/c0279fc9cafac91b7a6ec31c343bd8fbdd12d743) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: credit the branch that authored a merged row's end-of-validity Ending a row's validity is authored state, but a branch whose only change to a row was its window could contribute the instant the merge committed and still be absent from the merge's provenance. An identity is staged once, so a branch that merely moved an inherited edge's window had its staged copy skipped whenever another branch's property edit already staged that edge — and the provenance for edges is derived from the staged copies. Nodes were worse: a window change had no provenance path at all, so a window-only node ending was credited to nobody even when no other branch touched the row. The credit now comes from the window resolution itself, which is the only phase that knows whose claim was committed. It credits exactly the branches whose claim IS the resolved end — a claim that lost the least-claim rule contributed nothing to the committed row, and remains visible in `MergeReport.validityEnds` under `claimedBy`. A branch that both edited a row's properties and moved its window stays one contribution. The staged copy that carries a window-only ending is no longer credited for carrying it: that copy exists only to give the ending a row to write, its properties are the base's, and the branch holding it is whichever sorted first — possibly one whose claim the merge discarded. Which branch carries the row is left exactly as it was, because that branch also labels the base's properties in the repoint fold's property union, where a relabelled contribution can change which value a fold commits. Merge outcomes are unchanged; only the provenance is. ## 0.45.0 ### Minor Changes - [#355](https://github.com/nicia-ai/typegraph/pull/355) [`2882b23`](https://github.com/nicia-ai/typegraph/commit/2882b23ef6daed20041ccea56e3cfa76a8435c7a) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.nodes..updateWhere()` for typed, transactional set-based node updates selected by property and independent relationship predicates. The operation validates complete after-images and atomically maintains uniqueness, fulltext, vector, history, and revision state on SQLite and PostgreSQL. Its cross-backend storage primitive returns every updated after-image and provides bind-budgeted, graph- and concrete-kind-scoped uniqueness cleanup so rebuilding reservations cannot clear same-id nodes of another kind. - [#352](https://github.com/nicia-ai/typegraph/pull/352) [`872d196`](https://github.com/nicia-ai/typegraph/commit/872d1960b470352cdcf5412a1303e59b776ce1ed) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.repairContributions()`, a privileged, idempotent repair pass for strategy-owned contribution storage. It re-audits declarations from the active persisted graph, non-destructively retries `missing-marker` and `failed-materialization` findings, reports `stale` and `orphaned-marker` as `requires-rebuild`, and returns a fresh post-repair diagnostic result. Repair targets remain backend-owned so callers do not need access to TypeGraph-managed tables, physical names, or DDL. - [#353](https://github.com/nicia-ai/typegraph/pull/353) [`c225605`](https://github.com/nicia-ai/typegraph/commit/c22560576e2e22f5808ead72e552ead4b7f8743c) Thanks [@pdlug](https://github.com/pdlug)! - Add Store-level heterogeneous bulk edge reads that keep database round trips independent of schema breadth. ### Patch Changes - [#351](https://github.com/nicia-ai/typegraph/pull/351) [`ff8e428`](https://github.com/nicia-ai/typegraph/commit/ff8e4280456b21985033977ec9be553bad06d63c) Thanks [@pdlug](https://github.com/pdlug)! - Ensure the kind-removal status table before `evolve()` checks it, so databases created before TypeGraph 0.44 can evolve without manual backend initialization. Concurrent PostgreSQL focused-table ensures also retry the catalog uniqueness race that `CREATE TABLE IF NOT EXISTS` can surface during replica startup. - [#350](https://github.com/nicia-ai/typegraph/pull/350) [`f752543`](https://github.com/nicia-ai/typegraph/commit/f752543749ac6914592e5f765f6e153b69e72518) Thanks [@pdlug](https://github.com/pdlug)! - Clarify that schema-managed Stores are immutable schema snapshots. After `evolve()` changes the schema, callers must use the returned Store or the updated `StoreRef.current` for subsequent work; a previously captured Store is not mutated and its managed writes are rejected by the schema-version fence. Document how long-lived caches detect schema commits from other processes with `getCommittedSchemaVersion()` and refresh through a verified Store open. Correct the `StoreRef` contract to say that the replacement is installed before a successful schema-changing call resolves, rather than claiming that the in-memory ref update is atomic with the persisted schema commit. ## 0.44.0 ### Minor Changes - [#331](https://github.com/nicia-ai/typegraph/pull/331) [`a1f1fde`](https://github.com/nicia-ai/typegraph/commit/a1f1fdec1be98dcbf243db738fb43cf634cae278) Thanks [@pdlug](https://github.com/pdlug)! - Add batched multi-source edge reads: `bulkFindFrom` / `bulkFindTo` `EdgeCollection` could only read the edges of ONE endpoint at a time, so rendering a page of N nodes with their relationships cost N statements. The new `store.edges..bulkFindFrom(froms, options?)` and `bulkFindTo(tos, options?)` read a whole SET of endpoints in set-oriented statements per endpoint kind and bind-budget chunk, returning the edges grouped per input (index `i` holds the edges of input `i`, empty array when an endpoint has none). This widens the predicate rather than batching the calls: `from_id = ?` becomes `from_id IN (...)`, the same prefix seek on the edge relation's system index. Temporal semantics are identical to `findFrom` / `findTo` — same default mode, same `temporalMode` / `asOf` options, same soft-delete filtering, same per-endpoint ordering — and a `StoreView` exposes both methods pinned to its coordinate. Pass `limitPerInput` to bound each endpoint's fan-out (applied in SQL via `ROW_NUMBER()` where the backend supports window functions). Inputs larger than the backend's bound-parameter budget are split across statements transparently. Backend authors: this adds a new **optional** `GraphBackend` operation, `findEdgesByEndpointSet(params)`, with its own `FindEdgesByEndpointSetParams`. `FindEdgesByKindParams` is unchanged. It is a separate operation rather than optional fields on `findEdgesByKind` so that a backend which does not implement it cannot degrade silently. Optional params would have left an existing backend type-correct while it ignored the id list and returned every edge of the kind — which the collection would rebucket into a correct-looking answer at unbounded cost. Support is now detected by the method's presence, before any read is issued, and `bulkFindFrom` / `bulkFindTo` refuse with a typed `ConfigurationError` on a backend without it rather than looping `findFrom` per input. The parameter shape also makes the previously-validated illegal states unrepresentable: one `side` instead of two id lists, no scalar `fromId` / `toId` to disagree with a set, and no `limit` / `offset` / `after` to slice a read the backend splits into bind-budget chunks. - [#334](https://github.com/nicia-ai/typegraph/pull/334) [`7a2e16b`](https://github.com/nicia-ai/typegraph/commit/7a2e16b70685609b19cda6683dc1d545d5aa5f9a) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.verifyContributions()`, an owner-agnostic diagnostic that crosses each contribution currently expected by the active graph and backend strategies against its durable marker and the physical catalog. Nothing on the open path probes the catalog — boot and the runtime asserts short-circuit on a per-instance signature cache and then on the marker row alone — so a database whose strategy-owned tables were dropped out of band opened completely clean and failed at the first fulltext or vector read. The diagnostic reports detected problems as `orphaned-marker` (marker records a success, table absent), `missing-marker` (table present, nothing attests it), `failed-materialization` (the marker records a failed attempt and no table was produced — marker and catalog agree, and it is broken anyway), or `stale` (marker recorded at a different shape), with the `owner` / `logicalName` / `physicalName` and, for vector slots, the `kind` and `fieldPath` needed to route to the state-specific repair without reconstructing internal marker strings. For vector slots, `missing-marker` and `failed-materialization` use the non-destructive forced ensure; only `orphaned-marker` and `stale` rebuild vector storage with `store.reembedVectorField`. `lastError` carries the reason the marker recorded, when it recorded one: `state` says which repair to run, `lastError` says why it broke. A contribution with neither a marker nor a table was never attempted and is omitted, as are retired markers and unsupported vector slots, so an empty result is not proof of initialization. It is read-only (one existence query per contribution table, no DDL, no writes) and deliberately not a boot step; the fast-path caching stays the default. Backends that cannot probe their own catalog throw `ConfigurationError` rather than reporting a clean bill of health. - [#335](https://github.com/nicia-ai/typegraph/pull/335) [`7950bb0`](https://github.com/nicia-ai/typegraph/commit/7950bb0de0b2cb93fd71c4cf6644df0b65126b6b) Thanks [@pdlug](https://github.com/pdlug)! - Support list-valued parameters in `in()` / `notIn()` `field.in(param("ids"))` now binds the whole list at `.prepare().execute({ ids: [...] })`, so the canonical "fetch these ids" query can finally be prepared. The list rides on a single bound parameter that the dialect unpacks (`json_each` on SQLite, `jsonb_array_elements_text` on PostgreSQL), which keeps arity out of the SQL text: one compiled statement serves every list length, and a list of any size costs one bound parameter instead of one per element. An empty list is valid — `in([])` matches nothing, `notIn([])` matches everything. A `ParameterRef` passed among the _elements_ of a literal list (`in(["a", param("b")])`) was previously coerced to a literal and silently produced wrong results. It now throws `UnsupportedPredicateError` naming the supported form. A name used both as a list and as a scalar in one query is rejected at `prepare()`. List elements are validated against the field's type before binding, so `[1, "a"]` against a number field is rejected with a `ConfigurationError` rather than failing on PostgreSQL and silently matching nothing on SQLite. This matches the literal form, which already refuses a mixed list. Non-finite numbers (`NaN`, `±Infinity`) are now rejected in any parameter binding, list or scalar. `JSON.stringify` turns them into `null`, so a list binding became SQL NULL — `notIn(param("x"))` with `[NaN]` filtered out every row — and SQLite binds a scalar `NaN` as NULL, so `eq(param("x"))` with `NaN` quietly matched nothing. Both now throw. `DialectAdapter` gains two members, `inListParameter` and `packListValue`; custom dialect adapters must implement them. - [#329](https://github.com/nicia-ai/typegraph/pull/329) [`03e87bd`](https://github.com/nicia-ai/typegraph/commit/03e87bd1d370408af7dcb640729e9579508db29f) Thanks [@pdlug](https://github.com/pdlug)! - Export the committed-schema reads from the package root. `getActiveSchema`, `isSchemaInitialized`, and the `SerializedSchema` type now sit next to `getCommittedSchemaVersion` in `@nicia-ai/typegraph`, so answering "what kinds does this database already have?" no longer requires finding the `@nicia-ai/typegraph/schema` subpath or querying `typegraph_schema_versions` by hand. `getActiveSchema` and `getCommittedSchemaVersion` now cross-reference each other in their docstrings. - [#332](https://github.com/nicia-ai/typegraph/pull/332) [`2fb8925`](https://github.com/nicia-ai/typegraph/commit/2fb89258b2047fe8ae6d2ce97cb7f219200285e5) Thanks [@pdlug](https://github.com/pdlug)! - Fix `migrateSchema()` silently dropping runtime-committed kinds `migrateSchema(backend, graph, currentVersion)` committed `graph` verbatim. It did not fold the persisted graph extension, so kinds committed at runtime by `Store.evolve()` — which live in `schema_doc.extension`, not in the compile-time graph — were erased from the active schema document while their rows stayed in `typegraph_nodes` / `typegraph_edges`, reachable by nothing. The persisted `deprecatedKinds` set was erased the same way. This was reachable by following the library's own advice: the `MigrationError` raised for a breaking change told callers to "use `getSchemaChanges()` to review, then `migrateSchema()` to apply", and doing so with the graph they passed to `createStoreWithSchema` destroyed every `evolve()`-committed kind. Two changes: - **The persisted graph extension (and deprecated-kind set) is now folded in**, exactly as `createStoreWithSchema` and `getSchemaChanges` already did. `migrateSchema` was the last commit path that did not. Callers pass the graph they have; runtime-committed kinds survive. - **A commit that would drop a kind still holding rows is refused** with a `MigrationError` whose `details.reason` is the new `"kind-removal"` discriminant and whose `details.droppedKinds` names them. Pass `{ discardDroppedKindRows: true }` if losing those rows is the intent — the name says what the flag does, because the next reconcile deletes them. The guard fires on the actual harm — rows the next reconcile would delete — not on kind removal as such. Dropping an _empty_ kind is unaffected, so the documented three-deploy removal flow (stop writing → delete the rows → drop from `defineGraph()` and migrate) still works exactly as written; Deploy 2 is now what makes Deploy 3 legal instead of being merely advisory. Live rows only, matching the `excludeDeleted` default of the equivalent probe in `Store.evolve()`. Breaking property changes — the documented reason to reach for `migrateSchema()` — are unaffected. `MaterializeRemovalsEntry` gains a `"skipped"` variant, carrying `reason: "kind-is-live"`. `materializeRemovals()` returns it when a queued removal names a kind the active schema declares again, so the decline is reported rather than leaving the queue at a non-zero depth with nothing explaining why. Consumers that switch exhaustively on `status` must handle it. The type is now a discriminated union, so `"failed"` carries a required `error` and `"skipped"` a required `reason`. `Store.evolve()` refuses to re-add a kind whose data cleanup is still pending, with a `ConfigurationError` naming the kind and pointing at `materializeRemovals()`. Reads filter only by `(graph_id, kind)`, so re-adding before cleanup made the previous incarnation's rows visible alongside the new ones — and the cleanup was then declined because the kind was live, so they were never reclaimed. The documented cycle (remove → `materializeRemovals` → re-add) is unaffected. Two further corrections found while reviewing the above: - **A stale store can no longer resurrect a removed kind.** The fold now strips the supplied graph's own extension slice before applying the persisted one, so the committed document is a function of the database alone. Previously `migrateSchema(backend, store.graph, v)` — `store.graph` is public and returns the merged graph — unioned a stale slice back in and silently undid `Store.removeKinds()`, leaving a kind the schema called live while its `typegraph_kind_removals` row stayed queued for a later hard-delete. `Store.#catchUpToStored` has stripped for this exact reason; the schema layer now matches it. - **`discardDroppedKindRows`'s documentation was wrong.** It claimed the dropped kind's rows stay and that `materializeRemovals` "will never clean them up". `materializeRemovals` re-derives removals by walking schema-version history, so the next reconcile hard-deletes them regardless. The flag buys a committed schema, not retained data; the docstring now says so and points callers at copying the rows out first. ### Patch Changes - [#328](https://github.com/nicia-ai/typegraph/pull/328) [`dc2a386`](https://github.com/nicia-ai/typegraph/commit/dc2a386fa2b6275b7d0f3d6d80e2959ea094365b) Thanks [@pdlug](https://github.com/pdlug)! - State `store.batch()`'s real cost where callers see it. `batch()` runs its queries in sequence, keeping at most one in flight — at least one statement each, and two for a query whose selective-field mapping falls back after its statement has already executed. So it caps concurrency at best and will not fix an N+1. It is also not a snapshot: PostgreSQL's default read-committed isolation lets a later query in the batch observe a commit the earlier ones did not, and there is no public way to get one across fluent queries, since a transaction context exposes no query builder. The docstrings for `batch()`, `BatchableQuery`, `executeOn`, and the edge `batchFind*` methods now lead with that, and point at the set-oriented and chunked alternatives, described by what they actually do: `.traverse()` compiles a chain to one statement, `store.subgraph()` costs 2 statements on SQLite and 3 on PostgreSQL, `getByIds()` issues one statement per bind-limit chunk (falling back to concurrent per-id lookups where the backend exposes no batch read), and `bulkFindByIndex()` costs a probe plus that same chunked hydration. The docs site is corrected to match, including claims that `batch()` "minimizes round-trips for reads", that `batchFind*` collapses N reads into "a single transactional round-trip", that `subgraph()` is a single statement, and that `getByIds()` is a single query. Transaction support no longer implies a transport shape anywhere: Durable Objects use an ambient transaction with no framing statements, and the non-transactional path may still reuse one client. The changelog entry that shipped `batch()` carries a correction note rather than a silent rewrite. Execution semantics are unchanged. One public diagnostic changes: the `ConfigurationError` message for a batch endpoint read on a read-only `StoreView` no longer calls `batch()` a "batch loader". - [#344](https://github.com/nicia-ai/typegraph/pull/344) [`ea05d0d`](https://github.com/nicia-ai/typegraph/commit/ea05d0da9f4956062b895138be03ad3b36fa289b) Thanks [@pdlug](https://github.com/pdlug)! - Fence deferred kind cleanup against concurrent schema re-adds. Removal now rechecks the active schema and atomically deletes live rows, recorded-time intervals, vector storage, and contribution markers under the schema lock. Custom backends that implement the optional `schemaWriteTransaction` capability must expose transaction-bound statement execution, table-existence probing, schema DDL, and vector-contribution marker deletion on its callback target. - [#347](https://github.com/nicia-ai/typegraph/pull/347) [`1616e93`](https://github.com/nicia-ai/typegraph/commit/1616e9380de834afa4912c91b79e45ed8edcd122) Thanks [@pdlug](https://github.com/pdlug)! - Fence schema-version commits against concurrent schema-managed Store writes. SQLite uses its immediate writer transaction; PostgreSQL locks the active schema row in shared mode for managed writes and exclusive mode for schema commits. Managed writes revalidate their Store schema version while holding the fence, so stale queued writes fail instead of landing against a schema that no longer accepts them. Snapshot-isolated PostgreSQL transactions may raise the database's native serialization failure; callers retry the whole transaction, and graph merge does so automatically. Schema-managed Stores on non-transactional or custom backends without the fence now fail closed on writes. Raw `createStore()` instances, direct backend writes, and Stores whose schema metadata was reset by `clear()` remain outside the versioned guarantee. - [#342](https://github.com/nicia-ai/typegraph/pull/342) [`d481054`](https://github.com/nicia-ai/typegraph/commit/d481054ebbac7fd6ec06b8d0b5cfd28313efab89) Thanks [@pdlug](https://github.com/pdlug)! - Make the documented store query hooks fire for query-builder statements, including prepared queries, batched queries, and selective-projection retries. Each submitted statement now reports its SQL, parameters, row count, duration, and failures through the existing `StoreHooks` callbacks. - [#343](https://github.com/nicia-ai/typegraph/pull/343) [`347d5e3`](https://github.com/nicia-ai/typegraph/commit/347d5e3ba1c634697b873044a9770436a119604b) Thanks [@pdlug](https://github.com/pdlug)! - Avoid repeated selective-projection fallback queries. Smart-select planning now covers common high-value threshold branches, and prepared queries remember a missing-field fallback so later executions fetch the full row directly. ## 0.43.0 ### Minor Changes - [#320](https://github.com/nicia-ai/typegraph/pull/320) [`010132a`](https://github.com/nicia-ai/typegraph/commit/010132a6ce2625b83f6256ef78bbc9bbd78867ee) Thanks [@pdlug](https://github.com/pdlug)! - Allow idempotent endpoint-based edge writes to set application-time validity. `getOrCreateByEndpoints` now accepts `validFrom` and `validTo`, while `bulkGetOrCreateByEndpoints` accepts them per item. Creation applies both fields, updates and resurrections apply `validTo`, and pure found results leave the existing window unchanged. ## 0.42.1 ### Patch Changes - [#317](https://github.com/nicia-ai/typegraph/pull/317) [`8024711`](https://github.com/nicia-ai/typegraph/commit/80247111bde9282dbbe1a9ef3c31ca66bd16ae39) Thanks [@pdlug](https://github.com/pdlug)! - Prevent large PGlite bulk writes from silently leaving the connection unable to return rows. PGlite backends now advertise their safe 32,767-parameter limit, PostgreSQL batch sizes follow the active backend capability, and over-budget statements fail before driver dispatch. ## 0.42.0 ### Minor Changes - [#313](https://github.com/nicia-ai/typegraph/pull/313) [`a797a8b`](https://github.com/nicia-ai/typegraph/commit/a797a8b1ebe077869e42a334816b172314cb0132) Thanks [@pdlug](https://github.com/pdlug)! - Expose the schema-commit surface's decisions as data instead of prose, so callers can pre-flight a proposal and classify a failure without matching message text. `MigrationError` now carries a stable `details.reason` discriminant — `"schema-behind" | "breaking-change" | "no-active-version" | "version-not-found"` (exported as `MIGRATION_FAILURE_REASONS`) — plus the structured `details.diff` for the outcomes that computed one. Branch on `details.diff.hasBreakingChanges` to tell an additive change from an incompatible one, with no re-query and no substring matching. Note that `MigrationErrorDetails.reason` is now required rather than an optional free-text string. For pre-flight, `classifySchemaChanges(diff)` reduces a diff to `"identical" | "additive" | "incompatible"`, and the existing SELECT-only `getSchemaChanges` is now reachable from a store handle: `store.schemaChanges()` returns the diff and `store.requiresMigration()` answers the boolean predicate (also `true` when nothing has been committed yet). A least-privilege runtime can detect that it needs the privileged bootstrap instead of discovering the migration wall partway through a request. Documents two operational facts that were previously invisible at the call site: kinds are scoped to the `graph_id` (a namespace _is_ a graph id — separate declaration sites do not isolate kinds), and running many `graph_id`s with divergent schemas in one database is a supported multi-tenant pattern, including the one cross-graph coupling (SQL index names are database-global, so identical kind+index shapes share a physical index and divergent shapes fail loudly). Also fixes `getSchemaChanges` to fold in the persisted graph-extension before diffing, matching what the commit path already does. Without it a compile-time graph was compared against a stored schema that also contains runtime-committed kinds, so those kinds read as removals and an unchanged schema was reported as requiring a breaking migration. ### Patch Changes - [#313](https://github.com/nicia-ai/typegraph/pull/313) [`a797a8b`](https://github.com/nicia-ai/typegraph/commit/a797a8b1ebe077869e42a334816b172314cb0132) Thanks [@pdlug](https://github.com/pdlug)! - Stop reporting a reordered declaration as a schema change. Restating a kind with its properties, enum members, or edge endpoints listed in a different order is a semantic no-op, but the diff compared those arrays positionally and reported the kind as `modified` — forcing callers into a privileged migration for a schema that had not actually changed. A reordered `enum` was even classified `breaking`, i.e. a pure reordering demanded a destructive-migration decision. `required`, `enum`, and edge `fromKinds` / `toKinds` are now compared as the sets they are, in both the modified-vs-unmodified decision and the breaking-change severity classification. Genuine changes — added or removed properties, newly required properties, changed enum members, different edge endpoints — are detected exactly as before. The normalization is deliberately scoped to diff comparison and is **not** applied to the canonical form behind `computeSchemaHash`, so no schema hash already committed to a database changes. The normalization walks the document as JSON Schema rather than as plain JSON, because a key's meaning depends on where it appears. Recursion is an **allowlist** of known schema-valued keywords; everything else is preserved verbatim: - Instance data (`default`, `const`, `examples`) and unknown extension keys — Zod's `.meta()` merges arbitrary keys straight into the generated schema — are compared verbatim. Recursing into them would sort a nested key merely _named_ `required`, silently normalizing away a real change to a stored value. - Keys under `properties`, `patternProperties`, `dependentSchemas`, `$defs`, and `definitions` are user-chosen field names, not keywords, so a field _named_ `default` still has its subschema normalized like any other. - `dependentRequired` maps a name to a set of names, so each set is order-normalized. The allowlist fails in the safe direction: an unrecognized schema-valued keyword is left unsorted, so a reordering inside it reads as a change rather than being hidden. ## 0.41.0 ### Minor Changes - [#311](https://github.com/nicia-ai/typegraph/pull/311) [`008fa20`](https://github.com/nicia-ai/typegraph/commit/008fa2008d199f997f45af967ba2f1a1fbad4970) Thanks [@pdlug](https://github.com/pdlug)! - Make verified adapter stores reusable across connections so serverless/edge deployments that open a fresh database connection per request can verify once per isolate instead of paying a schema-reconcile round-trip on every request. `AdapterStore` now exposes `reconciledSchema`, an opaque snapshot of a store's reconciled (compile-time + runtime-committed) graph and committed schema version. Pass it to a synchronous `createAdapterStore(graph, backend, { reconciled })` — which issues **zero** database queries and still validates reads and writes against runtime-committed kinds — or call `store.withBackend(freshBackend)` to rebind an already-verified store onto a new connection with no re-verify (the store's connection is captured immutably, so this returns a new equivalent store rather than mutating in place). The new `getCommittedSchemaVersion(backend, graphId)` reads the committed version with a single indexed SELECT, the cheap cross-isolate probe for detecting when another process committed a schema change and the cached snapshot must be refreshed. ## 0.40.0 ### Minor Changes - [#308](https://github.com/nicia-ai/typegraph/pull/308) [`db2dc31`](https://github.com/nicia-ai/typegraph/commit/db2dc31e2a66f3195c0d5d7e3df19864cb64672c) Thanks [@pdlug](https://github.com/pdlug)! - Replace timestamp-only `RecordedInstant` values with versioned anchors that encode a strict per-graph logical revision alongside a non-decreasing physical wall-time high-water mark. Recorded relations store numeric revisions while the public anchor remains one durable string. Upgrade timestamp-only preview tables with `migrateLegacyRecordedTime()` and remap external checkpoints with `migrateRecordedAnchor()`. Driver timestamps are normalized without host-local timezone parsing, migration integrity failures are typed, and the retained anchor map can be dropped automatically after its final graph is cleaned up. History-enabled async store factories now reject an unmigrated recorded schema at open, including when the legacy tables are empty. ## 0.39.0 ### Minor Changes - [#306](https://github.com/nicia-ai/typegraph/pull/306) [`cd4e0eb`](https://github.com/nicia-ai/typegraph/commit/cd4e0ebf1fa8a51e8a965e667801941a5360e097) Thanks [@pdlug](https://github.com/pdlug)! - Add bounded, deterministic `scan()` pagination to recorded-time node and edge collections so adapters can reconstruct complete historical snapshots without retaining a separate identity inventory. - [#305](https://github.com/nicia-ai/typegraph/pull/305) [`4349766`](https://github.com/nicia-ai/typegraph/commit/4349766fc9d5db7046f05cefab6e56f4dd4d655a) Thanks [@pdlug](https://github.com/pdlug)! - Harden adapter capability surfaces and document their migrations. This is source-breaking for adapter code that reads `tx.sql` without first narrowing `tx.sqlAvailability === "available"`: non-available union arms now omit `sql` instead of exposing it as an optional `never`/`undefined` property. The runtime history and revision-tracking guards remain fail-loud for JavaScript and type-suppressed callers. Add `openProvenanceStore(targetStore)` as the preferred graph-merge provenance API while retaining `openProvenanceStore(backend, targetGraphId)` for standalone inspection tools. On Cloudflare D1 and Durable Object SQLite, ignore only a recognized `SQLITE_AUTH` rejection of the performance-only `analysis_limit` PRAGMA and continue with scoped `ANALYZE`; unexpected maintenance failures stay visible. ## 0.38.0 ### Minor Changes - [#297](https://github.com/nicia-ai/typegraph/pull/297) [`474afe6`](https://github.com/nicia-ai/typegraph/commit/474afe602f44ef60e313015d853c4705b30e8790) Thanks [@pdlug](https://github.com/pdlug)! - Add global and weighted multi-seed personalized PageRank with induced-subgraph, temporal-view, direction, convergence-tolerance, and working-memory options. - [#303](https://github.com/nicia-ai/typegraph/pull/303) [`58855f7`](https://github.com/nicia-ai/typegraph/commit/58855f7b9e4f808805aecb4d1169fadd8e6aaab1) Thanks [@pdlug](https://github.com/pdlug)! - Move query compilation behind TypeGraph-owned backend and SQL-fragment abstractions so strict consumers no longer typecheck unused Drizzle dialect declarations. Add a Drizzle-free `core` entrypoint and managed full-Store entrypoints for local SQLite and PGlite, with packed TypeScript 5 and 6 regression coverage for both databases. The portable `@nicia-ai/typegraph/indexes` entrypoint is also Drizzle-free. Direct Drizzle index-builder helpers moved to `@nicia-ai/typegraph/adapters/drizzle/indexes`. Advanced adapter APIs now use TypeGraph's `SqlFragment` instead of Drizzle `SQL`: this includes query `compile()` results, custom `GraphBackend` implementations, and custom fulltext/vector strategies. Use `toSQL()` for a dialect-rendered `{ sql, params }` result, or `renderSqlite()` / `renderPostgres()` when rendering a fragment directly. Custom backend and strategy authors can import the complete, Drizzle-free contract vocabulary from `@nicia-ai/typegraph/backend`. This entrypoint names every operation parameter, row, strategy payload, dialect port, SQL fragment chunk, and supporting schema/index type referenced by those contracts. API Extractor enforces zero forgotten exports for new entrypoints and fingerprints the complete pre-existing debt set so added or removed leaks cannot pass silently. The default `Store` is now the portable TypeGraph surface. It keeps the full graph API and graph-owned transactions while omitting adapter-native handles and caller-owned transaction adoption. Drizzle integration entrypoints return `AdapterStore` when precisely typed `tx.sql`, `withTransaction`, or `withRecordedTransaction` interoperability is required. `createStore`, `createStoreWithSchema`, and `createVerifiedStore` now return the portable contract; use their `createAdapterStore`, `createAdapterStoreWithSchema`, and `createVerifiedAdapterStore` counterparts when the application deliberately needs adapter-native interoperability. Migration map: | 0.37 use | 0.38 replacement | | ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | `createStoreWithSchema(...)` followed by `tx.sql` | `createAdapterStoreWithSchema(...)` | | `Store` | `AdapterStore` | | `TransactionContext` | `AdapterTransactionContext` | | `HistoryTransactionContext` | `TransactionContext` for portable history stores, or `AdapterHistoryTransactionContext` for adapter history stores | | `AdoptedTransaction` | The adapter's concrete native-handle type, such as `AnySqliteDatabase` or `AnyPgTransaction`, passed as the adapter generic | `HistoryTransactionContext` and `MeasurableHistoryTransactionContext` were removed rather than retained as aliases. Portable stores have one SQL-free transaction context across live and history modes. Adapter contexts expose `sql` only after `sqlAvailability === "available"`; the other discriminated union arms omit the property. Store evolution is now generic over the exact Store flavor, so adapter/history/recorded-read surfaces and compatible `StoreRef` values are preserved without downstream casts. Remove casts that existed only to restore the old widened evolution result. Requiring the `sqlAvailability` check is source-breaking for adapter code that previously read `tx.sql` from the unnarrowed union. Narrow on the discriminant before passing the handle to even an `unknown`-typed sink. `GraphBackend` is now the portable TypeGraph backend port. Native transaction adoption lives on `AdapterBackend`, so a capability-less backend cannot be passed to an adapter-store factory. Portable transaction contexts and adapter transaction contexts both expose the same runtime-enforced read-only `TransactionReadBackend`. Adapter contexts add only the precisely typed native `sql` handle; TypeGraph internals reach the full transaction backend through a non-public, non-enumerable runtime port. Backend functions are now receiver-free (`this: void`); custom backends must close over their state instead of depending on method receivers. PostgreSQL adapter stores now expose and accept `AnyPgTransaction` for native transaction interoperability. A root Drizzle PostgreSQL database is rejected at compile time and runtime; pass only the transaction handle received by a caller-owned `db.transaction(...)`. SQLite adoption remains database-handle based so its documented manual-`BEGIN` integration continues to work. Public `TransactionOptions` now contains only caller-selectable isolation and access modes. TypeGraph's temporary-write authorization is an internal, globally branded capability and is no longer expressible through the public transaction contract. Fulltext and vector strategy members are readonly function properties, closing TypeScript's method-bivariance loophole for third-party implementations. Dialect adapter members use the same receiver-free function-property contract, and `TransactionOptions` is exported from the root entrypoint for portable transaction consumers. Managed SQLite and PGlite factories preserve the precise live, history, or recorded-read Store flavor selected by their options, including when options are widened before the call. This keeps unavailable write and native-adapter capabilities unrepresentable instead of relying on runtime failures. Every Store flavor exposes the safe, Drizzle-free `store.capabilities` descriptor for runtime feature checks without exposing backend operations. `AdapterHistoryStore.backend` exposes the narrower `HistoryStoreBackend`, which omits raw SQL, native import, graph clearing, and nested backend transactions so capture-bypassing writes are absent at both type and runtime levels. Backend capability narrowing now uses an exhaustive runtime allowlist instead of default-forwarding proxy overlays. New `GraphBackend` members must be classified explicitly, preventing adapter capabilities from leaking through a history wrapper. Store evolution also preserves each refined Store flavor and accepts invariant `StoreRef` values for that exact replacement surface. Add checked-in API Extractor reports derived from every package export. CI now fails when the emitted public declaration surface changes without an intentional report update. Direct SQL fragment values now pass through the same dialect binding normalization as placeholders and compiled queries. Runtime store, transaction, schema, and recorded-read ports use versioned global symbols so mixed ESM/CJS or duplicated bundle instances interoperate safely. Dialect policies outside the compiler are exhaustive records or switches, so adding a new SQL dialect cannot silently inherit SQLite behavior. Remove the transitional `SQL`, `SqlRenderDialect`, and `AdoptedTransaction` aliases. Import `SqlFragment` and `SqlDialect` directly. The constructors that brand arbitrary fragments as executable SQL are now internal; public compiled SQL values come from TypeGraph's query compiler. Managed local stores now live at `/sqlite/local` and `/postgres/pglite`; bring-your-own-connection APIs live under `/adapters/drizzle/sqlite...` and `/adapters/drizzle/postgres...`. The old `@nicia-ai/typegraph/sqlite` and `@nicia-ai/typegraph/postgres` entrypoints were removed. They are not compatibility aliases because those names now distinguish the managed Store API from bring-your-own-connection adapters. Move imports as follows: | 0.37 entrypoint | 0.38 entrypoint | | ------------------------------------------------- | ----------------------------------- | | `/sqlite` | `/adapters/drizzle/sqlite` | | `/sqlite/local` for `createLocalSqliteBackend` | `/adapters/drizzle/sqlite/local` | | `/sqlite/libsql` | `/adapters/drizzle/sqlite/libsql` | | `/postgres` | `/adapters/drizzle/postgres` | | `/postgres/pglite` for `createLocalPgliteBackend` | `/adapters/drizzle/postgres/pglite` | The managed `/sqlite/local` and `/postgres/pglite` entrypoints keep Drizzle out of their public declarations, but their built-in database implementation still uses Drizzle internally. `drizzle-orm` therefore remains a required peer of the 0.38 package; declaration isolation does not imply installation isolation. - [#298](https://github.com/nicia-ai/typegraph/pull/298) [`f178663`](https://github.com/nicia-ai/typegraph/commit/f1786634e7867b892805349dc922269e155d1d65) Thanks [@pdlug](https://github.com/pdlug)! - Add deterministic synchronous label propagation with exact binary tie-breaking, induced node-kind scope, temporal and recorded-time views, bind-independent neighbor voting, early period-two oscillation detection, and an `onMaxIterations` completion contract: `"throw"` (default) returns only a converged labeling, while `"return"` yields the exact fixed-round Graphalytics CDLP labeling. ## 0.37.1 ### Patch Changes - [#292](https://github.com/nicia-ai/typegraph/pull/292) [`0152c3b`](https://github.com/nicia-ai/typegraph/commit/0152c3bf4a0bd3c047931041d1505b69a25fa05a) Thanks [@pdlug](https://github.com/pdlug)! - Restore graph algorithms on Cloudflare Durable Objects SQLite. The auto-detected `do-sqlite` profile now marks temporary-table graph analytics as unsupported, routes shortest-path and reachability algorithms through their inline fallback, and rejects temporary-table-only algorithms with the existing typed capability error instead of leaking workerd's `SQLITE_AUTH` failure. - [#293](https://github.com/nicia-ai/typegraph/pull/293) [`9309ec3`](https://github.com/nicia-ai/typegraph/commit/9309ec3f839b53474d389b9376f7851c87753e28) Thanks [@pdlug](https://github.com/pdlug)! - Speed up exact weakly connected components with indexed changed-label frontiers, changed-row-only writes, and one fewer working-table join. Preserve synchronous convergence across bind-limited edge-kind chunks, and align shortest-path identity tie-breaks with portable binary ordering. ## 0.37.0 ### Minor Changes - [#269](https://github.com/nicia-ai/typegraph/pull/269) [`92479d4`](https://github.com/nicia-ai/typegraph/commit/92479d44ba7f0fc76985d51174cb9801c055fd4d) Thanks [@pdlug](https://github.com/pdlug)! - Vector storage now rides the [#135](https://github.com/nicia-ai/typegraph/issues/135) durable-contribution machinery, so the runtime never issues DDL on the embedding hot path. Previously every vector op (`upsertEmbedding` / `deleteEmbedding` / `vectorSearch` / `createVectorIndex`) lazily ran `CREATE TABLE IF NOT EXISTS` for its per-`(kind, field)` table on whatever connection it executed on. On a least-privilege Postgres role (USAGE on `public`, full DML, but no `CREATE`) this failed with `permission denied for schema public` (SQLSTATE 42501) — even when the table already existed, because Postgres runs the schema aclcheck before the `IF NOT EXISTS` short-circuit. The fulltext path already avoided this via durable markers; vectors now do too. What changed: - **Boot (privileged):** `createStoreWithSchema` provisions every embedding `(kind, field)` table + a durable contribution marker, enumerated from the graph. `evolve()` provisions any embedding fields it introduces. A slot already provisioned at a _different_ shape (the declared dimension changed) is warned about and left untouched — boot stays reachable so `store.reembedVectorField()` can recreate it; until then, writes to that field fail with a `stale` `StoreNotInitializedError` that points at `reembedVectorField`. - **Runtime writes (DML-only):** `upsertEmbedding` (single and batch) and `deleteEmbedding` assert the durable marker with a cached, signature-checked SELECT and run DML — never DDL. `createVerifiedStore` verifies vector markers at attach, alongside fulltext. - **Vector reads are not marker-gated:** `store.search.vector`, `store.search.hybrid`, and query-builder `.similarTo()` predicates compile to SQL against the per-field table directly (searches may override the metric at query time, so their slot legitimately differs from the provisioned shape); against an un-provisioned database they surface the engine's missing-relation error, which `createVerifiedStore` catches at attach. - `reembedVectorField` re-stamps the marker after recreating storage at a new dimension; vector-field reclaim (`materializeRemovals`) clears the marker when it drops a table. **Breaking:** vector ops now require a prior privileged `createStoreWithSchema` (exactly as fulltext already does). A plain `createStore` + embedding write with no provisioning step throws `StoreNotInitializedError` instead of lazily creating the table. **Migration:** after upgrading, run `createStoreWithSchema(graph, adminBackend)` once under the schema-owner role. It creates the per-field vector tables + markers; least-privilege runtimes then assert markers (SELECT) and run vector DML with zero DDL — no `GRANT CREATE` required. Consumers that boot manually (raw DDL + the sync `createStore` attach + `backend.ensureRuntimeContributions`) provision vectors the same way: the new `resolveGraphVectorSlots(graph)` export enumerates every embedding `(kind, field)` slot, and `backend.ensureVectorSlotContribution(slot)` materializes each — the exact step `createStoreWithSchema` performs. Batch counterparts (`backend.ensureVectorSlotContributions(slots)` / `backend.assertVectorSlotsInitialized(slots)`) resolve every slot's markers with one graph-scoped query — what boot and verified attach use, and the right choice for many embedding fields over a remote connection. - [#284](https://github.com/nicia-ai/typegraph/pull/284) [`26f5b4a`](https://github.com/nicia-ai/typegraph/commit/26f5b4a3f129353e0ec92040497ddca685563c7f) Thanks [@pdlug](https://github.com/pdlug)! - TypeGraph's base-relation indexes are now **system-index declarations** — a single declared list (`SYSTEM_INDEX_DECLARATIONS`) that both dialect schemas derive from and that materializes onto already-initialized databases. Previously the base indexes were hand-written twice (once per dialect schema) and applied only by first-boot bootstrap DDL, so an index added in a newer library version never reached an existing database without manual DDL (the gap [#282](https://github.com/nicia-ai/typegraph/issues/282) exposed). Now: - **Single source, parity by construction.** `createSqliteTables` / `createPostgresTables` build their node/edge/recorded-relation indexes from the same declarations, and a cross-dialect extraction test asserts the two generated DDL scripts' full index sets stay identical. - **Upgrade path.** `createStoreWithSchema` brings a database's system indexes up to the running library version at boot — `CREATE INDEX CONCURRENTLY` on PostgreSQL, riding the same status table, drift signatures, invalid-leftover healing, and cross-caller claim protocol as graph-declared indexes. A database whose indexes all exist settles from three concurrent catalog/status reads (scoped to the session `search_path`, so schema-per-tenant databases never observe each other's indexes) with no index DDL and no status writes — the only DDL on that warm path is the idempotent status-table `CREATE TABLE IF NOT EXISTS` ensure step every materialize verb runs. A system index that is physically absent or invalid is rebuilt even when a stale success row survives (dump/restore, manual drop). Failures — including status-table infrastructure errors — degrade to a warning: indexes are a performance concern and the store still boots. Deployments that must not run index builds inline at boot pass `systemIndexes: "skip"` to `createStoreWithSchema` and materialize out-of-band. - **New API: `store.materializeSystemIndexes()`** for deployments that boot without `createStoreWithSchema` (zero-DDL attach) — call once under a DDL-capable role after upgrading. Strict where the boot path is lenient: throws `ConfigurationError` on backends without DDL/status primitives. - `IndexEntity` gains a `"system"` member; system status rows carry the relation key (e.g. `"recordedNodes"`) in their `kind` column. Generated DDL is unchanged for default and short custom table names — same index names, columns, and order — so existing databases and drizzle-kit migrations are unaffected. Names that would exceed PostgreSQL's 63-char identifier bound (very long custom table names) are now deterministically truncated + hash-suffixed instead of being silently truncated by the engine into collisions. System index names are reserved: a graph-declared index using one is rejected at table definition and by `materializeIndexes()` (previously its `CREATE INDEX IF NOT EXISTS` silently no-opped against the differently-shaped system index while recording success). Legacy databases that predate the recorded relations skip those indexes cleanly instead of attempting failing DDL at every boot. - [#273](https://github.com/nicia-ai/typegraph/pull/273) [`42f6941`](https://github.com/nicia-ai/typegraph/commit/42f6941fdb601629c0f45ec54fa9f3e50bd028ae) Thanks [@pdlug](https://github.com/pdlug)! - Add `trustedImportGraph` and `trustedImportGraphStream` for atomic initial loads into a fresh, dedicated database. The distinct trusted surface bypasses schema, reference, cardinality, and conflict validation; uses prepared SQLite writes or PostgreSQL `UNNEST` ingestion; defers rebuildable secondary indexes; refreshes planner statistics; and rolls the complete stream back on any failure. The first version rejects non-empty TypeGraph data tables, recorded history, revision tracking, uniqueness constraints, searchable fields, vector fields, and backends without the required native transactional path. - [#279](https://github.com/nicia-ai/typegraph/pull/279) [`c44eeac`](https://github.com/nicia-ai/typegraph/commit/c44eeac36805be144b25123fafb5c52ae30b73a4) Thanks [@pdlug](https://github.com/pdlug)! - Add exact `store.algorithms.weaklyConnectedComponents()` for transactional SQLite and PostgreSQL backends. Results include deterministic component representatives and sizes, honor valid/recorded temporal views, and fail with a typed convergence error instead of returning partial labels when the configured iteration budget is exhausted. Callers can restrict WCC to a `nodeKinds` induced subgraph, retaining isolated in-scope nodes without seeding unrelated node kinds. PostgreSQL iterative operations now refresh temporary-table planner statistics after sufficiently large seeds and multiplicative growth, avoiding plans based on the engine's initial one-row estimate. The policy also covers growing BFS working tables and is a no-op on SQLite. Set-based reachability now deduplicates edge targets before target-node visibility checks and avoids computing unused predecessor paths. This reduces dense-frontier work while preserving minimum-depth results and cross-backend semantics. - [#288](https://github.com/nicia-ai/typegraph/pull/288) [`17a3f83`](https://github.com/nicia-ai/typegraph/commit/17a3f83d806e0b77748c346f5c320d8d55080b91) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.algorithms.weightedShortestPath` — a minimum-total-weight path search weighting each traversed edge by a numeric edge property (LDBC Interactive IC14 shape). Runs frontier-based relaxation on the shared iterative substrate with best-target pruning, works on both execution paths (temporary working table and inline fallback), and honors valid-time and recorded-time coordinates including pinned StoreViews. Edge weights are audited up front: negative, non-numeric, out-of-range, or (without `defaultWeight`) missing weights throw the new typed `InvalidEdgeWeightError`. Weight arithmetic is IEEE 754 double precision on both backends, so total weights are backend-identical; among equal-total-weight paths the returned node sequence is too, except when the `edges` list exceeds the backend's bind-parameter budget (hundreds of edge kinds in one call). ### Patch Changes - [#282](https://github.com/nicia-ai/typegraph/pull/282) [`923219d`](https://github.com/nicia-ai/typegraph/commit/923219d6854a0a97cc186c2f4f27e6564bb935be) Thanks [@pdlug](https://github.com/pdlug)! - Add a `(graph_id, id)` index to the live and recorded node tables so bare-id lookups (a node's `id` without its `kind`) seek instead of scanning the graph's node partition — the composite keys lead with `kind`, so they can't serve that probe. `store.algorithms.degree()`'s node-kind subquery is the main consumer: ~95 ms → sub-millisecond at LDBC SNB SF1 (3.16M nodes) on SQLite, at the live and recorded coordinates alike. New databases get both indexes at bootstrap. Existing databases adopt them with a one-time `await backend.bootstrapTables()` — every statement is `CREATE … IF NOT EXISTS`, so the call is idempotent and only creates what's missing. On PostgreSQL this issues a plain `CREATE INDEX` (briefly locks writes on large tables); schedule it, or apply the equivalent `CREATE INDEX CONCURRENTLY` statements manually. - [#265](https://github.com/nicia-ai/typegraph/pull/265) [`35ab2a0`](https://github.com/nicia-ai/typegraph/commit/35ab2a02af728df9059750518ddbdd12e489450e) Thanks [@pdlug](https://github.com/pdlug)! - Docs: scope the `coalesceUnchangedUpserts` benefit correctly. Coalescing eliminates _re-delivery_ churn (an already-applied change delivered again, value-identical to the live row). It does not make a full replay-from-zero free when the stream supersedes values in place: re-applying an older value over the live row is a genuine change, and restoring the current value afterwards is another, so such a replay still writes — and leaves a spurious back-and-forth band in the live store's recorded history. Churn-free rebuilds replay into a fresh store instead. Clarified in the option's TSDoc and in the "Materializing external event logs" guide; no behavior change. - [#289](https://github.com/nicia-ai/typegraph/pull/289) [`199b33a`](https://github.com/nicia-ai/typegraph/commit/199b33aabffa304b6a11f34ad7ca9b0e5f449218) Thanks [@pdlug](https://github.com/pdlug)! - Cap SQLite-backed Durable Object statements at Cloudflare's 100-bound-parameter limit. Structural client detection now makes platform identity authoritative over stale execution hints, and capability overrides cannot raise the hard ceiling. Recorded-history capture and every capability-driven SQLite batch path chunk large writes before workerd rejects the query, while SQLite literal list predicates use one JSON-bound parameter instead of one bind per element. - [#285](https://github.com/nicia-ai/typegraph/pull/285) [`9949562`](https://github.com/nicia-ai/typegraph/commit/99495623057610e502671515709c29d8f5139ae2) Thanks [@pdlug](https://github.com/pdlug)! - Cut three overheads out of the iterative graph algorithms, root-caused with `EXPLAIN (ANALYZE, BUFFERS)` against LDBC SNB SF1 on PostgreSQL. Weakly connected components no longer re-validates node visibility per edge in its propagate rounds. The working table is seeded through the same graph/kind/temporal filters inside the same snapshot and both edge endpoints are already joined against it, so membership is the visibility proof; the per-edge `typegraph_nodes` index loops (hundreds of thousands per round on SF1) added nothing. Results are byte-identical. Traversal rounds now carry their own bookkeeping instead of issuing follow-up statements: seeding returns the frontier through `INSERT … RETURNING`, and bidirectional shortest-path rounds detect the frontier meeting inside the expansion statement rather than with a separate probe per round. A shortest-path traversal that used to issue two to three statements per round now issues one, roughly halving round-trip latency on latency-bound connections. The working-table `ANALYZE` policy is unchanged in its thresholds but no longer runs when no further round will read the table. When several equal-depth meetings exist, the tie now breaks by node id then kind in code-unit order on both backends — previously the selection followed the database collation, so a PostgreSQL cluster with a linguistic default collation could pick a different (equally shortest) path. New option: iterative algorithm calls (`reachable`, `shortestPath`, `canReach`, `neighbors`, `weaklyConnectedComponents`) accept `workingMemory?: string`, an opt-in, transaction-scoped override of the session's `work_mem`, applied on PostgreSQL with `SET LOCAL` semantics via parameterized `set_config`. By default (option omitted) operations inherit the server's configured `work_mem` — nothing is overridden. `work_mem` is a threshold each sort/hash operator (and each parallel worker) may allocate up to, not a per-operation budget, and concurrent calls multiply it; set it deliberately (e.g. `"64MB"`) for large single-tenant analytical runs where the configured default spills whole-graph sorts to disk (measured ~106MB external merges per WCC round on SF1). The override never touches the session or server setting, is validated as `kB|MB|GB` within PostgreSQL's accepted `work_mem` range (64kB–2147483647kB) with the same typed error on both backends, and is ignored by SQLite. - [#290](https://github.com/nicia-ai/typegraph/pull/290) [`247c1b7`](https://github.com/nicia-ai/typegraph/commit/247c1b77b8c30d2f03520d52214bfaebbf1a0e6c) Thanks [@pdlug](https://github.com/pdlug)! - Fix PostgreSQL pointer-level `pathIsNull()` / `pathIsNotNull()` predicates misclassifying two stored value shapes. The previous text-comparison form (`#>> path = 'null'`) went three-valued on a stored JSON `null` — so `pathIsNull()` silently failed to match those rows on PostgreSQL while matching them on SQLite — and misread the JSON _string_ `"null"` as null, falsely matching it with `pathIsNull()` and excluding it from `pathIsNotNull()`. Both predicates are now type-based (`jsonb_typeof`) and never SQL NULL, converging on SQLite's (correct) semantics. Field-level `isNull()` / `isNotNull()` predicates were already correct and are unchanged. Behavior change on PostgreSQL for affected data: rows holding a JSON `null` now match `pathIsNull()`, and rows holding the string `"null"` no longer do. - [#283](https://github.com/nicia-ai/typegraph/pull/283) [`8306680`](https://github.com/nicia-ai/typegraph/commit/830668032328e5fdc031deccfaf0c452f466a15c) Thanks [@pdlug](https://github.com/pdlug)! - Selective `ORDER BY … LIMIT` queries now compile with late materialization: the query sorts and limits a lean candidate set carrying only identity, sort keys, and predicate columns, then re-fetches the deferred projection columns by primary key for only the surviving rows — instead of extracting every projected column for every candidate and discarding all but the `LIMIT` survivors after the sort. At LDBC SNB SF1, IC9's top-20 over a 1.18M-comment fan-out stops extracting `content` 1.18M times, ~30–37% faster on SQLite. The transform fires only on the selective `.select()` path with `ORDER BY` and a positive `LIMIT` at the live coordinate. Aggregates, vector/fulltext, optional (LEFT JOIN) traversals, edge-field projections, non-selective queries, and recorded-time reads keep the flat plan unchanged. - [#274](https://github.com/nicia-ai/typegraph/pull/274) [`2a889aa`](https://github.com/nicia-ai/typegraph/commit/2a889aaea095a842fdd6f1b4a97feba5e6026d82) Thanks [@pdlug](https://github.com/pdlug)! - Replace path-enumerating recursive CTEs in `reachable`, `neighbors`, `shortestPath`, and `canReach` with set-based breadth-first search. Transactional SQLite and PostgreSQL backends now execute graph iterations against a connection-local temporary working table, de-duplicated by node kind and ID on every round. Non-transactional backends retain parity through a bind-limit-aware inline frontier. Traversals run in one snapshot where the backend supports transactions, preserve temporal filtering, and clean up temporary state on success or failure. ## 0.36.0 ### Minor Changes - [#261](https://github.com/nicia-ai/typegraph/pull/261) [`5bc7b53`](https://github.com/nicia-ai/typegraph/commit/5bc7b5333d30392c31605161436441b3e8602447) Thanks [@pdlug](https://github.com/pdlug)! - Return a receipt from `store.withRecordedTransaction`, and add scoped write measurement with `tx.measure`. - **`store.withRecordedTransaction(externalTx, fn)` now returns `Promise>`** instead of `Promise`. The adopted path is the only way to get exactly-once cursors and graph writes atomically on a history store, and it now surfaces the same receipt `transactionWithReceipt` does: `receipt.writes` for dropped-change detection and `receipt.recorded` as the per-transaction replay anchor (`undefined` for a read-only callback or a non-history store). **BREAKING:** the adopted path now returns the result under `.result`. Migrate by destructuring: ```typescript // Before const x = await store.withRecordedTransaction(externalTx, fn); // After const { result: x } = await store.withRecordedTransaction(externalTx, fn); ``` - **Scoped receipts — `tx.measure((scoped) => ...)`.** On the receipt-enabled contexts (`transactionWithReceipt`, `withRecordedTransaction`), `tx.measure` runs its callback with a **scoped context** — a second view over the same transaction — and returns a `TransactionOutcome` whose receipt counts exactly the writes made **through that scoped context** (`scoped.nodes` / `scoped.edges`). So a framework can attribute writes to user code it invoked (e.g. a materializer measuring `project(scoped, change)` to detect a dropped change) while its own bookkeeping — written through the outer `tx` — stays out of the count. Attribution is by which context you write through, not by timing, which makes overlapping and concurrent measures safe by construction (two scopes racing under `Promise.all` never cross-count). Nesting composes; measured writes still count in the outer receipt; a scoped receipt's `recorded` is always `undefined`. Plain `store.transaction()` contexts have no `measure` (that path runs no recorder and stays zero-overhead). New exported types: `MeasurableTransactionContext`, `MeasurableHistoryTransactionContext`, `ScopedMeasure`. - **Adopted contexts seal on return.** A transaction context retained and written through _after_ its `withRecordedTransaction` callback resolves now fails loud on both paths — the history path's capture guard is checked _before_ the live write (so a swallowed error can no longer commit an uncaptured row), and the non-history path seals its receipt-tracked collections (so a post-return write can't persist a row the already-returned receipt never counted). - [#262](https://github.com/nicia-ai/typegraph/pull/262) [`34468a0`](https://github.com/nicia-ai/typegraph/commit/34468a04f9cb6abff34177282d31f1240f1254d1) Thanks [@pdlug](https://github.com/pdlug)! - Add an opt-in `coalesceUnchangedUpserts` store option for at-least-once / replay materializers. Idempotent event-log projectors converge live state correctly, but every re-delivery of a byte-identical value still performed a real write: `upsertById` on an existing id called `updateNode` unconditionally, allocating a fresh recorded instant and a new history row. A full replay of an N-event log therefore rewrote every row and grew recorded history by N — the recovery / rebuild workload inflates history the most. With `createStore(graph, backend, { coalesceUnchangedUpserts: true })`, an `upsertById` (or `bulkUpsertById` item) whose validated props are value-identical to the existing **live** row performs **no write at all**: no `updateNode`, no recorded-time capture, no history row, no revision-anchor advance, and no `update` operation hooks. It resolves with the existing node. The dirty-check compares the storage-normalized representation (props run through the kind's Zod schema, key-order-independent), so it answers exactly "would the persisted value differ?". A write still happens (never coalesced) when the row is soft-deleted (an upsert resurrects it), when an explicit `validFrom` / `validTo` is passed, or when any prop differs. Default off, because some consumers want an audit row per re-delivery. Covered symmetrically for edge `bulkUpsertById` (props only — endpoints are the edge's identity). Receipt semantics are unchanged and need no new signal: a coalesced upsert still counts as one write intent (`writes.total`) but captures nothing (`recorded` stays `undefined`) — the same two-signal shape as a no-op delete, which at-least-once consumers already handle by carrying the prior anchor forward. - [#260](https://github.com/nicia-ai/typegraph/pull/260) [`35d03ae`](https://github.com/nicia-ai/typegraph/commit/35d03ae0fd4647286927d617210f20a7e47df4b6) Thanks [@pdlug](https://github.com/pdlug)! - Make the store transaction surface tell the truth about raw SQL and history capture. - **New `tx.sqlAvailability` discriminant.** Every transaction context now carries a required `sqlAvailability: "available" | "history" | "revisionTracking" | "unavailable"` field. Branch on it instead of truthiness-testing `tx.sql`: under `history: true` / `revisionTracking: true` the raw handle is present-but-throwing (so `if (tx.sql)` read truthy and then threw), and it is `undefined` only on the non-transactional fallback. `"available"` means `tx.sql` is a usable raw handle; `"history"` / `"revisionTracking"` mean raw SQL is disabled here; `"unavailable"` means the backend has no transactions (`tx.sql === undefined`, no atomicity). - **`store.withTransaction()` on a history-enabled store is now a compile error.** It always threw at runtime; the call site now rejects the argument with a message pointing at `store.withRecordedTransaction()`. The runtime guard is unchanged for suppressed calls. - **Branchable recorded-capture guard codes.** The `ConfigurationError`s these guards throw carry a stable `details.code` (`RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION`, `RECORDED_CAPTURE_RAW_SQL_DISABLED`, `REVISION_TRACKING_RAW_SQL_DISABLED`), now exported as `RECORDED_CAPTURE_GUARD_CODES` with a `RecordedCaptureGuardCode` type and an `isRecordedCaptureGuardError(error, code?)` type guard — so a portable caller can distinguish "history forbids raw SQL here" from "this backend has no transactions" without substring-matching the message. - **Fixed `withRecordedTransaction`'s JSDoc**, which incorrectly promised `tx.sql`; on the adopted path you already hold the pinned connection, so write your own relational tables through the external transaction handle you passed in. ## 0.35.0 ### Highlights TypeGraph 0.35 improves bulk ingestion, search, and traversal performance across SQLite and PostgreSQL. Bulk creation and import batch their validation and side effects, large autocommit loads refresh planner statistics, and built-in hybrid search combines retrieval, fusion, and hydration in one SQL statement. Search gains property filters, pagination, subclass scope, and explicit approximate vector retrieval; exact retrieval remains the default. Graph branches can now use durable revision origins and counters to validate their base without fingerprinting every live row. Streaming interchange supports larger graph copies, and `transactionWithReceipt()` exposes completed collection write intents and recorded commit coordinates. Index declarations gain GIN, trigram, and system-column keys. Correctness fixes keep repeated and prepared queries bound to a fresh read instant, preserve temporal metadata through branch copies, and enforce exact vector results even when an approximate index exists. Current temporal reads now use the application clock consistently with writes. ### Upgrade notes - Audit persisted schemas as well as source definitions for endpoint-incompatible `implies()` relations. These are now rejected when constructing or loading a registry. - Backend rows expose `props` as `RowProps`, which may be JSON text or a parsed object. Direct backend consumers should use `rowPropsToObject()` or `rowPropsToJsonText()` instead of unconditionally calling `JSON.parse()`. - Custom vector strategies must declare the new filtered approximate-search capability, including whether they guarantee a full result page. - Existing databases retain their old edge traversal indexes until those indexes are explicitly rebuilt. The detailed entries below include the PostgreSQL and SQLite rebuild procedures; rerunning `CREATE INDEX IF NOT EXISTS` alone does not widen an existing index. - New records without an explicit `validFrom` now begin at their creation instant. Existing open-left records retain their stored bounds, and interchange preserves those bounds through branch copies. Approximate vector retrieval requires `{ approximate: true }`; applications that previously received approximate results on the default path may see a different cost for the corrected exact search. ### Minor Changes - [#231](https://github.com/nicia-ai/typegraph/pull/231) [`839f536`](https://github.com/nicia-ai/typegraph/commit/839f53621998d41704537e45408872d49452cf1c) Thanks [@pdlug](https://github.com/pdlug)! - Aggregate queries now support `.orderBy()`. Previously `ExecutableAggregateQuery` exposed `limit()` but no way to order results, so `.aggregate({...}).limit(n)` returned an arbitrary `n` groups rather than the top `n` — the most common aggregate shape ("top N groups by count/sum") required fetching every group and sorting in JS. `.orderBy(key, direction?)` takes any output name from `.aggregate({...})` — either a grouped field or an aggregate alias — and can be chained for multi-key sorts: ```typescript store .query() .from("Author", "a") .traverse("wrote", "e") .to("Book", "b") .groupByNode("a") .aggregate({ author: field("a", "name"), bookCount: count("b") }) .orderBy("bookCount", "desc") .limit(2) .execute(); ``` Ordering resolves against the projected SELECT-list output alias rather than recompiling the underlying expression, so it works uniformly for grouped fields and aggregates on both SQLite and PostgreSQL with no dialect-specific handling. - [#212](https://github.com/nicia-ai/typegraph/pull/212) [`dcdd542`](https://github.com/nicia-ai/typegraph/commit/dcdd54246fef1e93839196d7029e4dbadbc72b42) Thanks [@pdlug](https://github.com/pdlug)! - Autocommit `bulkCreate` and `bulkInsert` calls (nodes and edges) now refresh planner statistics automatically when a single call writes 1,000 rows or more, closing the stale-statistics window after bulk loads where the planner keeps pre-load row estimates until ANALYZE runs (observed 25-200x slowdowns on traversal and fulltext shapes). Tune the threshold or disable with the new `autoRefreshStatistics` store option (`createStore(graph, backend, { autoRefreshStatistics: 5000 })` or `false`). Bulk writes inside a caller-provided transaction never auto-refresh — statistics cannot see uncommitted rows — and a refresh failure degrades to a warning without failing the committed write. `importGraph()` keeps its existing built-in refresh. - [#195](https://github.com/nicia-ai/typegraph/pull/195) [`e48dfa2`](https://github.com/nicia-ai/typegraph/commit/e48dfa2531148892ca7f5432a3ced6068b464807) Thanks [@pdlug](https://github.com/pdlug)! - `bulkCreate` now batches its round trips end to end instead of degenerating into per-row statements around one multi-row INSERT. - Validation probes: per-row existence checks collapse into one `getNodes` per kind, and per-row uniqueness pre-checks into one `checkUniqueBatch` per (constraint, kind) — the batch validation caches are primed up front, so the per-row checks run against memory. Validation now runs as a synchronous first pass, so a later row's validation error can surface before an earlier row's constraint error (both fail the whole batch). - Side effects: uniqueness entries write through a new `insertUniqueBatch` (multi-row conditional upsert with the same per-entry `UniquenessError` semantics), fulltext sync goes through the existing `upsertFulltextBatch`, and embedding sync through a new `upsertEmbeddingBatch` per (kind, field) — implemented for pgvector, sqlite-vec, and libSQL native vectors via an optional `VectorStrategy.buildUpsertBatch` seam with a per-row fallback for custom strategies. Measured on the write bench (in-memory SQLite, 100-row batches of nodes with searchable + embedding fields): ~1,600 → ~4,100 rows/s (~2.6×). The win compounds on per-statement-networked engines (Turso, D1, Neon), where each eliminated statement is a network round trip. - [#194](https://github.com/nicia-ai/typegraph/pull/194) [`b3668c9`](https://github.com/nicia-ai/typegraph/commit/b3668c96db58127f983695fa6df8f39662ed761b) Thanks [@pdlug](https://github.com/pdlug)! - Default-path performance tuning for SQLite and bulk maintenance verbs. - `createLocalSqliteBackend` now applies connection pragmas at open: `journal_mode=WAL`, `synchronous=NORMAL`, and a 5s `busy_timeout`. On file-backed databases this makes single-operation writes roughly 5× faster than the better-sqlite3 driver defaults (rollback journal, `synchronous=FULL`), because each write no longer pays a full-durability fsync in journal mode. Override individual values via the new `pragmas` option, or pass `pragmas: false` to keep driver defaults. - The SQLite backend now detects the connection's real bound-parameter budget instead of assuming the historic 999: better-sqlite3 compiles in `SQLITE_MAX_VARIABLE_NUMBER=32766` (probed via `PRAGMA compile_options`, with a `sqlite_version() >= 3.32` fallback), Cloudflare D1 is capped at its documented 100, and undetectable async drivers keep the conservative 999 floor. Batch chunk math derives from the detected budget, so bulk inserts on better-sqlite3 use ~33× fewer statements (111-row chunks → 3,640-row chunks), and batched writes on D1 no longer exceed its per-statement limit. `capabilities.maxBindParameters` reports the detected value and remains overridable. - `importGraph()` now refreshes planner statistics (`ANALYZE`) automatically after an import that created or updated rows, and `store.materializeIndexes()` does the same on SQLite after creating indexes. Stale statistics after bulk loads previously degraded traversals ~10× on PostgreSQL and some FTS5 queries ~30× on SQLite until the engine caught up on its own. Both verbs accept `refreshStatistics: false` to opt out. On PostgreSQL, `materializeIndexes()` builds with `CREATE INDEX CONCURRENTLY` and skips the automatic refresh (concurrent same-index builds from two callers can deadlock when a refresh shifts their timing) — call `store.refreshStatistics()` after materializing. - PostgreSQL `refreshStatistics()` now issues one `ANALYZE (SKIP_LOCKED)` per table instead of a single multi-table `ANALYZE`. A multi-table ANALYZE is one transaction acquiring several ShareUpdateExclusive locks in sequence, and ANALYZE's lock class conflicts with in-flight `CREATE INDEX CONCURRENTLY` builds — the old shape could deadlock against concurrent index DDL; the new one can never join a lock-wait cycle (a locked table is skipped and covered by the next refresh or autovacuum). - [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Declare, as a typed capability, whether a backend's filtered approximate vector search can silently return a short page. Every approximate (ANN) search TypeGraph issues carries at least one row filter — the liveness predicate that hides soft-deleted and out-of-validity rows — and a `.where(...)` predicate narrows it further. Where the engine applies that filter relative to the index traversal decides whether the page fills: - **`sqlite-vec`** pushes the filter into the `vec0` KNN candidate set. Exact — the only engine here that guarantees a full page. - **`pgvector` ≥ 0.8** re-enters the index for more candidates (`hnsw.iterative_scan` / `ivfflat.iterative_scan`, applied automatically). Much better recall than a post-filter, but **not** a guarantee: the iterative scan stops at `hnsw.max_scan_tuples` / `ivfflat.max_probes`, and on **pgvector < 0.8** there is no iterative scan at all — the backend detects that at runtime, warns once, and the search stays `ef_search`-bounded. - **`libsql-native`** cannot do either: DiskANN's `vector_top_k` is a table function with no filter pushdown. TypeGraph over-fetches `4 × (limit + offset)` neighbors and post-filters, so once more than that headroom is filtered out the search returns **fewer than `limit` rows even though more matches exist**. Heavy tombstone drift — routine in a temporal store — is what makes this real rather than theoretical. That asymmetry was previously only a code comment. `VectorCapabilities` now carries a required `filteredApproximateSearch: { mode, guaranteesFullPage }`. **Read `guaranteesFullPage`, not `mode`** — `mode` (`"filter-pushdown" | "iterative-scan" | "post-filter"`) names the mechanism the strategy asks for, but only `guaranteesFullPage` reflects the runtime-dependent, scan-bounded reality (it is `true` for `sqlite-vec` alone). It is documented in the backend parity matrix, and boundary tests execute the difference against real libSQL, sqlite-vec, and pgvector: the same 200-vector fixture, the same filter, the same `limit`. **Breaking for custom vector strategies only.** `VectorCapabilities` gained a required field, so a hand-written `VectorStrategy` must now declare both its mode and whether it guarantees a full page. That is deliberate: an omitted declaration would inherit an engine promise the strategy may not keep. - [#198](https://github.com/nicia-ai/typegraph/pull/198) [`a9477bb`](https://github.com/nicia-ai/typegraph/commit/a9477bb28ee887a1a93c103a64912e8563de9d76) Thanks [@pdlug](https://github.com/pdlug)! - Property filters that a btree can never serve now have a declarative index story: `defineNodeIndex` / `defineEdgeIndex` accept `method: "gin" | "trigram"` (default `"btree"`, unchanged). - `method: "gin"` emits a PostgreSQL expression GIN (`jsonb_path_ops`) over the field's jsonb extraction, serving the array containment predicates (`contains` / `containsAll` / `containsAny` on array fields). Verified to match TypeGraph's compiled `(props #> ARRAY[…]) @> $1` form under parameterized prepared statements — note that a hand-written whole-column `GIN (props)` never matches these expressions (the previous docs guidance recommended one; corrected). - `method: "trigram"` emits an expression GIN with `gin_trgm_ops` over the field's text extraction, serving substring and case-insensitive matches (`contains` / `startsWith` / `endsWith` / `like` / `ilike` on string fields). `materializeIndexes()` installs `pg_trgm` (`CREATE EXTENSION IF NOT EXISTS`) on first use. Both are materialize-only (like vector ANN indexes) and PostgreSQL-only: `materializeIndexes()` reports them as `skipped` on SQLite, whose substring-search story is FTS5 fulltext. GIN-family declarations take exactly one field and reject `unique`, `coveringFields`, and `where`; `method: "btree"` is canonicalized by absence so existing stored schema documents and materialization signatures are unchanged. `bulkFindByIndex` rejects GIN-family indexes (it compiles equality probes, which only btree declarations serve). - [#204](https://github.com/nicia-ai/typegraph/pull/204) [`94eea90`](https://github.com/nicia-ai/typegraph/commit/94eea90ead38c69c0ac5b55bad34036f45578b87) Thanks [@pdlug](https://github.com/pdlug)! - perf: `store.search.hybrid` now runs as a single SQL statement on the built-in backends — both sources, weighted RRF fusion, liveness, and node hydration composed into one round trip (previously two search statements plus an id-hydration fetch, with fusion in JS). Results are identical to the previous path; the saving scales with per-statement cost (serverless drivers, D1/Durable Objects, remote databases). `GraphBackend` gains an optional `hybridSearch` member; backends without it (custom backends, capability profiles without window functions) keep the multi-statement fallback. - [#223](https://github.com/nicia-ai/typegraph/pull/223) [`a161d70`](https://github.com/nicia-ai/typegraph/commit/a161d70895da5101706602ac13e4cca4b7fc6a62) Thanks [@pdlug](https://github.com/pdlug)! - Add `asNodeId` and `asEdgeId` constructors for branding persisted ids that round-trip through untyped storage before being passed back to read, update, or delete APIs. - [#241](https://github.com/nicia-ai/typegraph/pull/241) [`8f3e772`](https://github.com/nicia-ai/typegraph/commit/8f3e7727dde0d46415b90e715138c0a9766cd2b5) Thanks [@pdlug](https://github.com/pdlug)! - Fixes `implies(edgeA, edgeB)` silently accepting endpoint-incompatible edge pairs. Previously an ontology declaration like `implies(about, writes)` — where `about` connects `Paper -> Topic` and `writes` connects `Author -> Paper` — was accepted without complaint, and `expand: "implying"` query traversal would then silently fold `about` rows into a `writes` traversal even though the two edges connect entirely different node kinds. `implies()` relations are now validated wherever a query-capable `KindRegistry` is built — `createStore()`/`createStoreWithSchema()` for a live graph definition, and `deserializeSchema(...).buildRegistry()` for a persisted schema — including relations authored through `store.evolve({ ontology })`. A relation is accepted when every kind the implying edge allows on a side (`from`/`to`) is assignable — equal, or a `subClassOf` descendant — to at least one kind the implied edge allows on that same side; otherwise construction throws a `ConfigurationError` describing the incompatible kinds and how to fix the declaration. **Breaking change — two things to know before upgrading.** _It breaks the load path, not just graph definition._ `deserializeSchema(...)` runs the same endpoint check inside `buildRegistry()`, so a schema **already persisted** under 0.34 that carries a now-rejected `implies()` relation throws at the first `buildRegistry()` after the upgrade — no code change of yours required to trigger it. Audit persisted schemas before rolling out, not only the graph definitions in source. _It rejects superset domains, not only disjoint ones._ A relation is accepted only when every kind the implying edge allows on a side is assignable to at least one kind the implied edge allows on that side. So `implies(a, b)` where `a` is declared `from: [Person]` and `b` is declared `from: [Employee]` (with `Employee subClassOf Person`) is **rejected**, even though every `a` row on disk might in fact start at an `Employee`: `Person` is not assignable to `Employee`. The declaration, not the data, is what the traversal folds on, and a `Person`-rooted `a` row folded into a `b` traversal would be unsound. The same rule is what makes the previously-silent disjoint case (`Paper -> Topic` implying `Author -> Paper`) an error. Fix such relations by narrowing the implying edge's endpoints, adding a `subClassOf` relation to bridge the mismatch, or removing the `implies()` declaration. - [#195](https://github.com/nicia-ai/typegraph/pull/195) [`e48dfa2`](https://github.com/nicia-ai/typegraph/commit/e48dfa2531148892ca7f5432a3ced6068b464807) Thanks [@pdlug](https://github.com/pdlug)! - `importGraph` now processes each `batchSize` slice with batched round trips instead of fully single-row statements. Nodes: one `getNodes` per kind for existence, one `checkUniqueBatch` per (constraint, kind) for uniqueness pre-checks, one multi-row insert, and one batched side-effect pass (uniqueness entries, fulltext, embeddings) for the accepted creates. Edges: one `getNodes` per endpoint kind for reference liveness, one `getEdges` for existence, and one multi-row insert. Per-row semantics are unchanged: conflicts route by `onConflict`, a uniqueness conflict is recorded as a per-row error entry (the rest of the import proceeds), reference validation still rejects missing or tombstoned endpoints, and rows repeating an id within a slice fall back to the per-row path so they observe the first occurrence's row exactly as before. Measured on the write bench (in-memory SQLite, 500 nodes + 500 edges per import): ~26k → ~96k entities/s (~4×). The win compounds on per-statement-networked engines (Turso, D1, Neon), where the old path paid one round trip per row and the new one pays a handful per slice. - [#236](https://github.com/nicia-ai/typegraph/pull/236) [`31aee82`](https://github.com/nicia-ai/typegraph/commit/31aee82608518411e4e9f905c96c52348f7cf08f) Thanks [@pdlug](https://github.com/pdlug)! - `defineNodeIndex` accepts a new `keySystemColumns` option: system columns (e.g. `"id"`) to include in the index key, positioned after the `scope` prefix and before `fields`/`coveringFields`. `fields` is now optional (was a required non-empty tuple) — an index must declare at least one of `fields`, `coveringFields`, or `keySystemColumns`. This closes a real gap: a covering index can only serve a query's join index-only (avoiding a heap fetch per candidate row) if the index's key matches the join's actual predicate. Queries that join on a system column directly (e.g. TypeGraph's compiled `n.id = e.from_id` for a reverse traversal) had no way to declare a matching index, since `fields`/ `coveringFields` only ever accept the node's own schema properties. `keySystemColumns: ["id"]` (plus `coveringFields` for whatever the query also projects) now lets that same join be served index-only. Rejects edge-only system columns (`from_kind`/`from_id`/`to_kind`/ `to_id`) on a node index, and rejects any column already implied by `scope`. Not supported with `method: "gin" | "trigram"` (same restriction as `coveringFields`). Also rejects `unique: true` combined with `keySystemColumns: ["id"]` — every node's `id` is already unique per row, so a unique index keyed on `id` plus other columns can never enforce a meaningful constraint across those other columns. Canonicalized by absence, like `method`: indexes that don't use it produce byte-identical names/hashes to before this field existed, so existing stored schema documents and materialization signatures are unaffected. - [#208](https://github.com/nicia-ai/typegraph/pull/208) [`586b2b0`](https://github.com/nicia-ai/typegraph/commit/586b2b05f3f501f3d53db1dbb2ec247e17a67294) Thanks [@pdlug](https://github.com/pdlug)! - fix: `materializeIndexes` serializes same-index builds across callers on PostgreSQL via a durable claim in the status table (two concurrent same-name expression-index `CREATE INDEX CONCURRENTLY` builds can deadlock — no safe-snapshot exemption). Losers wait and converge as `alreadyMaterialized`; a crashed builder's claim expires after a 15-minute lease and the takeover drops the INVALID index leftover before rebuilding (relational indexes now self-heal instead of requiring manual repair). With same-index builds serialized, the automatic post-create `ANALYZE` is re-enabled on PostgreSQL. - [#201](https://github.com/nicia-ai/typegraph/pull/201) [`b52ae3b`](https://github.com/nicia-ai/typegraph/commit/b52ae3b3358435de9774f348fc94ab7140bdc7eb) Thanks [@pdlug](https://github.com/pdlug)! - perf: eliminate the PostgreSQL JSONB parse→stringify→parse round trip per row. **Public backend row contract change:** rows returned by `GraphBackend` read methods now carry `props` as `RowProps = string | Readonly>` — JSON text on SQLite, the driver-parsed object on PostgreSQL. Code that consumed backend rows directly with `JSON.parse(row.props)` must switch to the new `rowPropsToObject(row.props)` (or `rowPropsToJsonText` when text is required); both helpers and the `RowProps` type are exported from the package root. Store-level APIs (`store.nodes.*`, `store.query()`, search, export) are unaffected — they already return parsed objects. - [#249](https://github.com/nicia-ai/typegraph/pull/249) [`d2a6feb`](https://github.com/nicia-ai/typegraph/commit/d2a6feb8a99aaafa247c7bf97f9670c56608a870) Thanks [@pdlug](https://github.com/pdlug)! - Add revision-anchored graph branches and streaming interchange. Stores can opt into `revisionTracking: true` (or use `history: true`) so branch and merge validation read a durable per-graph origin and revision instead of fingerprinting every live row or accepting a coincident revision from another store. Physical branch clones now stream bounded interchange batches, enabling large branch copies, exports, and imports without materializing the full graph in memory. Direct backend writes remain outside the revision-tracking contract; tracked stores fail loudly if `tx.sql` would bypass that contract. - [#203](https://github.com/nicia-ai/typegraph/pull/203) [`801768d`](https://github.com/nicia-ai/typegraph/commit/801768d2e1a63a0d3bda9d40a46a7f03deddffbd) Thanks [@pdlug](https://github.com/pdlug)! - feat: facade search scoping — `store.search.{vector,fulltext,hybrid}` accept `where` (a property predicate compiled by the shared query compiler into the search statement's candidate set), `offset` (rank-relative pagination pushed into the engine), and `includeSubClasses` (search `subClassOf` descendants and merge into one ranking). Filters compile into the search statement's candidate set — exact on pgvector, sqlite-vec, tsvector, and FTS5, where a filtered search returns `limit` hits whenever enough matches exist; libSQL DiskANN post-filters a 4× over-fetched ANN set, so its recall against the filter is bounded by that headroom. Search now applies full current-read semantics (validity windows, not just tombstones), matching `find()`. - [#205](https://github.com/nicia-ai/typegraph/pull/205) [`17bbe54`](https://github.com/nicia-ai/typegraph/commit/17bbe5419a246c95bbab9f6bc7da64f6691e159e) Thanks [@pdlug](https://github.com/pdlug)! - feat: `.similarTo(vector, k, { approximate: true })` — opt-in approximate retrieval for the inline vector predicate. Each declaring kind's relevance branch compiles to the engine's native ANN search form (vec0 `MATCH … k=`, libSQL `vector_top_k`, pgvector's index-eligible scan), scoped to the query's candidate nodes via the same pushdown the search facade uses, so composed predicates and traversals still constrain results. Never applied silently: the default remains the exact distance scan, and slots declared `indexType: "none"` keep it even with the opt-in. - [#245](https://github.com/nicia-ai/typegraph/pull/245) [`ef6def6`](https://github.com/nicia-ai/typegraph/commit/ef6def6b67e306a9cdb40e78723dad6d36f89647) Thanks [@pdlug](https://github.com/pdlug)! - `createLocalSqliteBackend`'s `pragmas` option accepts two new fields: `cacheSizeKib` (`PRAGMA cache_size`) and `mmapSizeBytes` (`PRAGMA mmap_size`). Both default to `undefined`, leaving SQLite's own built-in defaults (a 2MiB page cache, mmap disabled) untouched — existing callers are unaffected. SQLite's 2MiB default cache is fine for a small embedded database, but once a database's working set exceeds it, every page a query touches past that point pays a fresh disk read instead of a cache hit — including pages an otherwise fully covering index would have served from cache alone. Set `cacheSizeKib` (and optionally `mmapSizeBytes`) once a database's working set is known to exceed the default, the same way you'd size a page cache for any other embedded or server database engine. - [#197](https://github.com/nicia-ai/typegraph/pull/197) [`f420a92`](https://github.com/nicia-ai/typegraph/commit/f420a922a1f168891ee4de54e91cc9ca1638deed) Thanks [@pdlug](https://github.com/pdlug)! - SQLite CRUD statements now reuse the prepared-statement cache. The operation backend's read/write helpers previously executed through drizzle's `db.all()` / `db.run()`, which re-prepares every statement on every call — only the query engine's `backend.execute` path used the prepared-statement LRU. On synchronous drivers (better-sqlite3, bun:sqlite) CRUD statements and the per-write transaction frames (`BEGIN IMMEDIATE` / `COMMIT` / `ROLLBACK`) now route through the execution adapter's compiled path, so a repeated operation shape re-binds parameters against a cached prepared statement. A warmed CRUD cycle re-prepares nothing. Async drivers (remote libsql/Turso, D1) have no statement cache and keep the existing execution path. Measured on the write bench (in-memory SQLite, order-controlled A/B): single-op creates ~18.3k → ~28.8k ops/s (~1.6×), transaction-batched creates ~23.9k → ~36k ops/s (~1.5×). - [#251](https://github.com/nicia-ai/typegraph/pull/251) [`f23f7a5`](https://github.com/nicia-ai/typegraph/commit/f23f7a5d15fd5fb59de3667c7d7b10e1975690d4) Thanks [@pdlug](https://github.com/pdlug)! - `createLocalSqliteBackend`'s `pragmas` option accepts a new field: `walAutocheckpointPages` (`PRAGMA wal_autocheckpoint`). Defaults to `undefined`, leaving SQLite's own built-in default (1,000 pages, ~4MiB) untouched — existing callers are unaffected. SQLite's default checkpoints WAL back into the main database file every ~4MiB. That's fine for a normal read/write mix, but a large bulk load pays increasingly expensive checkpoints as the database file grows over the course of the load — each checkpoint has to flush WAL frames into a B-tree that's larger, and less page-cache-resident, than the one before it. A local repro (real `bulkInsert()` calls, 100K/500K/2M synthetic rows) confirmed this: raising `walAutocheckpointPages` cut a 2M-row bulk load's wall-clock time by over 50% at the largest scale tested, with the effect growing at larger row counts. Set `walAutocheckpointPages` for a bulk-insert-heavy workload; `0` disables automatic checkpointing entirely for callers that would rather run one explicit `PRAGMA wal_checkpoint` after the load finishes. - [#222](https://github.com/nicia-ai/typegraph/pull/222) [`7588634`](https://github.com/nicia-ai/typegraph/commit/758863402ec69b3724acb93f07a51eaf23132dc7) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.transactionWithReceipt()`, which runs a transaction and returns a receipt summarizing completed collection write intents and, for history-enabled stores, the recorded commit instant allocated by the transaction. - [#233](https://github.com/nicia-ai/typegraph/pull/233) [`e0e6304`](https://github.com/nicia-ai/typegraph/commit/e0e6304bc17b6d9c004376fd77ddf8cc3b0cc252) Thanks [@pdlug](https://github.com/pdlug)! - Every `TypeGraphError` subclass with a fixed-shape `details` payload now declares a narrowed `readonly details` type (e.g. `RestrictedDeleteError.details` is `RestrictedDeleteErrorDetails`, not the base class's `Readonly>`), so reading structured fields like `error.details.edgeCount` no longer requires a cast. The new `XxxErrorDetails` types (`NodeNotFoundErrorDetails`, `EdgeNotFoundErrorDetails`, `KindNotFoundErrorDetails`, `NodeConstraintNotFoundErrorDetails`, `NodeIndexNotFoundErrorDetails`, `EndpointNotFoundErrorDetails`, `EndpointErrorDetails`, `UniquenessErrorDetails`, `CardinalityErrorDetails`, `DisjointErrorDetails`, `RestrictedDeleteErrorDetails`, `VersionConflictErrorDetails`, `SchemaMismatchErrorDetails`, `MigrationErrorDetails`, `EagerMaterializationErrorDetails`, `StaleVersionErrorDetails`, `SchemaContentConflictErrorDetails`, `StoreNotInitializedErrorDetails`, `DatabaseOperationErrorDetails`, `EmbeddingDimensionChangedErrorDetails`) are exported from the package root alongside the existing `ValidationErrorDetails`. Classes with intentionally open, per-call-site details (`ConfigurationError`, `UnsupportedPredicateError`, `CompilerInvariantError`, `BackendDisposedError`) are unchanged. - [#206](https://github.com/nicia-ai/typegraph/pull/206) [`995b964`](https://github.com/nicia-ai/typegraph/commit/995b9643927f498deabb157113bcb5ecb5883ca9) Thanks [@pdlug](https://github.com/pdlug)! - perf: cascade deletes batch their edge removals — new optional `GraphBackend.deleteEdgesBatch` / `hardDeleteEdgesBatch` members issue one statement per bind-budget chunk instead of one per connected edge (50-edge cascade on local PostgreSQL: 24.4ms → 3.6ms), with recorded-time capture preserved. `getOrCreate` variants no longer run the full Zod parse twice on the create leg. ### Patch Changes - [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: the synthetic CTE column names that carry selectively-extracted `props` fields are now bounded to PostgreSQL's identifier limit. A selected top-level `props` field is extracted once inside the CTE that owns it, under a generated column name encoding the query alias and the field name. The encoding was unambiguous but unbounded, and PostgreSQL silently truncates identifiers at 63 **bytes** — so two distinct `(alias, field)` pairs sharing a long prefix could collapse onto one column name after truncation, yielding an ambiguous-column error or the wrong value. Long names are now truncated on a UTF-8 character boundary and disambiguated with a hash of the full, untruncated pair — the same guard the sibling subgraph projection path already used, now extracted into one shared helper. Names that already fit are emitted unchanged, so compiled SQL for ordinary queries is byte-for-byte what it was. - [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Document a semantic consequence of batched writes: within **one backend batch call**, every row whose timestamp TypeGraph generates shares a single instant, sampled once for that call — not once per row, and not once per bind-budget chunk. `bulkCreate()` and `bulkInsert()` issue one such call, so all of their rows tie. Creating the same rows one at a time through `create()` gives each its own timestamp, so `ORDER BY created_at` was a total order there and is only a partial one after a bulk write. Two things it is **not** safe to conclude: - **`importGraph()` is not one instant.** It slices nodes and edges into `batchSize` batches and drives one backend call per slice, so each slice samples its own timestamp. Rows that carry an explicit `validFrom` in the import payload keep it verbatim; only generated defaults are affected. - **Ids are not a sequence.** The default generator is a random NanoID, and callers may supply arbitrary ids, so `ORDER BY id` is not insertion order. `(created_at, id)` is a _deterministic_ tiebreak, not a chronology. If input order matters, persist an explicit sequence column. One instant per batch call is the intended semantics — it is what makes a bulk write a single point in valid time rather than a smear — and it is the same choice `valid_from` already made. Nothing changes in behavior; this note exists because the batching work that landed this release moved several paths onto it. - [#248](https://github.com/nicia-ai/typegraph/pull/248) [`c379045`](https://github.com/nicia-ai/typegraph/commit/c37904505ee0cf17a9a62f4f7e6769be61319670) Thanks [@pdlug](https://github.com/pdlug)! - Perf: cache compiled query SQL across executions again, without freezing the read instant. The read-freshness fix recompiled a query's full AST to SQL on every `execute()` so a reused or prepared query would always see the latest rows. That kept results fresh but made the recommended `.prepare()`-once-`.execute()`- many pattern pay a full compile per call (a point lookup ~58µs, a three-hop traversal ~450µs of pure JS compilation). Only the bound "current" read instant varies between two compilations of the same query; the SQL text is identical. So a query now compiles once into a cached statement whose read instant is a reserved execution-time placeholder, and each execution fills a fresh instant into it and runs the cached text directly. Repeated point-query execution drops from ~47µs to ~2.4µs (near the raw-execution floor) while staying just as fresh — a row created after `prepare()` or the first `execute()` is still visible on the next call. The cache applies to `ExecutableQuery`, prepared queries, aggregate queries, and set operations, on backends that can compile and run raw SQL text (synchronous SQLite and PostgreSQL backends); other backends — including async SQLite profiles that do not expose `executeRaw` — fall back to per-call recompilation unchanged. Statements whose execution depends on the compiled SQL object — pgvector approximate-scan GUC tuning and parameter-blind-plan avoidance — keep running through the standard execution path. `param()` now rejects the reserved read-instant name, and aggregate queries (which have no `.prepare()`) reject `param()` with clear guidance instead of a downstream binding error. - [#244](https://github.com/nicia-ai/typegraph/pull/244) [`b38a537`](https://github.com/nicia-ai/typegraph/commit/b38a537d1f0abcb5925a94b7e0845fb1184509ff) Thanks [@pdlug](https://github.com/pdlug)! - Fix: "current" temporal reads now evaluate validity against the application clock, not the database clock — repairing a read-after-write consistency violation on Postgres. `valid_from` is stamped from the application clock (`Date.toISOString()`) on write, but a "current" read compiled its validity filter against the database clock (`valid_from <= NOW()` on Postgres). On any deployment where the application-server clock runs ahead of the database-server clock — i.e. the app and database on separate hosts, which is the norm — a freshly-created node or edge could be missing from the very "current" read that immediately followed its creation, until the database clock caught up. SQLite (a single in-process clock) was never exposed. The "current" read now binds the application clock (`nowIso()`) as a parameter — the same clock `valid_from`, the facade search-currency filter, and the recorded/logical clock already use — across every current-read path (standard and recursive queries, subgraph extraction, graph algorithms, and recorded-time reads). The temporal-visibility clock is now a single source. Because the current-read instant is no longer dialect-specific, the internal `DialectAdapter.currentTimestamp()` seam has been removed. **Know the consistency model this buys you.** Reads and writes now share one clock — _the clock of the process that issued them_. Read-after-write consistency therefore holds **per application process**: a node you just created is visible to the very next current read from that same process, which is the guarantee the bug broke. It does **not** extend across processes. Two application servers with skewed clocks, writing to one PostgreSQL database, can still miss each other's fresh rows: a row stamped `valid_from = T` by the server that runs ahead stays invisible to a current read from the server that runs behind until its own clock passes `T`. The window equals the skew between the two application hosts, not between an application host and the database. If you need cross-process read-after-write consistency, keep application clocks disciplined (NTP), or read at an explicit `asOf` coordinate rather than `current`. - [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: `store.algorithms.degree()` undercounted edges written before an endpoint declaration changed. To let the composite edge indexes seek — both lead with the endpoint kind column, so a bare `from_id = ?` cannot — the direction filter supplied the missing kind equality by enumerating the endpoint kinds the _graph declaration_ permits for the counted edge kinds. That enumeration is complete only for rows written under the current declaration. Narrow `knows` from `from: [Person]` to `from: [Employee]`, and every `Person`-rooted `knows` edge already on disk drops out of the filter: `degree()` silently returns a number too small, with no error and no warning. The filter now derives the kind from the counted node itself, via an uncorrelated scalar subquery. This is exact by construction: an edge row stores the _actual_ kind of each endpoint node (the write path copies it off the endpoint reference) and a node's kind is immutable for the life of its id, so for any edge incident to a node, the endpoint kind on that node's side is that node's kind and nothing else — however the declaration later evolves. It is also a better filter. An equality on one kind replaces an `IN` list over every declared endpoint and its `subClassOf` descendants, and both engines hoist the uncorrelated subquery to a constant (a Postgres InitPlan, a SQLite one-shot scalar subquery), so the seek is unchanged. `EXPLAIN QUERY PLAN` still shows `typegraph_edges_from_idx` / `_to_idx` seeks with no partition scan. `degree()` of an id that names no node is `0`, as before. - [#200](https://github.com/nicia-ai/typegraph/pull/200) [`472ac1c`](https://github.com/nicia-ai/typegraph/commit/472ac1c20a6751a52121da8732f6c562fe5124c8) Thanks [@pdlug](https://github.com/pdlug)! - `degree()` direction filters are now shaped for the default edge indexes. The filters previously compiled to bare `from_id = ?` / `to_id = ?`, which neither composite edge index can seek (both lead with the endpoint kind column) — so degree counts relied on engine-specific rescue: SQLite skip-scan (only with fresh statistics) or PostgreSQL 18's new btree skip scan, and degenerated to partition scans everywhere else (PostgreSQL ≤ 17, SQLite with stale statistics). The filters now enumerate the endpoint kinds the graph declaration permits for the counted edge kinds, expanded through the subClassOf closure — the same set edge writes validate against — making `edges_from_idx` / `edges_to_idx` structurally seekable on every engine and version. Measured on PostgreSQL 18 (where the old form was already skip-scan rescued): 0.30ms → 0.06ms per call; on older PostgreSQL the old form could not use these indexes at all. An edge set that declares no endpoint kinds on the required side now returns 0 without a round trip. Behavior note: because the counted set is now restricted to edges whose stored endpoint kind falls within the declaration's `subClassOf` closure, `degree()` no longer counts an edge whose stored `from_kind` / `to_kind` lies _outside_ that closure — e.g. a row written before the endpoint declaration was narrowed, or written directly through the backend bypassing endpoint validation. This matches how typed traversals already treat such rows (invisible to a schema-consistent read), but it is a change from the previous "count every edge touching this node regardless of stored kind" behavior. - [#220](https://github.com/nicia-ai/typegraph/pull/220) [`7b48543`](https://github.com/nicia-ai/typegraph/commit/7b4854310fc042410e31f2e14abc19a9e61e44a2) Thanks [@pdlug](https://github.com/pdlug)! - Edge delete, edge hard delete, and node hard delete no longer re-read the row inside the write transaction. The in-transaction preflight was pure round-trip fat on these paths: nothing consumed the row, and the writes are already concurrency-correct on their own — the tombstone UPDATE is guarded by `deleted_at IS NULL` and the hard deletes are id-keyed and idempotent, so a row deleted concurrently between the outside gate and the write lock degrades to a 0-row no-op with identical observable behavior (verified including recorded-time history under a deliberately staled gate). One less statement per delete (~20% of the per-op round trips on client/server engines). Node SOFT delete keeps its preflight deliberately: its pipeline consumes the pre-image for uniqueness-key cleanup, now documented in place. - [#227](https://github.com/nicia-ai/typegraph/pull/227) [`09754a6`](https://github.com/nicia-ai/typegraph/commit/09754a6e4435425e8a55e9a0b991fcbd66daccbf) Thanks [@pdlug](https://github.com/pdlug)! - Batches edge creation's endpoint-existence checks in `bulkCreate`/`bulkInsert` into one `getNodes` call per distinct (kind) referenced across the whole batch, instead of an individual `getNode` probe per edge (mirroring the batched existence/uniqueness pre-check node creation already had via `primeBatchValidationCaches`). Found while investigating why a real LDBC SNB SF1 bulk load (millions of nodes and edges) was far slower than expected: a controlled 1M-row reproduction showed `bulkInsert` edge-batch time growing from ~90ms to ~630ms per 2,000-row batch as the graph grew, while an equivalent node-only batch (no edges) stayed roughly flat. The edge batch path validated each edge's `from`/`to` endpoints with a `getNode` call per edge — for a batch with mostly-unique endpoints, that's thousands of individual round trips per batch instead of one batched fetch per distinct node kind. With the fix, the same 1M-edge reproduction's per-batch time drops to roughly ~90-160ms and its growth curve flattens substantially (the residual growth matches the same mild index-maintenance cost already seen on plain node inserts). No behavior change: this is a pure internal optimization to `executeEdgeCreateNoReturnBatch`/ `executeEdgeCreateBatch`; callers observe identical results, just fewer round trips. - [#245](https://github.com/nicia-ai/typegraph/pull/245) [`ef6def6`](https://github.com/nicia-ai/typegraph/commit/ef6def6b67e306a9cdb40e78723dad6d36f89647) Thanks [@pdlug](https://github.com/pdlug)! - The default edge traversal indexes (`{table}_from_idx` / `{table}_to_idx`, created for every graph on both SQLite and PostgreSQL) were missing two things a traversal join needs to be served fully index-only: - **`valid_from`** — one of the three system columns every compiled query's soft-delete / temporal-validity predicate checks (`deleted_at` and `valid_to` were already covered; `valid_from` wasn't). - **The join's target-id column** — a compiled traversal reads `n.id = e.to_id` for an outgoing traversal, or `n.id = e.from_id` for an incoming one (`standard-builders.ts`), but neither index carried the _other_ endpoint's id column, so the join to the target node still required a heap-row fetch even once the predicate columns above were covered. Both gaps produce the same symptom: SQLite's plan reads `USING INDEX`, never `USING COVERING INDEX`, so every candidate edge pays a heap-row fetch. That fetch is free while the table fits in the page cache. Once it doesn't — a real LDBC SNB benchmark run measured this at 10x data volume, where the nodes table outgrew available cache — every one of those fetches becomes a genuine random disk read, and with thousands of candidates per traversal that alone produced a multi-second/minute latency cliff on an otherwise sub-millisecond query shape. Both indexes now carry all five columns beyond their existing seek prefix (`deleted_at`, `valid_from`, `valid_to`, plus the other endpoint's id), confirmed via `EXPLAIN QUERY PLAN` against the actual SQL `execute()` sends (not `toSQL()`'s wider, unoptimized output) to flip to `USING COVERING INDEX`. **Existing databases get none of this until you rebuild the indexes.** The widened indexes materialize on **fresh databases only**. `generateSqliteMigrationSQL()` / `generatePostgresMigrationSQL()` emit `CREATE INDEX IF NOT EXISTS` under the _same index name_, and that is a no-op against an index that already exists — regardless of how the column list changed. An upgraded deployment silently keeps its narrow index, and keeps the latency cliff, until it runs the rebuild below. Upgrading the package is not enough; there is no automatic migration. ```sql -- SQLite: no CONCURRENTLY equivalent; drop and let the next migration -- run (generateSqliteMigrationSQL(), or a createStoreWithSchema boot, -- which re-issues idempotent DDL) recreate them. DROP INDEX IF EXISTS typegraph_edges_from_idx; DROP INDEX IF EXISTS typegraph_edges_to_idx; -- PostgreSQL: CREATE INDEX CONCURRENTLY does not block writes, but it -- cannot run inside a transaction and needs its own connection. Rename -- the old index out of the way first so the new one can use the -- production name without a window where neither exists. ALTER INDEX typegraph_edges_from_idx RENAME TO typegraph_edges_from_idx_old; CREATE INDEX CONCURRENTLY "typegraph_edges_from_idx" ON "typegraph_edges" ("graph_id", "from_kind", "from_id", "kind", "to_kind", "deleted_at", "valid_from", "valid_to", "to_id"); DROP INDEX CONCURRENTLY typegraph_edges_from_idx_old; ALTER INDEX typegraph_edges_to_idx RENAME TO typegraph_edges_to_idx_old; CREATE INDEX CONCURRENTLY "typegraph_edges_to_idx" ON "typegraph_edges" ("graph_id", "to_kind", "to_id", "kind", "from_kind", "deleted_at", "valid_from", "valid_to", "from_id"); DROP INDEX CONCURRENTLY typegraph_edges_to_idx_old; ``` - [#217](https://github.com/nicia-ai/typegraph/pull/217) [`fce0a0f`](https://github.com/nicia-ai/typegraph/commit/fce0a0f18b90e7b6f5b5d395681231865b21fb52) Thanks [@pdlug](https://github.com/pdlug)! - Non-approximate `.similarTo()` is now genuinely exact when an ANN index exists. pgvector serves any `ORDER BY embedding <=> q LIMIT k` from a matching HNSW/IVFFlat index, so after `materializeIndexes()` the default (non-approximate) inline vector predicate silently returned approximate results — measured recall 0.980 unfiltered and 0.000 under a selective filter at 50k docs, where the index frontier starves at the default ef_search and returns entirely wrong rows. The exact branch now orders by `(distance + 0.0)`, which the index opclass cannot match, forcing the true flat scan on every engine (numerically identity; inert on SQLite/libSQL whose ANN forms are opt-in constructs). Behavior change: exact queries that were silently index-served get correct results and flat-scan latency (50k x 384 dims: ~39ms instead of ~23ms-but-wrong). The sanctioned fast path remains `similarTo(..., { approximate: true })`, which is unchanged. The `bench:vector` lane's `vector:exact-postindex-recall` and `vector:exact-filtered-postindex-recall` rows now read 1.000. - [#210](https://github.com/nicia-ai/typegraph/pull/210) [`76422c6`](https://github.com/nicia-ai/typegraph/commit/76422c64189baa2c83287a99a1fea6a13bbfe976) Thanks [@pdlug](https://github.com/pdlug)! - perf: PostgreSQL fulltext queries are now parsed with the kind's DECLARED language as a plan-time constant (the same winning-language rule the write path applies to rows), instead of referencing the per-row `language` column. The per-row form made every tsquery non-constant, so the GIN index on `tsv` could never serve a match and every search re-parsed the query per row — measured 12.9ms → 2.3ms at 5,000 docs for the parse elimination alone, with GIN service now possible as corpora grow. Applies to the facade and the inline `$fulltext` predicate; mixed-language subclass aliases and explicit per-query overrides behave as before. - [#207](https://github.com/nicia-ai/typegraph/pull/207) [`5cbcb35`](https://github.com/nicia-ai/typegraph/commit/5cbcb35f9df972a6f36975b43adad2d7b110bfd1) Thanks [@pdlug](https://github.com/pdlug)! - perf: recorded-time capture acquires the PostgreSQL graph-write advisory lock once per transaction instead of once per captured write (`pg_advisory_xact_lock` is reentrant and held to transaction end, so the repeats were pure round trips). A 50-write recorded transaction drops from N+1 lock round trips to 1; measured 1.7× on the transaction shape. - [#215](https://github.com/nicia-ai/typegraph/pull/215) [`0eb2fd8`](https://github.com/nicia-ai/typegraph/commit/0eb2fd8ba778fb6e9cf6469481805a1c8cd86647) Thanks [@pdlug](https://github.com/pdlug)! - The single-statement hybrid search now emits the candidates set (liveness/currency filter, or the compiled `where` predicate query) once, as a CTE shared by the vector and fulltext legs, instead of embedding — and re-executing — a private copy inside each leg. The duplicate evaluation was most expensive with a `where` filter, whose compiled candidates query ran twice per search: measured on PostgreSQL, filtered hybrid drops 26.5ms → 17.1ms at 5k docs (bench shape 11.8ms → 8.6ms; unfiltered 6.1ms → 4.9ms). This also removes a subtle inconsistency where each leg stamped its own currency instant. SQLite is unchanged within noise (in-process re-execution was cheap). - [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: hybrid search's two execution paths agreed on scores but not on ties, and neither was deterministic across PostgreSQL databases. Relevance ranking breaks a score tie on `node_id`. Left bare, PostgreSQL sorts that under the database's default text collation: an `en_US.UTF-8` database orders `a, A, b, B` where byte order gives `A, B, a, b`. So the same query returned different pages on two databases whose `datcollate` differed, and disagreed with SQLite (whose `BINARY` collation is byte order) throughout. Three seams had to move together, because a hybrid search's tiebreak decides the page twice — once in the per-source ranks, and again in the fused ordering the ranks produce: - The single-statement hybrid search now renders `node_id COLLATE "C"` in both per-source `ROW_NUMBER()` windows and in the final `ORDER BY`. - The standalone fulltext search's `ORDER BY … , node_id` is C-collated too, so the multi-statement fallback's fulltext ranks match. - The fallback now re-ranks each leg's rows before assigning ranks, rather than trusting the order the source SQL happened to return for a single kind. The vector source breaks a distance tie arbitrarily — it carries no `node_id` tiebreak, because a second sort key would cost pgvector its ordered index scan — so its arrival order was never a sound basis for a rank. That re-rank sorts with a new code-point comparator rather than JavaScript's UTF-16 code-unit `<`, which disagrees with byte order for astral characters such as emoji. All three orderings now coincide, and the single-statement and multi-statement paths return identical hits, ranks, and scores even when every score ties. Results only change where they were previously non-deterministic. - [#213](https://github.com/nicia-ai/typegraph/pull/213) [`a243f3b`](https://github.com/nicia-ai/typegraph/commit/a243f3bc323f8d7377454f06c2349fb87386963c) Thanks [@pdlug](https://github.com/pdlug)! - `importGraph`'s default `batchSize` is now 1,000 (was 100), and the default now actually applies: options are parsed through `ImportOptionsSchema` at the function boundary, so direct calls that omit fields with schema defaults (e.g. `{ onConflict: "error" }`) resolve them instead of reading `undefined`. `ImportOptions` is now the schema's input type — fields with defaults are optional for callers. Each import batch pays fixed per-round-trip costs (existence probe, unique pre-check, one multi-row insert), so the old default dominated import time on client/server engines: a 20k-node + 5k-edge import on PostgreSQL drops from 1,515ms to 781ms (16.5k → 32k entities/s). SQLite imports are insensitive to the value (in-process, no round trips). Explicit `batchSize` values are unaffected. Fulltext batch upserts and deletes are now split by the driver's bind-parameter budget in the backend wrappers, like node/edge/unique inserts already were. Previously a searchable import slice emitted ONE FTS5 (or tsvector) statement over every row — 6 binds per row, so a 1,000-row slice overflowed SQLite's 999-bind fallback ceiling and D1's ~100-bind cap, and 6,000-row slices overflowed even better-sqlite3's 32,766 budget ("too many SQL variables"). - [#221](https://github.com/nicia-ai/typegraph/pull/221) [`9b61809`](https://github.com/nicia-ai/typegraph/commit/9b618098b6c6f4917f79a23f4b1f0477428de0b3) Thanks [@pdlug](https://github.com/pdlug)! - Inline `.similarTo(..., { approximate: true })` now actually uses the ANN index on PostgreSQL. Two defects compounded: the candidates membership subquery carried a `DISTINCT` that kept the planner off the ordered index scan entirely (even `enable_seqscan = off` could not rescue it — duplicates are irrelevant to `IN` membership, so the DISTINCT bought nothing), and the inline path never applied the pgvector GUCs the search facade uses, so even an index-served filtered scan would have starved at the default ef_search frontier. The compiler now emits duplicate-tolerant membership candidates for the engine-form branch and brands ANN-bearing statements; the PostgreSQL backend wraps branded statements with the facade's GUC overrides (`hnsw.iterative_scan = strict_order` / `ivfflat.iterative_scan = relaxed_order` on transaction-capable drivers with pgvector >= 0.8; the settings are transaction-scoped, so non-transactional backends such as neon-http keep the plain bounded scan). Set operations merge operand brands onto the combined statement, so a union with an approximate operand is wrapped too. Measured at 50k x 384 dims: unfiltered approximate 174ms -> 2.1ms (recall 0.995), filtered approximate 3.8ms at recall 1.000 on filter-independent corpora. The JOIN consumers of the scoped candidates (exact branch, fulltext CTE) keep their DISTINCT — a join does multiply rows on duplicates — and the non-approximate path's exactness guarantee is untouched. - [#224](https://github.com/nicia-ai/typegraph/pull/224) [`b5886cd`](https://github.com/nicia-ai/typegraph/commit/b5886cdad183dcba80586344935278a79f9ed795) Thanks [@pdlug](https://github.com/pdlug)! - Document external event-log materialization patterns and verify the export/import bulk-copy path into graph-merge branches. - [#199](https://github.com/nicia-ai/typegraph/pull/199) [`d01d6c7`](https://github.com/nicia-ai/typegraph/commit/d01d6c76be56efb393a3cd5506e6a5690995c409) Thanks [@pdlug](https://github.com/pdlug)! - Subgraph extraction is ~4× faster on PostgreSQL. The final node/edge fetches filtered ids with `IN (SELECT id FROM included_ids)`; PostgreSQL pulls that form up into a join whose recursive-CTE row estimate (~10 rows for a single-row seed) drives the planner into a nested-loop join filter — measured at ~10 million discarded rows on the depth-3 benchmark shape. PostgreSQL now evaluates membership against the materialized closure ids with a parameterized `text[]` semi-join (`EXISTS (SELECT 1 FROM unnest($ids) AS t(id) WHERE t.id = column)`) rather than pulling the recursive CTE into that join; SQLite keeps `IN (subquery)`, which it already evaluates optimally. Measured (benchmark suite, 1,200 users / depth-3 stress shape): PostgreSQL subgraph full hydration 322ms → 82ms, depth-2 11.5ms → 7.1ms; SQLite unchanged. - [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: serialize the statements TypeGraph issues on a transaction's pinned Postgres connection, so its own graph writes never present two queries to one connection at once. A transaction pins one connection, and the PostgreSQL wire protocol carries one statement at a time. node-postgres hid that behind an internal queue, deprecated it in `pg@8.22` ("Calling client.query() when the client is already executing a query is deprecated and will be removed in pg@9.0. Use async/await or an external async flow control mechanism instead"), and removes the queue in `pg@9`. TypeGraph overlapped statements on a pinned connection in two ways: - **Always on, no user concurrency required.** The node write pipeline issues `Promise.all([syncEmbeddings, syncFulltext])` for any schema that has both a `searchable()` field and an `embedding()` field, so every single `create()`, `update()`, or resurrect on such a schema put two statements on the wire. - **User-driven.** `store.transaction(async (tx) => { await Promise.all([...]) })` is a documented, recommended pattern. Transaction-scoped backends now run every statement they issue through a per-connection queue. Concurrency at the API surface is unchanged — a `Promise.all` of graph writes still works, and on a pooled (non-transactional) backend the statements still run genuinely concurrently. The queue serializes only what already had to be serial. A multi-statement `SET LOCAL`-scoped vector search (snapshot / set / select / restore) runs as one exclusive group, so two concurrent searches can no longer interleave and apply each other's `efSearch`. The transaction boundary also **drains and closes** the queue before the driver emits `COMMIT` / `ROLLBACK`. Those control statements do not travel through the queue, so without the drain a rollback could overlap a live statement. And a callback that rejects out of a `Promise.all` leaves its siblings running: their statements would otherwise land on the connection _after_ the pool had reclaimed it, executing inside an unrelated transaction. Such a statement is now refused with a new `TransactionClosedError` (normally invisible — `Promise.all` has already rejected with the original failure and discards this one). **Scope: the queue mediates only TypeGraph's own statements.** The raw Drizzle handle exposed as `tx.sql` (for writing your own relational tables in the same atomic boundary) bypasses it. Running a raw statement concurrently with a graph write — or with another raw statement — still races on the one pinned connection, and `drainAndClose` cannot wait for a raw statement it never saw. Await each `tx.sql` statement before the next write; this is inherent to a single-connection transaction, not something TypeGraph can enforce over a handle it doesn't mediate. `adoptTransaction()` likewise serializes the statements it issues but never closes the queue — the caller owns that transaction's end. - [#219](https://github.com/nicia-ai/typegraph/pull/219) [`ee93b77`](https://github.com/nicia-ai/typegraph/commit/ee93b77581e6bcbddccf5256dbb2b321b827e361) Thanks [@pdlug](https://github.com/pdlug)! - Statements whose good plan depends on their parameter values (the subgraph id-array fetches, marked internally with the custom-plan brand) now opt out of statement preparation per call on the postgres-js driver too, via `sql.unsafe(text, params, { prepare: false })`. Previously postgres-js prepared them like everything else, so after five executions PostgreSQL flipped them to a generic, parameter-blind plan — the same cliff fixed for node-postgres in the subgraph shared-traversal change (measured there: 21ms → 310ms on the edge fetch). Scalar-parameter statements keep the driver's prepared default. - [#246](https://github.com/nicia-ai/typegraph/pull/246) [`d5aafe8`](https://github.com/nicia-ai/typegraph/commit/d5aafe845f95a503070ac485994afb46b3a82cac) Thanks [@pdlug](https://github.com/pdlug)! - **Critical fix**: `.prepare()`d queries, and any `ExecutableQuery`/`UnionableQuery`/`ExecutableAggregateQuery` instance whose `.execute()` was called more than once, could silently miss rows created after the query was first compiled. A "current" (live) temporal-validity read binds its read instant (`currentReadInstant()`) at SQL compile time. All four query-builder classes cached their compiled SQL text across calls — `.prepare()` compiled once and every subsequent `execute({...})` reused that same SQL text, and a reused `ExecutableQuery`/`UnionableQuery`/`ExecutableAggregateQuery` instance cached its first `.execute()`'s compilation the same way. Both patterns froze "now" at the moment of first compilation: any row created afterward had a `valid_from` later than the frozen instant, so `valid_from <= now` silently evaluated to false for it, for the query's entire remaining lifetime. This is a regression introduced by the `current-read-app-clock` fix (the [#242](https://github.com/nicia-ai/typegraph/issues/242) clock-skew correction): the prior behavior (`NOW()` / `strftime('now')`, evaluated fresh by the database on every execution) did not have this problem. It is more severe than [#242](https://github.com/nicia-ai/typegraph/issues/242) — that bug required app/DB clock skew across separate hosts; this one reproduces unconditionally, in a single process, on the very next insert after a query is prepared or first executed. `.prepare()`-once-`.execute()`-many is this library's own documented, recommended pattern, so this affected the common case, not an edge case. **Fix**: none of the four classes cache compiled SQL text across calls anymore — each `execute()`/`compile()`/`toSQL()` call recompiles fresh, so `currentReadInstant()` is re-evaluated every time. `.prepare()` still builds and structurally validates the query AST once (so a malformed query still fails fast, before the first `execute()`); only the SQL-text compilation moved from prepare-time to each execute-time call. `param()`-bound values are unaffected — those were already correctly re-bound per call. - [#209](https://github.com/nicia-ai/typegraph/pull/209) [`5e24882`](https://github.com/nicia-ai/typegraph/commit/5e24882536a242d75a2ec9973bfb0301027da92c) Thanks [@pdlug](https://github.com/pdlug)! - perf: facade search candidate handling planned poorly at scale. The hybrid statement's fused CTE is now MATERIALIZED (PostgreSQL inlines single-use CTEs, re-executing the fusion subtree once per candidate node row under a nested-loop join), and unfiltered facade searches use a flat, parameter-bound current-read candidates subquery instead of a compiled builder query whose per-row SQL clock calls dominated on SQLite. Semantics are unchanged — validity windows and tombstones are still enforced, with the instant bound as a parameter. Only searches with a `where` predicate compile a builder query as candidates; `includeSubClasses` expands at the store level and each concrete kind uses the flat form. - [#202](https://github.com/nicia-ai/typegraph/pull/202) [`b45cfc3`](https://github.com/nicia-ai/typegraph/commit/b45cfc3e6d141a6f037544572f862f00c27d5571) Thanks [@pdlug](https://github.com/pdlug)! - fix: facade search (`store.search.vector` / `fulltext` / `hybrid`) now computes top-k over live nodes in SQL. Previously the search statement ranked side-table rows alone and hydration dropped tombstoned ids afterward, silently returning fewer than `limit` hits under index drift. Liveness is pushed into the KNN/MATCH SQL on every engine — exact on pgvector ≥0.8 (HNSW via `hnsw.iterative_scan = strict_order`; IVFFlat via `ivfflat.iterative_scan = relaxed_order` with an in-statement re-sort), sqlite-vec (vec0 primary-key `IN` pushdown), tsvector, and FTS5; libSQL DiskANN over-fetches 4× and post-filters (documented recall bound). - [#237](https://github.com/nicia-ai/typegraph/pull/237) [`48f324b`](https://github.com/nicia-ai/typegraph/commit/48f324b905c9d0e2aa52371780e3c443b596040a) Thanks [@pdlug](https://github.com/pdlug)! - Fixes `.select()` query projections losing the `NodeId` brand on node `id` fields. Previously `ctx.alias.id` in a `.select()` callback was typed as plain `string`, so feeding a projected node id back into `getById`/`getByIds` required an unsafe cast (`as never` or worse). `SelectableNode.id` is now typed `NodeId`, matching what `getById`/`getByIds` already require — no runtime change, no cast needed. Edge ids from `.select()` stay plain `string` on purpose: `traverse()` defaults to `expand: "inverse"`, which can back an edge alias with a row of the registered _inverse_ edge kind, so the alias's static edge type doesn't reliably describe the row. Use `asEdgeId` to re-brand a projected edge id before a point read. - [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: a set operation now binds one "current" read instant across all of its operands. `UNION` / `INTERSECT` / `EXCEPT` compile each operand independently, and each operand compiled its own temporal-validity filter from a fresh `nowIso()` sample. A compound `SELECT` is evaluated against a single snapshot, so two samples microseconds apart let the two halves of an `INTERSECT` or `EXCEPT` disagree about whether a row created between them is current — a row could satisfy the left operand's `valid_from <= now` and not the right's. Compilation of a set operation (including nested ones) now runs under a single pinned instant. Ordinary single-leaf queries were already consistent — they bind one instant per compile — and are unaffected. - [#226](https://github.com/nicia-ai/typegraph/pull/226) [`4cd6b4c`](https://github.com/nicia-ai/typegraph/commit/4cd6b4ca8275c2dad53d85c085347814528b3074) Thanks [@pdlug](https://github.com/pdlug)! - Fixes a scaling bug in the SQLite backend's `refreshStatistics()` (the planner-statistics refresh `bulkCreate`/`bulkInsert` trigger automatically after a large autocommit write — see the `autoRefreshStatistics` store option). It ran a bare, unscoped `ANALYZE`, which does two things wrong on SQLite: it re-analyzes every table in the database file (not just TypeGraph's own tables — already fixed on the Postgres backend), and it does a full, unbounded table/index scan per call (Postgres's `ANALYZE` samples a fixed-size set of rows regardless of table size; SQLite's does not unless bounded). A caller streaming a bulk load through repeated `bulkInsert()` calls — the only practical way to load a multi-million-row dataset without holding it all in memory — re-triggers this once each batch's row count crosses the threshold; with unbounded per-call cost growing with total table size, total load time integrated to O(n²) instead of O(n) (observed: a 2M-row bulk load that never finished after 4.5+ hours). `refreshStatistics()` on SQLite now scopes ANALYZE to TypeGraph's own tables and sets `PRAGMA analysis_limit` first, bounding each call's cost the way Postgres's already was. A 100k-row reproduction of the original shape now completes in ~8s with load time growing log-ishly with table size (2x from first batch to last), not quadratically. - [#218](https://github.com/nicia-ai/typegraph/pull/218) [`b601484`](https://github.com/nicia-ai/typegraph/commit/b601484e95f11f61d4b086f493a95e2b0c4f9c18) Thanks [@pdlug](https://github.com/pdlug)! - Non-approximate `.similarTo()` on SQLite now routes through sqlite-vec's vec0 KNN form. vec0's KNN is brute-force in C — exact by construction — so the default path keeps identical results (pinned against JS-computed ground truth) while dropping from the SQL distance scan to engine speed: 489ms → 124ms for top-10 over 50k 384-dim embeddings. Declared via a new `searchIsExact` flag on the vector-strategy contract; pgvector and libSQL leave it unset (their engine forms are approximate) and are unchanged. The metric gate still applies: an explicit metric override that differs from the slot's declared metric falls back to the SQL scan, which is correct for any metric. - [#211](https://github.com/nicia-ai/typegraph/pull/211) [`a216569`](https://github.com/nicia-ai/typegraph/commit/a21656906eec3cfc532200b1709d6356e6047d71) Thanks [@pdlug](https://github.com/pdlug)! - Subgraph extraction on PostgreSQL now runs the recursive traversal once instead of twice. The node and edge fetches previously each embedded the full recursive CTE; the closure ids are now fetched in one statement and passed to both fetches as a single `text[]` parameter, filtered via an `EXISTS` semi-join over `unnest`. Those id-filtered fetches execute as unnamed statements so PostgreSQL plans them against the actual array on every call — a named prepared statement flips to a generic plan after five executions, which mis-plans array-cardinality-dependent filters (measured 21ms → 310ms on the edge fetch). Depth-3 stress subgraph (1,109 nodes / 4,513 edges, wide payloads): 82.9ms → 30.9ms full hydration, 72.3ms → 15.6ms with SQL projection. SQLite keeps its existing single-statement-per-fetch form, which is already optimal for an in-process engine. - [#234](https://github.com/nicia-ai/typegraph/pull/234) [`d042a30`](https://github.com/nicia-ai/typegraph/commit/d042a304979ea32f5777480b2cd28a8a02b1f339) Thanks [@pdlug](https://github.com/pdlug)! - perf: push selected top-level `props` field extractions into the start/traversal CTEs instead of carrying the whole raw `props` JSONB/JSON column outward for later extraction at the final projection. Each selected field is extracted once, inline, as its own typed CTE column (named from a length-prefixed encoding of its alias and field, so distinct alias/field pairs can never collide on the same column name); the outer projection and any matching `ORDER BY` on the same field just reference that column directly instead of re-extracting from a carried-forward `_props` column. Found while investigating why a covering index on a system column (see `keySystemColumns`) still couldn't get Postgres to serve an indexed join index-only: the compiled query was asking for the entire `props` column in the join step even though the final `.select()` only needed one extracted field, so the specific indexed expression was never actually what got read from the table. No behavior change: compiled query results are identical; this only changes which columns each CTE carries and where field extraction happens. - [#242](https://github.com/nicia-ai/typegraph/pull/242) [`6b884b6`](https://github.com/nicia-ai/typegraph/commit/6b884b66b3f642bfc2a65064f51c63ce317c4cc9) Thanks [@pdlug](https://github.com/pdlug)! - Fix: creating a node or edge without an explicit `validFrom` now stamps the operation's own creation timestamp instead of storing SQL `NULL`. `NULL` is interpreted by temporal filters as open-left validity ("valid since forever"), so a record created without `validFrom` was visible at _any_ historical `asOf` instant — including ones before the record existed. This contradicted the documented contract ("omitted `validFrom` defaults to now") and is fixed at the insert layer for every write path: `create`, `createFromRecord`, `upsertById`/`upsertByIdFromRecord` (create branch), `bulkCreate`, `bulkInsert`, `bulkUpsertById`, and get-or-create, for both nodes and edges. `branch()`'s working-copy clone now also exports with `includeTemporal: true`, so a fork's `validFrom`/`validTo` exactly match the base's — without this, the clone would re-stamp any implicit `validFrom` to the fork's own (later) creation time, narrowing the fork's valid-time window relative to the base it was cloned from. This includes rows that still have a `NULL` `valid_from` (predating this fix, or written directly via the backend): `exportGraph`/`importGraph` now round-trip a confirmed open-left window as an explicit `null` rather than silently dropping it, so a legacy row's "valid since forever" semantics survive a clone unchanged instead of being narrowed to the clone's own creation time. `exportGraph`/`importGraph` round trips still default `includeTemporal` to `false`; without it, imported records get a fresh `validFrom` at import time rather than the source's original value (see the Interchange docs). Custom `GraphBackend` implementations that build their own inserts (rather than reusing the bundled Drizzle operation builders) should apply the same rule: an omitted `validFrom` defaults to the row's creation instant, and an explicit `null` is preserved as SQL `NULL` (open-left). - [#214](https://github.com/nicia-ai/typegraph/pull/214) [`583fbb3`](https://github.com/nicia-ai/typegraph/commit/583fbb3782d78b16e07f92082da37ab299c3d966) Thanks [@pdlug](https://github.com/pdlug)! - PostgreSQL ANN index builds (`materializeIndexes()` on pgvector HNSW/IVFFlat) now retry serially when the parallel build exhausts shared memory. Parallel builds stage the index graph in dynamic shared memory, and resource-constrained hosts — e.g. containers with the 64MB `/dev/shm` default — reject the allocation with SQLSTATE class 53 (observed: 53100 from `dsm_impl_posix` on a 50k x 384-dim HNSW build). The retry drops the INVALID leftover from the failed CONCURRENTLY build, pins the vector table to `parallel_workers = 0`, rebuilds in local memory, and restores the setting. Non-resource failures still surface as before. Serial builds are slower — raise `/dev/shm` and `maintenance_work_mem` where you control the host — but a slow index beats a silently missing one. ## 0.34.0 ### Highlights TypeGraph 0.34 adds provenance-backed source retraction through `@nicia-ai/typegraph/provenance`. Applications map their graph kinds to sources, justifications, facts, premises, and derivations; retracting a source then makes unsupported reachable facts non-current while preserving their edges and recorded history. `unRetract` reverses the belief transition without treating it as a domain delete. The release also strengthens transaction and import behavior: operation hooks report durably committed writes, conflicting updates preserve uniqueness reservations, tombstones survive update-mode imports, and incremental merges refuse inherited-row changes that occurred after planning. SQLite transaction handling and deterministic keyset pagination receive correctness fixes. ### Upgrade notes - Provenance retraction requires `history: true` and captures TypeGraph-managed writes. Out-of-band SQL remains outside recorded capture. - `onOperationEnd` now runs after the enclosing transaction commits. Consumers using hooks for metrics, cache invalidation, or audit events should account for the later notification boundary. - Incompatible property-schema changes, including narrowed enums and nested type changes, are now classified as breaking migrations. - `importGraph(..., { onConflict: "update" })` skips soft-deleted target rows. Use `onUnknownProperty: "allow"` for fidelity-preserving imports and `"strip"` when schema normalization is intended. - An incremental merge that detects a changed inherited target row returns a retryable `BaseVersionMismatchError`; recompute the merge against current state. ### Minor Changes - [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Add the `@nicia-ai/typegraph/provenance` subpath for provenance-backed source retraction. The first slice maps user graph kinds to source, justification, fact, premise, and derivation roles; supports multiple source node kinds and terminal fact kinds; requires `{ history: true }`; applies TypeGraph-managed belief transitions by making unsupported facts non-current; and keeps recorded-time replay available before and after retraction. A transition only touches facts reachable from the flipped sources, and closing a fact's currency is a belief-status change rather than a domain delete — the fact's edges are left untouched (no `restrict`/`cascade`/`disconnect` enforcement), so `unRetract` is an exact inverse of `retract`. PostgreSQL transitions serialize with TypeGraph-managed history writes on the same graph; out-of-band SQL remains outside recorded capture. ### Patch Changes - [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Stop opening a write transaction on `getOrCreateByConstraint`'s found path. The single-item node getOrCreate wrapped its whole body — probe included — in a transaction, so the common "already exists" case paid for `BEGIN IMMEDIATE` on SQLite (and, under history capture, the per-graph advisory lock on Postgres), and the nested create's operation hooks fired inside that outer transaction, reporting success before a COMMIT that could still fail. The probe now runs as a pure read; the create and update/resurrect legs each open their own (hooked) transaction, so `onOperationEnd` means durably committed. A concurrent create that reserves the key between the probe and the insert surfaces as a uniqueness conflict and is converged by a single re-probe. The bulk variant keeps its one enclosing transaction (atomic batch, hooks skipped by design). Edge `getOrCreateByEndpoints` gets the same probe-first shape. - [#191](https://github.com/nicia-ai/typegraph/pull/191) [`2cad229`](https://github.com/nicia-ai/typegraph/commit/2cad2293f2d937aff7f53a1318525814eeb05533) Thanks [@pdlug](https://github.com/pdlug)! - Guard `mergeIncremental()` against inherited-row lost updates. The incremental commit path re-checked new-row identity resolution and per-row resurrect/strip hazards, but not whether a committed row the plan mutates still held the value the plan merged against — so a concurrent write to an inherited row between planning (reads taken outside the transaction) and commit was silently discarded. The commit now re-reads, in-transaction, every committed target row the plan will change and aborts with a retryable `BaseVersionMismatchError` if it drifted, matching the snapshot merge path's TOCTOU contract. This covers all four mutating paths: node writes and node deletions (checked by `version`), and edge upserts and edge deletions (checked by a content signature over endpoints, liveness, and canonical props, since edges carry no version column). - [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - `importGraph(..., { onConflict: "update" })` now skips soft-deleted target rows instead of failing. Import never resurrects a tombstone: a node or edge that exists only as a tombstone counts as `skipped`, keeps its tombstone, and gets no uniqueness/embedding/fulltext side effects (a uniqueness reservation held by a tombstoned node would block live creates of the same value). Previously the update path attempted a live-row update that threw and aborted the whole import. `onUnknownProperty: "allow"` is also pinned as the fidelity-preserving strategy: it validates known fields but persists the given properties byte-for-byte — no transform re-application, no default injection — so an export→import round trip cannot corrupt values whose schema transforms are not idempotent; use `"strip"` for a normalizing import. - [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Fix a uniqueness-reservation corruption on a conflicting node update. `updateUniquenessEntries` mutated one constraint's sidecar at a time — releasing the old key before proving the new one free — so a caller that catches the resulting `UniquenessError` and still commits the transaction (notably `importGraph(..., { onConflict: "update" })`, which reports the conflict per row) left the node's already-mutated sidecars in a corrupt state: an earlier constraint's old key released (letting a later create silently duplicate it) or a new key wrongly reserved, while the row itself stayed unchanged. The update now runs in two passes — preflight every changed constraint's new key first, then apply all sidecar deletes and inserts only after every key is proven free — so a conflict throws with zero partial writes, for every caller of the shared node-write pipeline and for nodes with any number of unique constraints. - [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Make in-memory libsql databases safe across transactions, and fail loud on re-entrant root access. Local `@libsql/client` connections (`file:` paths and `file::memory:`) now frame transactions with raw `BEGIN IMMEDIATE`/`COMMIT` on the client's single stable connection instead of `client.transaction()`, which permanently hands that connection to the transaction and lazily opens a fresh — for `:memory:`, empty — database afterwards (tursodatabase/libsql-client-ts#229). Remote Turso connections keep using the driver's per-stream transactions. Separately, a store-level operation awaited from inside a `store.transaction` callback on the same SQLite backend (root store instead of the `tx` context) used to deadlock permanently — the open transaction holds the backend's serialized execution slot — and is now rejected with a `ConfigurationError` that points at the transaction-scoped context. - [#189](https://github.com/nicia-ai/typegraph/pull/189) [`fe21158`](https://github.com/nicia-ai/typegraph/commit/fe2115836d084a86613ae94a4403651d8316713a) Thanks [@pdlug](https://github.com/pdlug)! - Classify incompatible property-schema changes as breaking schema migrations. The migration diff previously compared only the top-level JSON-Schema token of each property, so a changed property type (e.g. `string` → `number`), a changed array item type (`string[]` → `number[]`), a narrowed enum, or a type change nested inside an object all auto-migrated silently as a non-blocking warning, leaving stored rows that no longer satisfy the declared schema; edge property changes were unconditionally treated as safe. Node and edge property diffs now share one recursive, conservative classifier: a change is `safe` only when it can be proven non-breaking (a new optional property, a metadata-only edit, or an additive optional field nested inside an object). Everything else — a removed property, a newly required property, an in-place type change, a changed array item schema, an enum/const/composition change, a same-type constraint change, or a breaking change nested inside an object — is `breaking` and blocks auto-migration. The `warning` severity is no longer emitted for property changes. - [#190](https://github.com/nicia-ai/typegraph/pull/190) [`1bfa9c2`](https://github.com/nicia-ai/typegraph/commit/1bfa9c28d04f03b9f82e23bf0a97417aba544767) Thanks [@pdlug](https://github.com/pdlug)! - Fix two silent query-correctness bugs. Keyset pagination (`paginate`/`stream`) now appends a unique `id` tiebreaker to the ORDER BY so a non-unique sort no longer drops equal-key rows across pages. And every compiled `LIKE`/`ILIKE` now emits `ESCAPE '\'` — including the case-sensitive `like` path, which previously omitted it — so escaped `%`/`_`/`\` match literally on SQLite as they already did on PostgreSQL, in both the auto-escaped operators (`contains`/`startsWith`/`endsWith`) and raw `like`/`ilike` patterns, and whether the pattern is a literal or a bound parameter (previously SQLite had no default LIKE escape character, so the two backends — and the direct vs prepared paths — diverged). - [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Fix a uniqueness-reservation loss on node resurrection. Resurrecting a soft-deleted node through `getOrCreateByConstraint` (or any `clearDeleted: true` upsert) ran the diff-based uniqueness maintenance, which skips a key that did not change — but the soft delete had already removed the node's uniqueness entries, so the resurrected node held NO reservation and a later `create` with the same unique value silently succeeded, duplicating it. A resurrecting update now re-checks and re-inserts the entries for its new props, exactly as the provenance reopen path does. - [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Open SQLite business-write transactions with `BEGIN IMMEDIATE` on the sync (better-sqlite3) path, matching schema writes and the async libsql/Drizzle path. A deferred `BEGIN` acquired the reserved write lock only on the first write, so a read-then-write inside a transaction could fail with "database is locked" against a writer on another connection to the same file; taking the lock at the start of the transaction lets SQLite's busy timeout wait for it instead. The per-backend serialized write queue continues to order a single backend's own transactions. - [#192](https://github.com/nicia-ai/typegraph/pull/192) [`2af3a06`](https://github.com/nicia-ai/typegraph/commit/2af3a065d9d54b0ac89c32dc27d637a4eedc58cf) Thanks [@pdlug](https://github.com/pdlug)! - Type-check the remaining StoreView read-name buckets. `CURRENT_ONLY_READ_NAMES` and `EDGE_BATCH_READ_NAMES` were plain `as const` arrays while every sibling bucket carried a `satisfies readonly (keyof Collection)[]` guard, so a renamed or mistyped method in those two would have gone uncaught at compile time. All six buckets are now checked against the live collection keys. Compile-time only. - [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Operation hooks now mean "durably committed" everywhere. `onOperationEnd` previously fired when an operation completed, even when that operation ran inside an enclosing transaction whose COMMIT later failed — so hook consumers (metrics, cache invalidation, audit logs) were told a rolled-back write succeeded. Operations inside `store.transaction` now defer their success hooks until the transaction commits, and a failed transaction converts every completed operation's pending success into `onError`. Edge `getOrCreateByEndpoints` no longer wraps its write legs in an outer transaction (each leg commits — and reports — on its own, with a probe/create race converged by one retry), and provenance transitions route their source-flip and per-fact hooks through the same deferred lifecycle. Inside an adopted transaction (`withTransaction` / `withRecordedTransaction`) the commit belongs to the caller and cannot be observed; hooks there keep firing at operation completion, as documented. ## 0.33.0 ### Highlights TypeGraph 0.33 adds recorded/system time alongside valid time. Enabling `history: true` captures TypeGraph-managed node and edge writes into historical relations at a monotonic per-graph commit instant, making it possible to reconstruct values that were later corrected or deleted. `store.asOfRecorded()` exposes read-only reconstruction through point reads, queries, subgraph extraction, and graph algorithms. Applications can pin recorded and valid time independently to distinguish when a fact was true from when TypeGraph knew it. External history producers can bind compatible recorded relations through `recordedRelation()` and `recordedRead`. ### Upgrade notes - Capture is opt-in and does not backfill existing rows. Enable it on a fresh graph for complete history; existing entities are first captured on their next managed write. Capture requires a transactional backend with statement execution. - Use typed collection writes for captured transactions. Raw `tx.sql` is unavailable under `history: true`; adopt caller-owned transactions through `withRecordedTransaction(externalTx, callback)` so capture flushes before commit. - Use `recordedNow()` as a post-write reconstruction anchor after checking for `undefined`. Recorded instants can advance ahead of wall-clock time, and recorded coordinates must use canonical UTC ISO 8601. - Recorded views expose only reconstructible reads. Current full-text/vector indexes and unsupported broad collection reads are refused; `asOfRecorded(T)` pins both time axes to `T` unless composed with an explicit valid-time view. - Custom backend and SQL tooling must honor the new backend-role and row-versus-statement SQL intent brands. `recordedRead` descriptors must come from `recordedRelation()` and cannot be combined with `history: true`. ### Minor Changes - [#186](https://github.com/nicia-ai/typegraph/pull/186) [`655407a`](https://github.com/nicia-ai/typegraph/commit/655407a9c225e8eca0aff5f636ed17ca99f3e382) Thanks [@pdlug](https://github.com/pdlug)! - Add recorded / system-time capture — TypeGraph's second temporal axis. Where valid time (`validFrom` / `validTo`, queried via `asOf` / `includeEnded`) records _when a fact was true in the world_, recorded time records _when TypeGraph captured a managed node/edge write_. Together they answer "what did TypeGraph reconstruct as true, as of a captured commit instant?" — surfacing values that were later corrected (à la SQL:2011 system-versioned tables). Enable capture per store with `createStore(graph, backend, { history: true })`. TypeGraph collection writes through that store are then captured into recorded-time relations (`typegraph_recorded_nodes` / `typegraph_recorded_edges`), stamped with a per-graph monotonic commit instant from a `typegraph_recorded_clock` (serialized on PostgreSQL via a per-graph advisory lock). Capture is opt-in and has **no backfill** — enable it on a fresh graph, since an entity that already exists is first recorded the next time it is written. It requires a transactional backend with statement execution (the built-in SQLite / PostgreSQL backends). Read at a recorded instant with `store.asOfRecorded(T)`, which returns a narrow read-only `RecordedStoreView`. Direct `store.asOfRecorded(T)` is diagonal bitemporal sugar (recorded _and_ valid axes both at `T`); chain `store.asOf(validT).asOfRecorded(recordedT)` to pin the two axes independently, or `store.view({ mode }).asOfRecorded(recordedT)` to compose recorded time with any valid-time mode (e.g. `includeTombstones`). `store.recordedNow()` returns the recorded high-water mark; after guarding the `undefined` case, passing that value to `store.asOfRecorded(...)` is a deterministic "as things stand now" anchor. Recorded instants are monotonic and can run briefly ahead of wall-clock time under bursty writes, so the wall clock is not a reliable anchor right after a write. The recorded view is a **reconstructing** lens that exposes only reads which can be faithfully rebuilt from the history relations: point reads (`nodes..getById` / `getByIds` and the edge equivalents), a sealed `query()`, `subgraph()`, and the graph algorithms (`reachable` / `canReach` / `shortestPath` / `degree`). Broad collection reads (`find` / `count` / `findFrom`), `search`, and fulltext / vector predicates refuse with a `ConfigurationError` / `UnsupportedPredicateError` — those indexes reflect current state only. `T` must be a canonical UTC ISO-8601 timestamp (`YYYY-MM-DDTHH:mm:ss.sssZ`). The public live-read and algorithm option types explicitly reject internal recorded coordinates, while recorded internals use a branded `RecordedInstant` so only validated canonical recorded instants can flow through the reconstructing paths. Recorded read binding is now explicit without exposing TypeGraph's internal capture binding. `history: true` enables TypeGraph-managed capture and binds the built-in recorded relations internally, while the factory-branded `recordedRelation({ schema })` / `recordedRead` path is the external-read-source API for hosts that populate a row-compatible recorded relation outside TypeGraph's writer wrapper. The store validates that runtime `recordedRead` values come from `recordedRelation({ schema })`, rejects `recordedRead` combined with `history: true`, and factory-brands/freezes SQL schema and recorded-read descriptors so they cannot be structurally forged as plain objects. Store overloads reflect that split: history-enabled stores expose `HistoryStore`, read-bound live stores expose `RecordedReadStore`, and captured-history stores expose `HistorySafeBackend` / `HistoryTransactionContext` types that hide raw statement / DDL write seams from the typed `backend`, `transaction()`, and `withRecordedTransaction()` surfaces. Writes under `history: true` flush capture at transaction commit, so they must go through the typed collections: raw `tx.sql` is disabled (it would bypass capture), and `store.withTransaction(externalTx)` is replaced by the callback form `store.withRecordedTransaction(externalTx, async (tx) => ...)`, which gives capture a flush point before the caller commits. `store.clear()` clears the recorded relations alongside the live tables. Node creates now run atomically on transactional backends with uniqueness, vector, and fulltext finalization, and node delete cascades now run atomically even without `history: true`. A failed finalize step rolls back the node row instead of leaving a partially indexed row behind. Overlapping PostgreSQL cascades may hold locks longer, so callers should keep normal deadlock-retry handling around concurrent deletes. Backend and SQL execution contracts are more explicit for maintainers and extension authors: backend role brands separate graph-write paths from raw/bulk paths, `execute` / `executeStatement` now require row-vs-statement SQL intent brands, transaction backends are composed from explicit backend facets instead of `Omit`, and backend wrappers use an exact overlay helper that preserves prototype/proxy backends while catching typoed override keys at compile time. Exports `RecordedStoreView` and its collection types (`RecordedStoreViewNodeCollection` / `RecordedStoreViewNodeCollections`, `RecordedStoreViewEdgeCollection` / `RecordedStoreViewEdgeCollections`, `TypedRecordedStoreViewEdgeCollection`). **Performance:** recorded reads reconstruct from the history relations rather than the live tables, so they are slower than current-state reads — most noticeably for full-graph `subgraph` / algorithm reconstructions on PostgreSQL. Use `asOfRecorded` for audit and point-in-time reconstruction, not hot-path reads. ## 0.32.0 ### Minor Changes - [#182](https://github.com/nicia-ai/typegraph/pull/182) [`0f0e771`](https://github.com/nicia-ai/typegraph/commit/0f0e77161d473b5c3b2d2e224d930c611eb4b123) Thanks [@pdlug](https://github.com/pdlug)! - Close the TOCTOU windows in graph-merge commits. A merge resolves its plan from reads taken before the commit transaction, so a write landing on the target in between could previously be committed over. Now, inside the commit transaction: `merge()` and `mergeAgainstBase()` re-validate the target's base@V content fingerprint, and `mergeIncremental()` re-runs its new-vs-base identity resolution (the unique-constraint and block-index probes). All three fail with `BaseVersionMismatchError` — instead of committing a stale plan or a duplicate entity — when the target changed in that window. Merge commits run at `SERIALIZABLE` isolation with bounded retry on serialization failures and deadlocks, making the guards race-free on multi-writer Postgres. `Store.transaction()` accepts optional `TransactionOptions` (isolation level) and `TransactionContext` exposes the transaction-scoped `backend`. - [#185](https://github.com/nicia-ai/typegraph/pull/185) [`4e23be8`](https://github.com/nicia-ai/typegraph/commit/4e23be8d6af94b965bdcf90e911dc0e1c49d2bad) Thanks [@pdlug](https://github.com/pdlug)! - Add `StoreView`, a read-only `(mode, asOf)` lens over a `Store` that pins a temporal coordinate and routes every supported read through it (the as-of database value, à la Datomic `(d/as-of db t)` / SQL:2011 `FOR SYSTEM_TIME AS OF`). Construct one with `store.asOf(T)` (valid-time) or `store.view({ mode, asOf })` for the other public modes (`current` / `includeEnded` / `includeTombstones`). The view exposes pinned `nodes` / `edges` collections (`getById` / `getByIds` / `find` / `count`, edge `findFrom` / `findTo`), a pre-pinned `query()`, `subgraph()`, and the graph algorithms (`reachable` / `canReach` / `shortestPath` / `neighbors` / `degree`). It is read-only by construction — writes and temporally-unscoped reads refuse with a clear error — and `search` refuses on a non-`current` pin (the fulltext / vector index reflects current state only). Internally every pinned surface injects a single opaque `ReadCoordinate` through one helper, so a future temporal axis (recorded / system time) lands on every surface at once instead of splitting per surface. The view's read surface is derived from a read/write split of the live collection types (`NodeTemporalReads` / `NodeCurrentReads` / `NodeWrites` and edge equivalents, now exported) with a `test-d` conformance check, so a new collection read cannot silently bypass the view's pinning decision. - **`store.snapshot()`.** Sugar for `store.asOf(new Date().toISOString())` — a read-only view pinned to the current instant captured once at construction. Unlike `store.view({ mode: "current" })` (which tracks "now" live), a snapshot is a stable point-in-time value where every surface observes the same instant. Mirrors Datomic's `(d/db conn)`. - **Sealed pinned query.** `view.query()` now returns a query builder whose temporal axis is sealed — calling `.temporal(...)` on it throws — so a pinned view cannot be silently re-coordinated per query. - **Current-only reads.** Constraint / index lookups (`findByConstraint`, `bulkFindByConstraint`, `bulkFindByIndex`), which have no temporal axis, are now available on a `current` view (delegating to the live store) and refuse with a clear error on a temporal pin — instead of being unavailable on every view. **Breaking — `find` / `count` signature:** `store.nodes..find(...)` / `count(...)` and `store.edges..find(...)` / `count(...)` now take the temporal coordinate as a **second** argument rather than inline in the filter object: `find(filter?, temporal?)` / `count(filter?, temporal?)`. For example, `nodes.Person.find({ where, temporalMode: "asOf", asOf })` becomes `nodes.Person.find({ where }, { temporalMode: "asOf", asOf })`, and `edges.worksAt.count({ temporalMode: "includeEnded" })` becomes `edges.worksAt.count(undefined, { temporalMode: "includeEnded" })`. Old call sites that inlined `temporalMode` / `asOf` are now type errors. `getById` / `getByIds` / `findFrom` / `findTo` / node `count` are unchanged (they already took a trailing temporal argument). **Breaking — canonical `validFrom` / `validTo` on write:** `create` / `update` / `bulk*` now require canonical fixed-width UTC ISO timestamps (`YYYY-MM-DDTHH:mm:ss.sssZ`) for `validFrom` / `validTo`, rejecting date-only, zoned-offset, variable/missing-millisecond, and rollover values with a `ValidationError`. This makes the _stored_ values that temporal filters compare as text always sort chronologically — the same contract the `asOf` read coordinate already enforces, applied uniformly to every timestamp in the system. Convert non-canonical inputs with `new Date(value).toISOString()`. There is no migration: pre-existing non-canonical rows are left as-is (recreate them if affected) — acceptable pre-1.0. **Behavior change:** `store.edges..findFrom(...)` / `findTo(...)` / `findByEndpoints(...)` (and their `batchFindFrom` / `batchFindTo` / `batchFindByEndpoints` variants) now honor the temporal model like `getById` / `find` instead of returning every non-soft-deleted edge. With no temporal argument, the graph's default `temporalMode` applies — so under the default `"current"` mode, edges outside their `validFrom` / `validTo` window are now excluded. Pass `temporalMode` / `asOf` to read at another coordinate (e.g. `temporalMode: "includeEnded"` to recover the previous "all non-deleted" behavior). `findByEndpoints` / `batchFindByEndpoints` gain a trailing `temporal?` argument and are now pinnable on a `StoreView` (no longer refused on a temporal pin). The internal `getOrCreate*ByEndpoints` identity lookup is unaffected — it deliberately matches against all edges regardless of validity window. **Read coordinates:** `asOf`, `.temporal("asOf", T)`, algorithms, subgraph, and `StoreView` require canonical UTC ISO timestamps (`YYYY-MM-DDTHH:mm:ss.sssZ`) for the same lexicographic-comparison reason. ## 0.31.0 ### Highlights TypeGraph 0.31 introduces graph branching and semantic merge through `@nicia-ai/typegraph/graph-merge`. `branch()` creates an isolated working copy; `merge()` reconciles one or more branches using stable IDs, declared uniqueness constraints, blocking keys, and optional similarity scoring. Canonical survivors receive merged properties and repointed edges, with conflicts and source contributions recorded in the report. Snapshot merges check the branch's base token, while `mergeIncremental()` supports targets that have advanced since the fork. Applications can configure property and delete/modify conflict policies, use ontology-aware type reconciliation, and optionally persist provenance in a separate sidecar graph. ### Upgrade notes - Merge requires a transaction-capable backend and refuses non-atomic execution. Vector and hybrid entity-resolution strategies require a configured embedder. - Choose snapshot or incremental merge according to whether the target may advance after the fork. A base mismatch requires a fresh branch or an appropriate incremental merge workflow. - Optional `persistProvenance` runs after the graph commit. A persistence failure is a report warning and does not roll back the merged graph. ### Minor Changes - [#178](https://github.com/nicia-ai/typegraph/pull/178) [`6b6e418`](https://github.com/nicia-ai/typegraph/commit/6b6e4186642c65d58c939250458b6521efbc40c7) Thanks [@pdlug](https://github.com/pdlug)! - Add `@nicia-ai/typegraph/graph-merge`, a TypeGraph-native branch and semantic merge subpath for deterministic entity-resolution merges across graph forks. ## 0.30.0 ### Minor Changes - [#171](https://github.com/nicia-ai/typegraph/pull/171) [`f5defd3`](https://github.com/nicia-ai/typegraph/commit/f5defd35b331e56f282d4eb501b98d3b9affe562) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.nodes..bulkFindByIndex(indexName, items, options?)` — batched candidate retrieval against declared node indexes, including non-unique ones. For each input record it returns the live nodes that share that record's declared index key, for import reconciliation, dedup-candidate discovery, and joining records against the graph by a composite key. Each input yields its own array (candidate retrieval, not a uniqueness guarantee); buckets preserve input order and are ordered by node id. TypeGraph owns the index semantics: keys are computed from `index.fields` only (reusing the index's own extraction expressions), the partial `where` is applied to stored rows, and a missing/`undefined` indexed field matches a stored `NULL` via a new null-safe-equality dialect adapter. An optional `limitPerInput` caps each bucket — in SQL via `ROW_NUMBER()` when the backend supports window functions, otherwise capped in memory with the same result. Date-typed key fields are rejected with `ConfigurationError` because they can't compare identically across SQLite and PostgreSQL. Unknown index names throw `NodeIndexNotFoundError`. `createLocalSqliteBackend` also gains a `capabilities` override for simulating engine capability gaps (e.g. `windowFunctions: false`) in tests. - [#173](https://github.com/nicia-ai/typegraph/pull/173) [`bd96cfb`](https://github.com/nicia-ai/typegraph/commit/bd96cfbeadde11c6986fb667f9a86b0ba0b5b1bd) Thanks [@pdlug](https://github.com/pdlug)! - Add the `backend.capabilities.windowFunctions` capability and reject relevance-ranking queries before SQL generation when a custom backend profile disables SQL window functions. ## 0.29.0 ### Highlights TypeGraph 0.29 makes vector storage and hybrid search portable across pgvector, sqlite-vec, and libSQL/Turso through a pluggable `VectorStrategy`. Embeddings move into graph-scoped, fixed-dimension storage per field, with strategy-derived metric and index capabilities, legacy migration tooling, and field re-embedding after dimension changes. PGlite gains first-class backend support and a local factory for running PostgreSQL and pgvector in process. SQLite set operations now compile their operands through the full query compiler, preserving traversal, search, nested ordering, and pagination behavior across backends. ### Upgrade notes - Existing vector deployments must run `migrateLegacyEmbeddings()` after upgrading. Search no longer reads the shared legacy `typegraph_node_embeddings` table. Use `reembedVectorField()` when an embedding field's dimensions change. - Install the optional PGlite peers when using the local PGlite backend. Configure `vector: false` when the engine has no vector extension. - Custom capability consumers must remove the descriptive-only `jsonb`, `ginIndexes`, `partialIndexes`, `cte`, and `returning` flags. Query semantics are provided by the shared compiler and configured strategies. - Interchange payloads with `source.type: "typegraph-cloud"` must be retagged as `"external"`; the removed source variant now fails validation. ### Minor Changes - [#161](https://github.com/nicia-ai/typegraph/pull/161) [`9e86269`](https://github.com/nicia-ai/typegraph/commit/9e862695c6a3341af5d8acbd4f652738bd7727ca) Thanks [@pdlug](https://github.com/pdlug)! - Add cross-backend vector and hybrid search through a pluggable `VectorStrategy`, closing [#157](https://github.com/nicia-ai/typegraph/issues/157). TypeGraph now has first-class vector storage and search for libSQL/Turso, sqlite-vec, and pgvector behind the same semantic search APIs. Backend highlights: - libSQL/Turso stores fixed-dimension embeddings in `F32_BLOB(N)` columns, supports cosine/L2 search, and can use DiskANN through `libsql_vector_idx` and `vector_top_k`. - sqlite-vec uses `vec0` KNN tables instead of brute-force vector scans. - pgvector uses graph-scoped, per-field `vector(N)` tables with HNSW/IVFFlat materialization. - Backends advertise vector metrics, index types, and dimension limits from the active strategy, and `createSqliteBackend` / `createPostgresBackend` accept a custom `vector?: VectorStrategy`. The release also adds migration and lifecycle tooling for the new storage model: - `migrateLegacyEmbeddings(...)` copies existing rows out of the legacy shared `typegraph_node_embeddings` table. - `store.reembedVectorField(kind, fieldPath, { embed? })` recreates a field's storage after an embedding dimension change and can re-embed existing rows. - `store.materializeRemovals()` reclaims vector tables for removed embedding fields and reports them in `MaterializeRemovalsResult.reclaimedVectorFields`. **Breaking storage change:** vector embeddings now live in graph-scoped, fixed-dimension per-field storage instead of the shared `typegraph_node_embeddings` table. Search no longer reads the legacy table. Deployments with existing embeddings must run `migrateLegacyEmbeddings(...)` once after upgrading; deployments without stored embeddings need no migration. - [#165](https://github.com/nicia-ai/typegraph/pull/165) [`ae5bfdc`](https://github.com/nicia-ai/typegraph/commit/ae5bfdc55aae3531bcd75f0770cb2812ad9682d9) Thanks [@pdlug](https://github.com/pdlug)! - Reduce `BackendCapabilities` to the flags the library actually consumes: `transactions`, `vector`, and `fulltext`. The descriptive-only flags `jsonb`, `ginIndexes`, `partialIndexes`, `cte`, and `returning` were never read anywhere to gate a query feature or pick an index strategy. `jsonb`/`ginIndexes` additionally misrepresented SQLite, which has native JSON (`json_extract`/`json_each`) and supports B-tree expression indexes on scalar JSON properties at parity with PostgreSQL — the only real JSON difference (GIN containment acceleration) is a Postgres performance characteristic, not a gated capability. If you were reading any of these removed flags, branch on `backend.dialect === "postgres"` instead, or rely on the dialect layer (JSON-path predicates, `WITH` queries, `RETURNING`, partial indexes, and `defineNodeIndex`/`defineEdgeIndex` work the same on both backends). - [#163](https://github.com/nicia-ai/typegraph/pull/163) [`0175a25`](https://github.com/nicia-ai/typegraph/commit/0175a2585029aa1b6ceabc9889074a72b8895d03) Thanks [@pdlug](https://github.com/pdlug)! - Add first-class support for [PGlite](https://pglite.dev/) (Postgres-in-WASM), closing [#160](https://github.com/nicia-ai/typegraph/issues/160). - **Execution fast-path fix.** `createPostgresBackend` now detects a PGlite `db.$client` and routes it to the unnamed positional query wrapper. PGlite's `.query` has no node-postgres named-statement config form — passing one desyncs its single connection (`08P01`), so under the default `prepareStatements: true` every query previously failed. PGlite works unchanged with `createPostgresBackend(drizzle(pglite))` now. - **`createLocalPgliteBackend`** — a batteries-included helper under the new `@nicia-ai/typegraph/postgres/pglite` entry, the Postgres analog of `createLocalSqliteBackend`. It constructs an in-process PGlite engine (in-memory by default, or any `dataDir`), loads pgvector, runs the schema DDL, and returns `{ backend, db, client }` whose `close()` disposes the engine. Pass `vector: false` to skip the extension, or `vector: ` to bring your own pgvector build. `@electric-sql/pglite` (and, for vector support, `@electric-sql/pglite-pgvector` on PGlite ≥ 0.5) are optional peer dependencies. The biggest payoff: the Postgres dialect and pgvector path can now be exercised in plain `pnpm test` with zero Docker. - [#162](https://github.com/nicia-ai/typegraph/pull/162) [`48a6ffc`](https://github.com/nicia-ai/typegraph/commit/48a6ffc3e63459e7a2535a936a8c9c3fbcd29a99) Thanks [@pdlug](https://github.com/pdlug)! - Add `vector: false` to `createPostgresBackend` to disable the vector stack. The Postgres backend wires `pgvectorStrategy` by default, assuming a standalone Postgres server has the pgvector extension installed. An in-process Postgres (PGlite) built without that extension can't honor it — the default strategy's `vector(N)` DDL hard-fails the moment an embedding is written or `CREATE EXTENSION vector` runs. Passing `vector: false` turns the stack off: the backend advertises no `capabilities.vector` and omits the embedding/search methods, mirroring a SQLite connection without sqlite-vec, so the store never routes vector work to it. Real-Postgres behavior is unchanged — the default remains `pgvectorStrategy`. - [#158](https://github.com/nicia-ai/typegraph/pull/158) [`bc07847`](https://github.com/nicia-ai/typegraph/commit/bc07847cbde20eedd01781062e0403856cb46079) Thanks [@pdlug](https://github.com/pdlug)! - Export the ontology transitive-closure utilities (`computeTransitiveClosure`, `invertClosure`, `isReachable`) from the package root. These were previously internal-only. Exposing them lets consumers reason over `subClassOf` / `equivalentTo` hierarchies — e.g. reconciling node types when merging graphs from independent sources. - [#166](https://github.com/nicia-ai/typegraph/pull/166) [`a32d31f`](https://github.com/nicia-ai/typegraph/commit/a32d31f7bbe9fc4657eb956e86900eaf1c283ef9) Thanks [@pdlug](https://github.com/pdlug)! - Remove the `typegraph-cloud` source type from the interchange `GraphDataSourceSchema`. TypeGraph Cloud is not a publicly available product, so the `typegraph-cloud` variant has been dropped from the graph-data source discriminated union, and the corresponding interchange documentation has been removed. `GraphDataSource` now accepts only `typegraph-export` and `external`. **Breaking:** importing data whose `source.type` is `"typegraph-cloud"` now fails schema validation. Re-tag such payloads as `"external"` before importing. - [#165](https://github.com/nicia-ai/typegraph/pull/165) [`ae5bfdc`](https://github.com/nicia-ai/typegraph/commit/ae5bfdc55aae3531bcd75f0770cb2812ad9682d9) Thanks [@pdlug](https://github.com/pdlug)! - Support the full query feature set inside SQLite set operations (`UNION`/`UNION ALL`/`INTERSECT`/`EXCEPT`). Previously the SQLite set-operation compiler hand-rolled a thin subset of leaf compilation and rejected leaves that used traversals, `EXISTS`/`IN` subqueries, vector or fulltext predicates, `GROUP BY`/`HAVING`, or per-leaf `ORDER BY`/`LIMIT`/`OFFSET` — throwing `UnsupportedPredicateError` at execution time. PostgreSQL accepted all of these. The result was a portability cliff: a combined query developed against PostgreSQL could throw the moment the backend was switched to SQLite. Both dialects now compile every leaf with the full query compiler and only differ in how each operand is wrapped. SQLite forbids parenthesized compound operands, but it does allow a `WITH` clause inside a FROM-subquery, so each operand is emitted as `SELECT * FROM ()`. This keeps every leaf's CTEs (traversal joins, recursive expansions, vector/fulltext relevance) scoped to its own subquery and lets per-leaf `ORDER BY`/`LIMIT`/`OFFSET` live inside the wrap. Nested set operations are wrapped the same way, preserving the AST's grouping regardless of the dialect's native compound-operator associativity. As a side effect, vector/fulltext predicates in set-operation leaves now use the backend's configured relevance strategy instead of falling back to the dialect default. Note: `GROUP BY`/`HAVING` leaves are supported at the compiler level, but the query builder still does not expose `.union()`/`.intersect()`/`.except()` on aggregate queries — that builder gate is unchanged and applies equally to both backends. ### Patch Changes - [#165](https://github.com/nicia-ai/typegraph/pull/165) [`ae5bfdc`](https://github.com/nicia-ai/typegraph/commit/ae5bfdc55aae3531bcd75f0770cb2812ad9682d9) Thanks [@pdlug](https://github.com/pdlug)! - Fix `ORDER BY`/`LIMIT`/`OFFSET` being silently dropped on a nested set-operation operand. When a set operation was nested inside another — e.g. `a.union(b).limit(10).intersect(c)` — the inner compound's suffix clauses were applied only at the top level, so the inner `limit`/`offset` were ignored and the outer operation ran over the full (unlimited) inner result. The compiler now emits each nested compound's own `ORDER BY`/`LIMIT`/`OFFSET` inside its operand subquery on both SQLite and PostgreSQL. - [#165](https://github.com/nicia-ai/typegraph/pull/165) [`ae5bfdc`](https://github.com/nicia-ai/typegraph/commit/ae5bfdc55aae3531bcd75f0770cb2812ad9682d9) Thanks [@pdlug](https://github.com/pdlug)! - Validate set-operation leaf vector predicates against the configured vector strategy rather than only the dialect's fallback metric list, so a custom strategy's metric (e.g. `inner_product` on SQLite) is accepted inside `UNION`/`INTERSECT`/`EXCEPT` leaves exactly as it is in a standalone query. Reject a per-query fulltext `language` override on the query-builder path (`.$fulltext.matches(..., { language })`) when the strategy's tokenizer is fixed at table-create time (SQLite/FTS5), matching the store-level search guard instead of silently ignoring the option. ## 0.28.1 ### Patch Changes - [#154](https://github.com/nicia-ai/typegraph/pull/154) [`6703c88`](https://github.com/nicia-ai/typegraph/commit/6703c880d3d9047149f91d1db4a27b414983c632) Thanks [@pdlug](https://github.com/pdlug)! - Fix `isMissingTableError` missing DrizzleQueryError-wrapped Postgres "relation does not exist" errors, breaking fresh/partial Postgres boot ([#153](https://github.com/nicia-ai/typegraph/issues/153)). `isMissingTableError` (the shared "relation not bootstrapped yet" discriminant for `loadActiveSchemaWithBootstrap`, `readActiveSchemaPure`, and the [#135](https://github.com/nicia-ai/typegraph/issues/135) durable-marker gate) classified failures by inspecting only `error.message`. On Postgres, drizzle-orm wraps every query-builder call (`db.select()`, `db.insert()`, …) in a `DrizzleQueryError` whose `.message` is the failed SQL text; the real driver error — carrying both `relation "…" does not exist` and SQLSTATE `42P01` — is preserved on `error.cause`, which the helper never walked. So the helper returned `false` and a benign "not bootstrapped yet" surfaced as a hard fault. This regressed `createStoreWithSchema` after the [#149](https://github.com/nicia-ai/typegraph/issues/149)/[#152](https://github.com/nicia-ai/typegraph/issues/152) read-only pre-check: `ensureRuntimeContributions` now calls `getMarker` (a query-builder read) on the possibly-absent `typegraph_contribution_materializations` table _before_ `ensureMarkerTable()`. On Postgres that read throws a `DrizzleQueryError`, the helper missed it, and the open rethrew instead of materializing — breaking seed, first boot, and test global-setup on any fresh or partial Postgres database (base tables present, marker table absent — e.g. drizzle-kit-managed schemas). SQLite was unaffected because better-sqlite3 throws a raw error whose `.message` literally contains `no such table`. `isMissingTableError` now walks the `error.cause` chain (cycle-safe) and additionally keys on the locale-independent SQLSTATE `42P01`, rather than matching only the outermost `.message`. Existing message patterns are retained, so all prior matches still hold; the fix applies uniformly to all three call sites, including the latent slow-path blind spot in `loadActiveSchemaWithBootstrap` / `readActiveSchemaPure`. ## 0.28.0 ### Minor Changes - [#150](https://github.com/nicia-ai/typegraph/pull/150) [`f9b1300`](https://github.com/nicia-ai/typegraph/commit/f9b1300a031eb758ae456fcd97ba8cbfdf93a2b8) Thanks [@pdlug](https://github.com/pdlug)! - Add a per-search `efSearch` knob for tuning pgvector HNSW recall ([#148](https://github.com/nicia-ai/typegraph/issues/148)). `store.search.vector` and the vector half of `store.search.hybrid` now accept an optional `efSearch` — the HNSW search frontier (`hnsw.ef_search`, default 40). pgvector caps a single index scan at `ef_search` candidates, so the hybrid over-fetch (`vectorK = 4 * limit` by default) silently under-delivers once `vectorK` climbs past the session default; the floor is `efSearch >= vectorK` and ~2–4× is the high-recall target. Being per-search lets one connection pool serve both a latency-sensitive interactive path and a recall-sensitive batch path. The Postgres backend applies it transaction-locally (`SET LOCAL hnsw.ef_search`) around the vector `SELECT`, so it never leaks to the next query on a pooled connection — `SET LOCAL` issued in autocommit would roll off with the statement and the next pooled query would see the session default. Omitting `efSearch` opens no transaction and preserves today's behavior exactly. Validated as a positive integer ≤ 1000 (pgvector's ceiling). Scope: pgvector HNSW only. sqlite-vec has no equivalent frontier knob and treats it as a no-op; transaction-less Postgres drivers (`drizzle-orm/neon-http`) ignore it with a one-time warning. IVFFlat's `ivfflat.probes` is a follow-up. ### Patch Changes - [#152](https://github.com/nicia-ai/typegraph/pull/152) [`761c672`](https://github.com/nicia-ai/typegraph/commit/761c672a991ea75454e441a4baf5939792da9505) Thanks [@pdlug](https://github.com/pdlug)! - Fix `ensureRuntimeContributions` running marker-table DDL on every store open ([#149](https://github.com/nicia-ai/typegraph/issues/149)). `createStoreWithSchema` → `ensureRuntimeContributions` previously ran the `typegraph_contribution_materializations` marker DDL (`ensureMarkerTable()` → `CREATE TABLE IF NOT EXISTS …`) on **every** open for any graph with runtime contributions (e.g. `searchable()` fields), even when every contribution was already materialized. The per-materializer `initializedGraphIds` cache is per-instance, so a deployment that builds a fresh backend per request (the norm on serverless Postgres) got an empty cache each time and re-ran the DDL on every open — which intermittently fails on connections that can't run it (observed on Cloudflare Workers + the Neon serverless driver) and surfaces as a wrapped `DrizzleQueryError` rather than a clean `MigrationError`. `ensureRuntimeContributions` now does a read-only pre-check first, mirroring the SELECT-only `assertInitialized`: when every runtime contribution is already materialized (marker present, signature matches, no recorded error) it returns without `ensureMarkerTable()` / `materializeOne`. A missing marker table, or any missing/stale/failed contribution, still falls through to the unchanged privileged first-materialization path. Warm per-request opens are now DDL-free. Note: the canonical runtime attach for the least-privilege / per-request deployment model remains `createVerifiedStore` (zero DDL by construction); `createStoreWithSchema` also runs bootstrap and auto-migration DDL and is still intended to run once under a privileged role. This change is defense-in-depth for the marker DDL specifically. ## 0.27.0 ### Minor Changes - [#144](https://github.com/nicia-ai/typegraph/pull/144) [`30a1cfd`](https://github.com/nicia-ai/typegraph/commit/30a1cfdba6f55240f3251de1ebdb05d69a66ea4c) Thanks [@pdlug](https://github.com/pdlug)! - Add `createVerifiedStore` and `assertSchemaCurrent` — the runtime counterparts of `createStoreWithSchema` for the least-privilege deployment model. `createStoreWithSchema()` runs DDL (bootstrap, safe auto-migrations, durable contribution materialization) and must run under a role with `CREATE` privileges. For applications that want their runtime under a least-privilege, DML-only role, the previous options were `createStore` (zero-DDL attach with no schema gate — drift goes undetected until a hot-path operation trips) or hand-rolling a SELECT-only verification dance from `getActiveSchema` + `getSchemaChanges`. This release adds two cleanly named entrypoints that share the same zero-DDL verification path: - **`createVerifiedStore(graph, backend, options?)`** — a SELECT-only attach (zero DDL) with a verification gate. Reads the active schema row and contribution markers, folds the persisted graph extension, and refuses to construct the Store unless the database is at the same schema version as the code graph. Returns `Promise<[Store, SchemaValidationResult]>` mirroring `createStoreWithSchema`. Throws `MigrationError` on any drift (safe or breaking — the least-privilege runtime cannot migrate), `ConfigurationError` when no schema has been initialized, and `StoreNotInitializedError` when the schema is current but runtime-contribution markers (e.g. fulltext) are missing/stale. - **`assertSchemaCurrent(backend, graph)`** — the same verification gate exposed as a standalone predicate for readiness probes / healthchecks. Returns the `SchemaValidationResult` or throws the same errors. The recommended deployment shape is now: 1. **Migration step** (privileged role with DDL/`CREATE`): run `createStoreWithSchema()` once at startup, or apply `generatePostgresMigrationSQL` / `generateSqliteMigrationSQL` plus a one-shot `createStoreWithSchema()` to materialize runtime contributions. 2. **Runtime** (least-privilege, DML-only role): attach with `createVerifiedStore()`. Zero DDL on the runtime path; schema drift fails fast with a clean `MigrationError` instead of leaking into hot-path operations or 500ing on a permission error. Internal: factored a pure `mergeStoredGraphExtension` helper out of `loadAndMergeGraphExtensionDocument` so the SELECT-only verifier reuses the same parse + extension-merge + deprecated-kind logic without going through the bootstrap-capable loader. No behavior change for the existing schema entrypoints. Documentation: "Database roles & least privilege" in `backend-setup.md` now folds in `createVerifiedStore` as the canonical runtime attach; `schema-management.md` covers Basic / Managed / Verified stores side by side; `troubleshooting.md` adds entries for `MigrationError` from a verifying attach and `ConfigurationError` on uninitialized databases. ### Patch Changes - [#144](https://github.com/nicia-ai/typegraph/pull/144) [`30a1cfd`](https://github.com/nicia-ai/typegraph/commit/30a1cfdba6f55240f3251de1ebdb05d69a66ea4c) Thanks [@pdlug](https://github.com/pdlug)! - Surface `MigrationError` before runtime-contribution DDL on a pending breaking migration ([#143](https://github.com/nicia-ai/typegraph/issues/143)). `loadActiveSchemaWithBootstrap` ran `ensureRuntimeContributions` (fulltext contribution DDL) **before** `ensureSchema` computed the schema diff and threw `MigrationError`. Contribution DDL is derived from the current code graph, so against a database still on the old schema version it was applied to a stale table shape. On Postgres the first failing statement aborts the surrounding transaction, and the error that escaped was the idempotent marker-table `CREATE TABLE IF NOT EXISTS "typegraph_contribution_materializations"` (collateral damage), not a clean `MigrationError`. Consumers using the documented migrate-on-`MigrationError` recovery pattern never saw a `MigrationError`, so the first request after every breaking schema change 500'd until a concurrent boot won the migration race. `loadActiveSchemaWithBootstrap` no longer materializes runtime contributions. `createStoreWithSchema` remains the single canonical durable-marker writer and runs the materialization step **after** `ensureSchema`, so the breaking-change gate is always reached first and a pending breaking migration throws `MigrationError` on the first request — making the migrate-then-retry recovery path work as documented. The pre-[#129](https://github.com/nicia-ai/typegraph/issues/129) `ensureFulltextTable` fallback is preserved at the canonical writer. No API changes. ## 0.26.0 ### Highlights TypeGraph 0.26 lets graph writes share a transaction with application-owned relational writes. `store.withTransaction(externalTx)` binds graph collections to the caller's connection, while `tx.sql` exposes the transaction handle when TypeGraph owns the boundary. Cloudflare Durable Objects SQLite gains asynchronous storage-transaction support, allowing graph and application writes to roll back together across awaited operations. Strategy-owned tables now follow one `TableContribution` contract. Full-text initialization becomes a durable, signature-checked materialization fact, allowing runtime operations to verify readiness without issuing DDL inside a business transaction. ### Upgrade notes - Initialize the parent Store through `createStoreWithSchema()` before adopting business transactions. Full-text operations refuse missing, stale, or failed materialization markers with `StoreNotInitializedError`. - `withTransaction()` requires real rollback support and refuses transactionless backends such as D1 and Neon HTTP. On synchronous better-sqlite3 connections, use an explicit transaction boundary instead of an async Drizzle transaction callback; Durable Objects use the asynchronous storage transaction runner. - Inside a TypeGraph-owned transaction, use its `tx.sql` handle for relational writes so they share the same connection. Using the outer database object on pooled backends can escape the transaction. - Custom `FulltextStrategy` implementations must replace `generateDdl()` with `ownedTables()`, declaring each table and its supporting indexes. Custom backends should implement the contribution initialization and durable materialization contracts described below. ### Minor Changes - [#139](https://github.com/nicia-ai/typegraph/pull/139) [`f1ea17c`](https://github.com/nicia-ai/typegraph/commit/f1ea17cafab281d61741b1d2ad0b26a769efaa5a) Thanks [@pdlug](https://github.com/pdlug)! - Cross-store atomicity: share one transaction across the TypeGraph store and an external Drizzle connection ([#134](https://github.com/nicia-ai/typegraph/issues/134)). Applications that persist into the same database through two layers — Drizzle for relational rows and TypeGraph for graph nodes/edges — previously had no way to make a write that spans both layers all-or-nothing. `store.transaction()` and `db.transaction()` each opened a _separate_ transaction on a _separate_ connection, so a failure between the two writes left either a stray relational row or a committed graph node with a dangling foreign reference. **What ships (additive — no breaking changes):** - New `Store.withTransaction(externalTx): TransactionContext`. The caller owns the transaction; `store.withTransaction(sqlTx)` returns a transaction-scoped `{ nodes, edges }` bound to that _exact_ connection, so both layers commit or roll back together. It is driver-agnostic; how you open the transaction is not. Async drivers (node-postgres, `neon-serverless` Pool, libsql): ```ts await db.transaction(async (sqlTx) => { const connector = await createConnectorRow(sqlTx, input); // Drizzle const txStore = store.withTransaction(sqlTx); await txStore.nodes.ArtifactSource.create({ // TypeGraph connectorId: connector.id, }); }); // one COMMIT / ROLLBACK ``` Synchronous `better-sqlite3` cannot use `db.transaction(async …)` (its driver rejects an `async` callback); open the transaction with explicit `BEGIN`/`COMMIT`/`ROLLBACK` instead and pass the connection to `withTransaction`. See the "Cross-Store Transactions" recipe for both shapes. - New optional `GraphBackend.adoptTransaction(externalTx)` member, implemented by the Drizzle Postgres and SQLite backends, plus the new `AdoptedTransaction` type. **Guarantees.** The adopted context reuses the parent store's already-resolved schema: it runs no `createStoreWithSchema` / `evolve` / `migrateSchema` and emits **no DDL inside the caller's business transaction**. Building on [#135](https://github.com/nicia-ai/typegraph/issues/135), fulltext operations assert the durable materialization marker (a cached `SELECT`, never DDL) and throw `StoreNotInitializedError` on a missing/stale/failed marker rather than migrating mid-transaction — so boot the parent store via `createStoreWithSchema` once at startup. When the backend cannot provide real rollback (`backend.capabilities.transactions === false`: `drizzle-orm/neon-http`, Cloudflare D1, SQLite `transactionMode: "none"`), `withTransaction` throws `ConfigurationError` rather than silently degrading — a non-atomic fallback is safe for graph-only writes but dangerous for cross-store flows, where the caller's relational write _would_ still commit. - [#142](https://github.com/nicia-ai/typegraph/pull/142) [`02c98a9`](https://github.com/nicia-ai/typegraph/commit/02c98a9933c888fcd732053e8cb47991614d2ec9) Thanks [@pdlug](https://github.com/pdlug)! - Transactional writes for Cloudflare Durable Objects SQLite (`do-sqlite`) ([#140](https://github.com/nicia-ai/typegraph/issues/140)). A store backed by `drizzle(ctx.storage)` previously fell back to non-transactional behavior, so TypeGraph mutations could not be composed atomically with a product's own relational ledger tables (e.g. `document_versions`, `change_events`) inside a Durable Object. **What ships (additive — no breaking changes):** - New SQLite `transactionMode: "do-sqlite"`, **auto-detected** for `drizzle(ctx.storage)`. Such backends now advertise `capabilities.transactions: true`. - `store.transaction(async (tx) => …)` and the caller-owned `store.withTransaction(db)` shape both work on Durable Objects. TypeGraph delegates to the async storage runner `ctx.storage.transaction(async …)` (surfaced by Drizzle as `db.$client.transaction`), which rolls back SQL writes across `await`. Drizzle's own `db.transaction()` on DO is `ctx.storage.transactionSync` and cannot span an `await`, so it is deliberately not used. There is no Drizzle transaction handle on DO — the storage transaction is ambient on the object — so the tx-scoped backend binds the outer `db`. ```ts await ctx.storage.transaction(async () => { const txStore = store.withTransaction(db); await txStore.nodes.Document.update(documentId, props); await db.insert(documentVersions).values(versionRow); await db.insert(changeEvents).values(eventRow); }); // one storage-transaction COMMIT / ROLLBACK across both layers ``` - A latent detection bug is fixed: drizzle's Durable Objects session class is `SQLiteDOSession` (not the previously-checked `SQLiteDurableObjectSession`), so a real `drizzle(ctx.storage)` store was misclassified. - New `TransactionContext.sql` — the raw Drizzle handle bound to the same transaction — for graph-owned cross-store writes across **all** transactional backends (Postgres, libsql, better-sqlite3, do-sqlite): ```ts await store.transaction(async (tx) => { await tx.nodes.Document.update(documentId, props); // tx.sql is the AdoptedTransaction union — cast to your concrete // Drizzle database type at the call site. const sqlTx = tx.sql as NodePgDatabase; await sqlTx.insert(documentVersions).values(versionRow); await sqlTx.insert(changeEvents).values(eventRow); }); ``` This is the graph-owned counterpart of `store.withTransaction` (where the caller owns the boundary). On Postgres/libsql it is a correctness requirement — the outer `db` would write on a different connection and escape the transaction. `tx.sql` is `undefined` only on the non-transactional fallback. Its static type is the `AdoptedTransaction` union; cast to your concrete Drizzle database type at the call site. **Guarantees.** Building on [#135](https://github.com/nicia-ai/typegraph/issues/135), no schema/bootstrap/fulltext DDL ever runs inside the business transaction: `bootstrapTables` and the durable materialization marker run outside any storage transaction, while the schema-version commit uses the `do-sqlite` runner (data only). Boot the parent store via `createStoreWithSchema` once at object startup. **Out of scope.** Cloudflare D1 stays `transactionMode: "none"`: `D1Database.batch(...)` is transactional but not an interactive runner. A batch-only D1 mode is tracked separately. - [#138](https://github.com/nicia-ai/typegraph/pull/138) [`bcf1e48`](https://github.com/nicia-ai/typegraph/commit/bcf1e4819754f1839a236d350d70bab9103607ce) Thanks [@pdlug](https://github.com/pdlug)! - Durable, enforced fulltext materialization ([#135](https://github.com/nicia-ai/typegraph/issues/135)). Strategy-owned fulltext table/index DDL was materialized lazily, guarded by an **in-memory, per-backend-instance boolean latch** (`fulltextEnsured`), and interleaved into the read/write data path. That was correct only by accident (idempotent DDL + a warm process) and at the wrong durability scope; it was inconsistent with how vector indexes are tracked and it blocked cross-store transaction adoption ([#134](https://github.com/nicia-ai/typegraph/issues/134)). "Is this graph's fulltext storage materialized?" is now a **durable, queryable database fact** instead of a process boolean. **Breaking (behavioral): fulltext now requires an explicit boot step.** `createStore()` is a synchronous, zero-I/O _attach_ — it never creates tables, repairs DDL, or writes materialization markers. The durable marker is written exclusively by the async boot path, `createStoreWithSchema(graph, backend)`, which must run once at application startup (outside request handlers and adopted transactions). A fulltext read/write — or a transaction that touches fulltext — against a database with no valid marker now throws the new `StoreNotInitializedError` instead of lazily emitting DDL on the hot path. Consumers already using `createStoreWithSchema` need no changes; consumers relying on lazy fulltext creation via bare `createStore()` must add a `createStoreWithSchema` call at boot. **What ships:** - New `@nicia-ai/typegraph` exports: `StoreNotInitializedError` and the `StoreNotInitializedReason` (`"missing" | "stale" | "failed"`) it carries in `details.reason`. - New per-deployment table `typegraph_contribution_materializations`, a sibling of `typegraph_index_materializations` (the declared-index status table is deliberately left unchanged). Keyed by [#129](https://github.com/nicia-ai/typegraph/issues/129) contribution identity `(graph_id, logical_name, owner, table_name)`; `signature` is a separate content-hash column, so a same-identity row with a drifted signature is a loud error, never a silent re-materialize. Failed re-attempts preserve the prior success timestamp via the same COALESCE rule as index materializations. - New backend primitives (SQLite + Postgres): `ensureContributionMaterializationsTable`, `getContributionMaterialization`, `recordContributionMaterialization`, and `assertRuntimeContributionsInitialized`. `ensureRuntimeContributions` and `ensureFulltextTable` now take a `graphId` and route through the durable-marker writer (short-circuiting when the recorded signature already matches). `createStoreWithSchema` records the marker after the schema version is resolved, covering the cold-initialize path. - The six fulltext-touching methods (`upsertFulltext`, `deleteFulltext`, `upsertFulltextBatch`, `deleteFulltextBatch`, `fulltextSearch`, `hardDeleteNode`) stop ensuring and instead assert the durable marker (resolved once per backend instance, cached). The transaction path performs zero DDL: the tx-scoped backend's fulltext methods assert the cached marker at point of use (a `SELECT`, never `CREATE`), so a transaction that never touches fulltext requires no fulltext initialization and one that does runs pure DML on the adopted transaction. This makes [#134](https://github.com/nicia-ai/typegraph/issues/134) (cross-store transaction adoption) sound by construction: a transaction-adopting primitive consults the durable fact and refuses with a clear `StoreNotInitializedError` if the store was never initialized, instead of emitting `CREATE INDEX` inside the caller's business transaction. - [#136](https://github.com/nicia-ai/typegraph/pull/136) [`9aa2d31`](https://github.com/nicia-ai/typegraph/commit/9aa2d31b8beddbf8f0dea08c4d9435ab3255b580) Thanks [@pdlug](https://github.com/pdlug)! - Unified `TableContribution` contract for strategy-owned tables ([#129](https://github.com/nicia-ai/typegraph/issues/129)). "What tables does TypeGraph own?" was previously split across four uncoordinated surfaces (Drizzle named exports, tables-factory recursion, strategy raw DDL, per-table `ensureXTable` methods). Adding a new strategy- or backend-owned table without also wiring an `ensureXTable` + bootstrap probe re-opened the gap [#128](https://github.com/nicia-ai/typegraph/issues/128) closed. This refactor routes every owned table through one shape. **Breaking (custom `FulltextStrategy` implementers only):** `FulltextStrategy.generateDdl(tableName): string[]` is replaced by `ownedTables(primaryTableName): readonly StrategyTableContribution[]`. A strategy now _declares_ its tables, Drizzle-free, as already authoritative contributions (`logicalName`, `owner`, resolved `tableName`, idempotent `createDdl` for the table **and its supporting indexes**, `runtimeEnsure`). The two shipped strategies (`tsvectorStrategy`, `fts5Strategy`) and all internal callers are migrated; consumers using only the shipped strategies need no changes. **What ships:** - New `@nicia-ai/typegraph` export: `TableContribution` and `StrategyTableContribution` (its strategy-declaration alias). Each contribution carries a stable, deployment-independent `logicalName` plus the resolved physical `tableName` (distinct identity vs. drift-signature inputs) — the prerequisite that lets [#135](https://github.com/nicia-ai/typegraph/issues/135) make fulltext materialization a durable, decidable fact instead of an in-memory per-backend latch. - `postgresContributions()` / `sqliteContributions()` are the single source of truth for DDL generation and the bootstrap ensure. `generatePostgresDDL` / `generateSqliteDDL` iterate contributions; the `table === tables.fulltext` reference-identity hack is gone from DDL generation. drizzle-kit visibility for the default Postgres strategy comes from the schema barrel exporting the matching `tables.fulltext` object (one object, not two); a non-default strategy exports its own. - New backend method `ensureRuntimeContributions()`, which runs each `runtimeEnsure` contribution's full idempotent `createDdl` (table + supporting indexes) so a partial state (table present, index missing) self-heals — not a probe-and-skip. `loadActiveSchemaWithBootstrap` calls it scoped to `runtimeEnsure` contributions only (the strategy-owned fulltext table today), so startup does not regress into broad DDL/probing across every table. `ensureFulltextTable` is retained as a thin back-compat wrapper. DDL statement ordering changes from "all CREATE TABLE, then all CREATE INDEX, then fulltext" to per-contribution "table then its own indexes". Safe because TypeGraph's tables carry no cross-table foreign keys; raw migration SQL byte output differs accordingly. Prerequisite for [#135](https://github.com/nicia-ai/typegraph/issues/135) (durable fulltext materialization), which is in turn the prerequisite for [#134](https://github.com/nicia-ai/typegraph/issues/134) (cross-store transaction adoption). ## 0.25.1 ### Patch Changes - [#130](https://github.com/nicia-ai/typegraph/pull/130) [`dbe52dc`](https://github.com/nicia-ai/typegraph/commit/dbe52dc5d1346543b5aab5b4380df85bdbf66750) Thanks [@pdlug](https://github.com/pdlug)! - Fix drizzle-kit-managed fulltext bootstrap gap on both Postgres and SQLite ([#128](https://github.com/nicia-ai/typegraph/issues/128)). Consumers managing typegraph storage via `drizzle-kit push` / `drizzle-kit generate` (`export * from "@nicia-ai/typegraph/postgres"` or `…/sqlite"`) got every typegraph table EXCEPT `typegraph_node_fulltext`. The fulltext table was strategy-owned raw DDL — the schema modules exposed only `fulltextTableName: string`, not a Drizzle table — so drizzle-kit silently skipped it. The `bootstrapTables` fallback in `loadActiveSchemaWithBootstrap` only fires on a missing-table error from `getActiveSchema`; once drizzle-kit had created `typegraph_schema_versions`, that branch stopped triggering and `searchable()` writes failed at runtime with `relation/table "typegraph_node_fulltext" does not exist`. Two fixes ship together: - **`backend.ensureFulltextTable()` (both backends).** A focused narrow-ensure that mirrors the existing `ensureIndexMaterializationsTable` / `ensureKindRemovalsTable` / `ensureReconciliationMarkersTable` idiom — single-table `CREATE … IF NOT EXISTS`, no Postgres SHARE-lock deadlock under concurrent replica startup. The backend wraps every method that emits fulltext SQL (`upsertFulltext` / `deleteFulltext` and their batch variants, `fulltextSearch`, and `hardDeleteNode` whose cascade unconditionally deletes from the fulltext table) to call the ensure first. A per-backend latch makes the per-call cost a single boolean check after the first invocation, so the wrapping is safe on the hot path. `loadActiveSchemaWithBootstrap` also calls the ensure as a belt-and-suspenders for the `createStoreWithSchema` path. Together these cover both async schema-aware boot AND the sync `createStore` path — the bare bootstrap-load probe alone would miss the latter. This is the canonical fix and the **only** viable one for SQLite (FTS5 virtual tables aren't drizzle-kit-modelable). - **Typed Drizzle pg-core table for `tsvectorStrategy` (Postgres only).** `createPostgresTables()` now returns `tables.fulltext` — a typed `pgTable` for the default `tsvector` + GIN stack — alongside `tables.fulltextTableName`. The new `fulltext` named export is included in `@nicia-ai/typegraph/postgres`, so `export *` lets drizzle-kit generate migrations for the fulltext table the same way it does for `nodes`/`edges`/etc. Custom `tsvector`/`regconfig` column types are exported alongside the existing `vector` column. `generatePostgresDDL` deliberately skips the typed Drizzle table (the column-walker can't reproduce the `GENERATED ALWAYS AS (…) STORED`clause) and continues to defer to `tsvectorStrategy.generateDdl()` for the runtime DDL emit. The two paths agree byte-for-byte; a drift sentinel test catches any divergence. Alternate Postgres fulltext strategies (pg_trgm, ParadeDB, pgroonga) still own their own DDL via `FulltextStrategy.generateDdl()` and the bootstrap probe runs it. Drizzle-kit consumers using a non-default strategy must override `tables.fulltext` in their schema barrel with their strategy's own table. Documented the SQLite FTS5 virtual-table caveat and the new Postgres `tables.fulltext` export in `apps/docs/src/content/docs/integration.md`. ## 0.25.0 ### Highlights 0.25.0 is the runtime schema evolution release. It adds graph extensions, unified index declarations and materialization, dynamic queries over runtime-declared kinds, runtime access to compiled props schemas, and a safer transactional schema-version commit path. - Graph extensions let applications commit reviewed JSON schema proposals as durable TypeGraph schema versions without redeploying application code. - Compile-time, graph-extension, relational, and vector indexes now share one canonical declaration channel and flow through `Store.materializeIndexes()`. - Dynamic query builder methods let typed queries traverse runtime-declared node and edge kinds while still validating kind names, endpoints, and field predicates at query-build time. - `Store` now exposes compiled Zod props schemas for compile-time and graph-extension kinds through `getNodePropsSchema`, `getEdgePropsSchema`, and their `OrThrow` variants. - Node and edge definitions now accept JSON-serializable `annotations` for consumer-owned metadata such as UI hints, audit policy, and provenance. ### Upgrade notes - Existing deployments with manually managed schemas should add the one-active schema-version partial unique index: `typegraph_schema_versions_one_active_per_graph_idx` on `(graph_id)` where `is_active` is true (`TRUE` on Postgres, `1` on SQLite). - Manually managed schemas should also sync the generated DDL for the new TypeGraph status tables, including `typegraph_index_materializations`, `typegraph_kind_removals`, and `typegraph_reconciliation_markers`. - Run schema migrations from a transactional backend. Edge or HTTP-only non-transactional drivers can continue serving normal reads and writes after the schema is established. - Tests that deep-compare the full `SchemaValidationResult` object may need to switch to partial matching because `initialized` and `migrated` now include `committedRow`. #### Custom backends and index consumers These changes affect custom `GraphBackend` implementations and advanced index consumers; ordinary `createStoreWithSchema`, query, and collection callers should not need code changes. - `insertSchema` and `setActiveSchema` were removed from `GraphBackend`. Implement `commitSchemaVersion` and `setActiveVersion` instead. - `commitSchemaVersion` and `setActiveVersion` require transactional behavior. Non-transactional drivers such as Cloudflare D1, Durable Objects, `drizzle-orm/neon-http`, and SQLite backends configured with `transactionMode: "none"` refuse these primitives for schema commits. - `createFulltextIndex` and `dropFulltextIndex` were removed from `GraphBackend`; fulltext storage remains owned by the active backend fulltext strategy. - The old `NodeIndex`, `EdgeIndex`, and `TypeGraphIndex` types were removed from `@nicia-ai/typegraph/indexes`. Use `NodeIndexDeclaration`, `EdgeIndexDeclaration`, or `IndexDeclaration`. - Custom backends should add the new optional materialization/removal primitives when they want first-class support for index status loading, removal reconciliation markers, and vector index materialization. ### Minor Changes #### New APIs - `defineGraphExtension(input)` and `validateGraphExtension(input, options?)`. - `Store.evolve`, `Store.deprecateKinds`, `Store.undeprecateKinds`, `Store.removeKinds`, `Store.materializeRemovals`, and dynamic collection accessors for graph-extension kinds. - `defineGraph({ indexes })`, `defineNodeIndex`, `defineEdgeIndex`, `andWhere`, `orWhere`, `notWhere`, and the `@nicia-ai/typegraph/indexes` subpath for advanced index tooling. - `Store.materializeIndexes(options?)` plus `MaterializeIndexesResult` status reporting. - `embedding(dimensions, options?)` vector index options and exported vector index declaration/configuration types. - `fromDynamic`, `traverseDynamic`, `optionalTraverseDynamic`, and `toDynamic` on the query builder. - `SchemaValidationResult.initialized` and `.migrated` now include `committedRow: SchemaVersionRow`. - `SqlTableNames` now includes `uniques` so cleanup paths can honor custom physical table names. #### Performance and reliability - Schema commits now use a transactional `commitSchemaVersion` backend primitive instead of the old insert-then-activate sequence, fixing the orphan schema-row crash window. - `materializeIndexes` bulk-loads materialization status in one round trip and records per-index drift/failure state in `typegraph_index_materializations`. - `materializeRemovals` records a reconciliation watermark, honors custom table names, and cleans secondary embedding/fulltext/unique rows for removed node kinds. - Schema hash and parsed-schema caches avoid repeated serialization, SHA-256, and Zod parse work on no-change startup and repeated store creation. - Graph-extension merge/compile paths share caches and fast paths for idempotent or partially overlapping evolves. - Postgres vector-index drops now run per-metric DDL concurrently. #### Pull requests - [#103](https://github.com/nicia-ai/typegraph/pull/103) - Add per-kind `annotations`. - [#106](https://github.com/nicia-ai/typegraph/pull/106) - Add atomic schema version commits. - [#107](https://github.com/nicia-ai/typegraph/pull/107) - Add compile-time index declarations to graph definitions and serialized schemas. - [#112](https://github.com/nicia-ai/typegraph/pull/112) - Add `Store.materializeIndexes`. - [#117](https://github.com/nicia-ai/typegraph/pull/117) - Unify vector indexes with the index declaration channel. - [#118](https://github.com/nicia-ai/typegraph/pull/118) - Add graph extensions. - [#125](https://github.com/nicia-ai/typegraph/pull/125) - Add dynamic query traversal methods. - [#126](https://github.com/nicia-ai/typegraph/pull/126) - Expose runtime Zod props schemas. - [#127](https://github.com/nicia-ai/typegraph/pull/127) - Pre-release cleanup and performance pass. ## 0.24.1 ### Patch Changes - [#99](https://github.com/nicia-ai/typegraph/pull/99) [`755df5a`](https://github.com/nicia-ai/typegraph/commit/755df5a8d8114fbc72047f436132bfe105d02823) Thanks [@pdlug](https://github.com/pdlug)! - Internal: dependency bump pass (patch/minor only — TypeScript and `@types/node` held back as separate majors). Notable runtime/peer-relevant moves: `nanoid` 5.1.9 → 5.1.11 (only published runtime dep); dev/peer `zod` 4.3.6 → 4.4.3, `@libsql/client` 0.17.2 → 0.17.3. Also drops the `export` keyword on 14 types that were never reachable through any public entry point (`src/index.ts`, `./schema`, `./indexes`, `./sqlite`, `./postgres`, etc.) and had no internal importers. These were leaked-internal types surfaced by a sensitivity change in `knip` 6.11. No symbol on the documented API surface changed; consumers importing only via the package's declared `exports` paths are unaffected. ## 0.24.0 ### Minor Changes - [#97](https://github.com/nicia-ai/typegraph/pull/97) [`8747df8`](https://github.com/nicia-ai/typegraph/commit/8747df8c003589f985e86ca654cf796fa5230e34) Thanks [@pdlug](https://github.com/pdlug)! - SQLite: implement `backend.vectorSearch`, unblocking `store.search.hybrid()` on SQLite. The hybrid retrieval facade has been Postgres-only since [#88](https://github.com/nicia-ai/typegraph/issues/88): SQLite shipped fulltext (`fulltextSearch`) and embedding persistence (`upsertEmbedding` / `deleteEmbedding`), but never the `vectorSearch` method that `executeHybridSearch` requires for RRF fusion. `.similarTo()` on SQLite still worked because the predicate path goes through the query compiler, not the backend facade — but anyone reaching for `store.search.hybrid()` on SQLite hit `ConfigurationError: Backend does not support vector search`. This release wires up the SQLite half of that contract: - `buildVectorSearchSqlite` issues `vec_distance_cosine` / `vec_distance_l2` against the embeddings BLOB column, mirroring the Postgres SQL shape (same WHERE / ORDER BY / score expression / minScore semantics). - `createSqliteBackend` exposes `vectorSearch` on the backend object whenever `hasVectorEmbeddings` is true (parallel to the existing `upsertEmbedding` gate). - `inner_product` is rejected — sqlite-vec has no `vec_distance_ip` function. ```typescript import { createLocalSqliteBackend } from "@nicia-ai/typegraph/sqlite/local"; const { backend } = createLocalSqliteBackend(); // sqlite-vec auto-loaded const store = createStore(graph, backend); const ranked = await store.search.hybrid("Document", { limit: 10, vector: { fieldPath: "embedding", queryEmbedding }, fulltext: { query: "climate adaptation" }, }); ``` **Performance.** On the standard search-shapes bench (500 docs, 384-dim), SQLite hybrid clocks in at **0.8ms** — about 3× faster than PostgreSQL's 2.5ms on the same shape. The bench harness now measures it on both backends; the previously-blank SQLite cell in the search comparison table is filled in. ## 0.23.0 ### Minor Changes - [#95](https://github.com/nicia-ai/typegraph/pull/95) [`6f3bf30`](https://github.com/nicia-ai/typegraph/commit/6f3bf30b4ac7c51a5528e1001dc97e05146801b7) Thanks [@pdlug](https://github.com/pdlug)! - PostgreSQL: official postgres-js / Neon support, server-side prepared statements on the fast path, and a `refreshStatistics()` API. **Four drivers supported.** `createPostgresBackend` has always been driver-agnostic, but only `node-postgres` was covered in CI. This release adds: - **`drizzle-orm/postgres-js`** — full adapter + integration suite coverage (~250 tests run against both `pg` and `postgres-js` against a real PostgreSQL). - **`drizzle-orm/neon-serverless`** — `@neondatabase/serverless` Pool over WebSockets. Wiring smoke tests verify driver detection, fast-path routing, Date→string normalization, and capability surface; the shared code paths are exercised by the `pg` integration suite since this driver is pg-Pool-protocol-compatible. - **`drizzle-orm/neon-http`** — `@neondatabase/serverless` `neon(url)` over HTTP. Auto-detected so `capabilities.transactions` is set to `false` (HTTP can't hold a session); single-statement reads, writes, and migrations work normally. Smoke tests verify the detection and capability override. Same `createPostgresBackend(db)` entry point regardless of driver. ```typescript // postgres-js import postgres from "postgres"; import { drizzle } from "drizzle-orm/postgres-js"; const backend = createPostgresBackend( drizzle(postgres(process.env.DATABASE_URL)), ); // Neon serverless (edge runtimes) import { Pool } from "@neondatabase/serverless"; import { drizzle } from "drizzle-orm/neon-serverless"; const backend = createPostgresBackend( drizzle(new Pool({ connectionString: env.NEON_DATABASE_URL })), ); ``` **On Neon HTTP vs WebSockets:** both work. The HTTP driver (`drizzle-orm/neon-http`) is best for stateless edge workloads — TypeGraph auto-disables transactions since HTTP can't hold a session, and `store.transaction(...)` falls through to non-transactional sequential execution. Use the WebSocket driver (`drizzle-orm/neon-serverless`) when you need atomic multi-statement writes. **~6× faster on multi-hop traversals via server-side prepared statements.** The execution adapter now uses `node-postgres`'s named prepared statements transparently — each unique compiled SQL string gets a stable counter-derived statement name (cached by SQL text), so PostgreSQL caches the plan after first execution. Combined with routing `execute()` through the fast path directly (skipping Drizzle's session wrapper), this drops the 3-hop benchmark from ~7.5ms to ~0.8ms median, putting TypeGraph-on-PostgreSQL at parity with Neo4j on every single-query and multi-hop shape we measure. The change is invisible to callers; existing code keeps working. postgres-js is unchanged (it handles its own preparation internally). **New `store.refreshStatistics()` / `backend.refreshStatistics()` API.** Call once after a large initial import or bulk backfill. Without fresh stats, the planner can pick suboptimal execution plans — on PostgreSQL this is the difference between a 0.5ms and 5ms forward traversal; on SQLite it's the difference between 0.9ms and 23ms fulltext search. Autovacuum / background statistics catch up eventually, but explicit invocation gives correct latencies immediately. ```typescript for (const batch of batches) { await store.nodes.Document.bulkCreate(batch); } await store.refreshStatistics(); ``` Implementations: SQLite runs `ANALYZE`; PostgreSQL runs `ANALYZE` on TypeGraph-managed tables only. Costs ~20ms on SQLite, ~80ms on PostgreSQL at the sizes this library is designed for. **Type surface changes:** - `GraphBackend` now requires a `refreshStatistics(): Promise` method. `TransactionBackend` still excludes it (statistics refresh isn't meaningful inside a transaction). External `GraphBackend` implementations (uncommon) need to add a no-op or proper implementation. - `PostgresBackendOptions` adds an optional `capabilities?: Partial` for users who need to override capability flags (e.g., for custom HTTP-style drivers). - `PostgresBackendOptions` also adds `prepareStatements?: boolean` (default `true`) and `preparedStatementCacheMax?: number` (default `256`). The prepared-statement name cache is now LRU-bounded so high-cardinality SQL text doesn't grow unbounded in either the Node process or in PostgreSQL's per-session prepared-statement memory. Set `prepareStatements: false` when pooling through pgbouncer in transaction-pool mode. See [`backend-setup`](https://typegraph.dev/backend-setup#choosing-a-postgresql-driver) for the runtime-to-driver matrix, per-driver setup snippets, and post-bulk-load guidance. ## 0.22.0 ### Minor Changes - [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Add `countEdges(edgeAlias)` and `countDistinctEdges(edgeAlias)` — edge-count aggregators that skip the target-node join in the count aggregate fast path. The default `count(targetAlias)` counts edges whose target node is currently live under the query's temporal mode, which requires joining the edges to the target node table on every aggregation. For the common "how many follow relationships does this user have?" question, that join is unnecessary work: you want to count edges, not reach through each edge to validate the target. ```typescript import { count, countEdges, field } from "@nicia-ai/typegraph"; const result = await store .query() .from("User", "u") .optionalTraverse("follows", "e", { expand: "none" }) .to("User", "target") .groupByNode("u") .aggregate({ name: field("u", "name"), // Counts live edges, regardless of target-node validity. // Skips the typegraph_nodes join entirely — ~1.7x faster on // SQLite, ~1.35x on PostgreSQL at benchmark scale. followCount: countEdges("e"), // Counts edges to live targets. Keeps the target-node join // so the target's temporal window is honored. liveFollowCount: count("target"), }) .execute(); ``` **When to use which:** - `count(targetAlias)` — when the semantic question is "how many of this user's follows point to a live user?" The target-node join enforces the target's `validTo` / `deleted_at` filters. - `countEdges(edgeAlias)` — when the semantic question is "how many follow relationships does this user have?" The edge's own temporal and deletion filters are enforced; target validity is not consulted. - `countDistinctEdges(edgeAlias)` — same semantics as `countEdges` but with `COUNT(DISTINCT ...)`. Useful under ontology-driven expansions where the same edge can appear multiple times in join output. The two can be mixed in one aggregate. When present together, the compiler keeps the target-node join but switches it to a `LEFT JOIN` with node-side filters pushed into the `ON` clause so edge counts reflect all live edges while node counts only reflect edges to live targets. No change to existing `count(...)` behavior. This is purely additive — code that currently uses `count("targetAlias")` continues to count live targets exactly as before. ### Patch Changes - [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Push `LIMIT` past `GROUP BY` in the count aggregate fast path when it's safe. When `groupByNode(...).aggregate({ x: count(alias) })` is paired with an optional traversal and a `.limit(n)` that doesn't depend on the aggregate (no `ORDER BY`, or an `ORDER BY` restricted to group keys), the compiler now emits the `LIMIT` inside the start CTE. The `GROUP BY` runs over `n` rows instead of the full start set — `O(limit)` grouping work instead of `O(|start|)`. When `OFFSET` is also set, it rides along with the `LIMIT` into the start CTE and the outer `SELECT` drops its own `LIMIT`/`OFFSET` so neither clause is double-applied. The fast path also picks `INNER JOIN` over `LEFT JOIN` for the target-node join whenever a `whereNode()` predicate applies to the target alias, so those predicates constrain every aggregate — including `countEdges(...)`. `LEFT JOIN` remains the strategy when only temporal/delete filters apply to the target, so `countEdges` and `count(target)` can coexist in one query with divergent semantics. No change to query semantics — aggregate counts still reflect the same `count(target)` as before, including the target node's temporal and deletion filters. No change to aggregate queries without a `LIMIT`. No change on SQLite or PostgreSQL query shapes outside the fast path. Measured impact: scopes down group-by work for "top-N by count"-style aggregate queries. No impact on the blog-post benchmark's full-graph aggregate (which measures the ungrouped 1,200-user case and intentionally runs without a `LIMIT`). - [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Fix `generateSqliteDDL` and `generatePostgresMigrationSQL` emitting `(unknown, unknown, ...)` for indexes threaded through `createSqliteTables({}, { indexes })` or `createPostgresTables({}, { indexes })`. The DDL generator's SQL-chunk flattener didn't handle two cases that appear inside index expression keys: Drizzle column references nested inside a SQL stream (whose `.getSQL()` wraps the column back inside a self-referential SQL object, causing the previous logic to recurse and fall through to `"unknown"`), and `StringChunk` values stored as single-element arrays (`[""]`). Expression indexes now emit correctly in both dialects, e.g. ```sql CREATE INDEX IF NOT EXISTS "idx_tg_node_user_city_cov_name_…" ON "typegraph_nodes" ("graph_id", "kind", (json_extract("props", '$."city"')), (json_extract("props", '$."name"'))); ``` Added a regression test in `tests/indexes.test.ts` asserting that DDL from `createSqliteTables`/`createPostgresTables` never contains `(unknown` and includes the expected column and `json_extract` / `ARRAY['…']` expressions. - [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Emit `NOT MATERIALIZED` on PostgreSQL traversal and start CTEs so the planner can inline them and see their inner row statistics. PostgreSQL defaults to materializing any CTE referenced more than once. TypeGraph's traversal compilation references each CTE twice — once from the next hop's join, once from the final SELECT — which triggers materialization under the default rules. Materialized CTEs have opaque statistics to the planner, causing poor join orderings and wildly off row estimates on multi-hop queries over larger graphs. Introduces a `emitNotMaterializedHint` dialect capability (`true` for PostgreSQL, `false` for SQLite, which ignores the hint entirely) and threads it through the start-CTE and traversal-CTE emitters. The hint matches what an expert would write by hand for the same query shape. Impact on the TypeGraph benchmark suite: - Multi-hop traversal plans no longer carry opaque materializations, so the planner picks index-scan orderings appropriate to the starting row's selectivity. - No visible change on SQLite (the hint is not emitted). - Guards against regressions on larger graphs where materialized CTE plans degenerate into cross-product-plus-filter. - [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Persist vector embeddings on the SQLite backend when sqlite-vec is loaded. Previously, `store.nodes.X.create({ ..., embedding: [...] })` on SQLite validated the embedding and inserted the node, but the embedding itself was silently dropped — the SQLite backend didn't implement `upsertEmbedding`/`deleteEmbedding`, so the store's embedding-sync path quietly no-op'd. Vector predicates like `d.embedding.similarTo(q, 20, { metric: "cosine" })` then ran against an empty `typegraph_node_embeddings` table and returned zero rows without error. This release wires up both methods on the SQLite backend. They encode embeddings to `vec_f32('[...]')` BLOBs on write and rely on sqlite-vec at query time — same storage shape the existing `.similarTo()` compilation already targets. Activation is opt-in via a new `hasVectorEmbeddings` option on `createSqliteBackend` so callers that haven't loaded sqlite-vec don't hit `no such function: vec_f32` at write time. `createLocalSqliteBackend` best-effort-loads sqlite-vec at startup and flips the option automatically, so the common local setup works without configuration. ```typescript // Local backend: sqlite-vec is loaded automatically when installed. const { backend } = createLocalSqliteBackend(); // BYO drizzle connection: pass hasVectorEmbeddings after loading sqlite-vec. import sqliteVec from "sqlite-vec"; sqliteVec.load(sqlite); const backend = createSqliteBackend(drizzle(sqlite), { tables, hasVectorEmbeddings: true, }); ``` `getEmbedding` and the hybrid-search facade (`store.search.hybrid(...)`) remain PostgreSQL-only — decoding the raw BLOB back to `number[]` via `vec_to_json` and exposing a hybrid-search backend method are tracked separately. ## 0.21.0 ### Highlights TypeGraph 0.21 adds full-text search and hybrid vector/text retrieval. Mark string fields with `searchable()` and TypeGraph maintains native PostgreSQL tsvector/GIN or SQLite FTS5 storage. The node-level `$fulltext.matches()` predicate composes with property filters, graph traversal, and vector similarity; `store.search` provides full-text and hybrid helpers with reciprocal-rank fusion and optional snippets. Full-text behavior is supplied by a pluggable `FulltextStrategy`, with query-mode, language, prefix, and highlighting capabilities checked against the active strategy. Existing graph data can be indexed through the paginated `rebuildFulltext()` maintenance API. ### Upgrade notes - Node and edge property names beginning with `$` are now reserved for query accessors. Rename those fields before upgrading; graph definition rejects them with `ConfigurationError`. - Use `store.search.rebuildFulltext()` to index existing data after declaring searchable fields. It reports skipped invalid properties and is a maintenance operation rather than a transactionally frozen scan of the whole graph. - Query options must be supported by the configured strategy. SQLite FTS5 cannot honor per-query language overrides; unsupported modes, overrides, and snippet requests are refused. - `findNodesByKind` now breaks equal creation-time ties by ID. Callers should not rely on the previously unspecified order. ### Minor Changes - [#88](https://github.com/nicia-ai/typegraph/pull/88) [`6f681d5`](https://github.com/nicia-ai/typegraph/commit/6f681d59f16ef7d7651627999cce6cada01d024e) Thanks [@pdlug](https://github.com/pdlug)! - Add fulltext search and hybrid (vector + fulltext) retrieval. Declare `searchable()` string fields on any node schema and TypeGraph keeps a native FTS index in sync — `tsvector` + GIN on PostgreSQL, FTS5 on SQLite. Query it through a node-level `n.$fulltext.matches()` predicate that composes with metadata filters, graph traversal, and vector similarity in one SQL statement. ```typescript import { defineNode, searchable, embedding } from "@nicia-ai/typegraph"; const Document = defineNode("Document", { schema: z.object({ title: searchable({ language: "english" }), body: searchable({ language: "english" }), tenantId: z.string(), embedding: embedding(1536), }), }); // Fulltext + metadata filter in a single query const results = await store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext.matches("climate change", 20).and(d.tenantId.eq(tenant)), ) .select((ctx) => ctx.d) .execute(); // Hybrid: vector + fulltext fused with Reciprocal Rank Fusion at the SQL layer const hybrid = await store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext .matches("climate", 50) .and(d.embedding.similarTo(queryVector, 50)) .and(d.tenantId.eq(tenant)), ) .select((ctx) => ctx.d) .limit(10) .execute(); // Store-level helper with tunable RRF weights and snippets const tuned = await store.search.hybrid("Document", { limit: 10, vector: { fieldPath: "embedding", queryEmbedding: queryVector }, fulltext: { query: "climate change", includeSnippets: true }, fusion: { method: "rrf", k: 60, weights: { vector: 1, fulltext: 1.5 } }, }); ``` Query modes cover `websearch` (Google-style syntax — default), `phrase`, `plain`, and `raw` (dialect-native tsquery / FTS5 MATCH). Highlighting via `ts_headline` / `snippet()` is opt-in per query. No extensions required: Postgres uses the built-in `tsvector` + GIN (works on every managed provider); SQLite uses FTS5 which is statically linked into the standard `better-sqlite3` / `libsql` / `bun:sqlite` distributions. See `/fulltext-search` for the full guide. ### Added - `n.$fulltext` — node-level fulltext accessor; `.matches(query, k?, options?)` composes against the combined `searchable()` content. `$fulltext` is exposed on every `NodeAccessor`; a runtime guard throws a clear error if the node kind has no `searchable()` fields. `k` defaults to 50. - `store.search` facade — `store.search.fulltext()`, `store.search.hybrid()`, and `store.search.rebuildFulltext()` grouped under one namespace. Lazy-initialized and cached on first access. - `FulltextSearchHit`, `VectorSearchHit`, and `HybridSearchHit` are generic over the node type (`FulltextSearchHit`). `store.search.fulltext("Document", ...)` returns hits with `hit.node` narrowed to the Document node shape — no cast required. - `backend.upsertFulltextBatch` + `backend.deleteFulltextBatch` — symmetric batched fulltext primitives. Homogeneous batch shape, duplicate-nodeId dedupe last-write-wins, per-row fallback when unset. - `store.search.rebuildFulltext(nodeKind?, { pageSize?, maxSkippedIds? })` — rebuilds the fulltext index from existing node data using keyset pagination on `id` (stable under shared timestamps and light concurrent writes). Transacts per page; cleans stale rows for soft-deleted nodes; validates `pageSize` as a positive integer; counts corrupt / non-object props as `skipped` and surfaces offending IDs via `skippedIds` without aborting. `maxSkippedIds` (default 10,000) lets operators investigating systemic corruption collect the full list. Concurrent hard-deletes between pages may be missed — document as maintenance operation. - Keyset pagination on `findNodesByKind` via new `{ orderBy, after }` params. - `QueryBuilder.fuseWith({ k?, weights? })` — tunable RRF on the query-builder path. Flat `HybridFusionOptions` shape, identical to `store.search.hybrid`'s `fusion` option. Throws at compile time if the query lacks either a `.similarTo()` or `n.$fulltext.matches()`. Shares its validator with `store.search.hybrid({ fusion })` so `method`, `k`, and per-source weights are checked identically on both paths. - `FulltextStrategy` — pluggable abstraction (exported from the top-level entry) that owns the **entire** SQL pipeline for a dialect's fulltext support: DDL, upsert (single + batch), delete (single + batch), MATCH condition, rank expression, and snippet expression. Ships `tsvectorStrategy` (Postgres built-in `tsvector`) and `fts5Strategy` (SQLite FTS5); dialect adapters expose `fulltext: FulltextStrategy | undefined`. Alternate Postgres stacks (pg_trgm, ParadeDB / pg_search, pgroonga) choose their own column layout, index type, and projection — TypeGraph's operation layer just delegates to the active strategy. Strategies declare prefix-query support explicitly via `FulltextStrategy.supportsPrefix`, so capability discovery stays correct for strategies that support prefix matching via dedicated syntax without advertising raw-mode pass-through. - Backend-level fulltext strategy override: `createPostgresBackend(db, { fulltext })` and `createSqliteBackend(db, { fulltext })` accept a `FulltextStrategy` that takes precedence over the dialect default. Threaded through to compiler passes, backend-direct search SQL, all write SQL, DDL generation, and capability discovery — so a ParadeDB-backed Postgres `store.search.hybrid()` fuses the same way a tsvector-backed one does, without any call-site changes. - Option validation: `store.search.fulltext` and `store.search.hybrid` validate caller options against the active `FulltextStrategy` (falling back to `BackendCapabilities.fulltext.{phraseQueries, highlighting, languages}` when no strategy is attached). A `mode` outside `strategy.supportedModes` throws, `includeSnippets: true` on a strategy whose `supportsSnippets` is false throws, and a per-query `language` override on a strategy whose `supportsLanguageOverride` is false (e.g. SQLite FTS5) throws. Advisory warning for unknown languages on strategies that honor overrides. `$fulltext.matches()` is validated against the dialect strategy's `supportedModes` at compile time. - One-time `console.warn` when a node kind has multiple `searchable()` fields with conflicting `language` values. The first field's language wins on the stored row; the warning makes the silent collapse visible so users know to split multilingual content across dedicated node kinds. - Snippet highlighting uses `…` consistently across both shipped strategies (`ts_headline` on Postgres, `snippet()` on SQLite). One stylesheet applies everywhere. - `FulltextSearchResult.score` is always `number`. The Postgres adapter coerces `numeric`-as-string driver returns at the backend boundary so downstream code never sees a union type. - Hybrid SQL emitter uses a deterministic `COALESCE(fulltext.node_id, embeddings.node_id) ASC` tiebreak, matching the JS-side `localeCompare(nodeId)` tiebreak used by `store.search.hybrid` — both hybrid paths produce identical top-k under RRF score ties. - Postgres fulltext table schema: `language` is `regconfig` (not `TEXT`) and `tsv` is a `GENERATED ALWAYS AS (to_tsvector("language", "content")) STORED` column. Postgres owns the `content / language → tsv` invariant; the strategy's write SQL doesn't recompute `tsv` inline. The `content` column is populated verbatim, and the per-query `language` override path still accepts a text parameter (cast to `regconfig` at query time). SQLite's FTS5 virtual table is unchanged. ### Changed - **`defineNode()` / `defineEdge()` reject `$`-prefixed property names.** The `$` namespace is reserved for node-level accessors (starting with `$fulltext`). A `ConfigurationError` is raised at graph-definition time instead of silently shadowing user fields at query time. Rename any such fields before upgrading. - **`findNodesByKind` offset pagination now has a deterministic tiebreaker** (`ORDER BY created_at DESC, id DESC`). Row order was previously under-specified when `created_at` values collided; callers that happened to rely on an implementation-dependent order may see different tie-breaking. ## 0.20.0 ### Minor Changes - [#85](https://github.com/nicia-ai/typegraph/pull/85) [`12055d0`](https://github.com/nicia-ai/typegraph/commit/12055d053b22cfadd1439c9a667307fae77af6a2) Thanks [@pdlug](https://github.com/pdlug)! - Add Tier 1 graph algorithms on `store.algorithms.*`: `shortestPath`, `reachable`, `canReach`, `neighbors`, and `degree`. ```typescript // Find the shortest path through a set of edge kinds const path = await store.algorithms.shortestPath(alice, bob, { edges: ["knows"], maxHops: 6, }); // Enumerate reachable nodes within a depth bound const reachable = await store.algorithms.reachable(alice, { edges: ["knows"], maxHops: 3, }); // Fast existence check const connected = await store.algorithms.canReach(alice, bob, { edges: ["knows"], }); // k-hop neighborhood (source always excluded) const twoHop = await store.algorithms.neighbors(alice, { edges: ["knows"], depth: 2, }); // Count incident edges const total = await store.algorithms.degree(alice, { edges: ["knows"] }); ``` All traversal algorithms compile to a single recursive-CTE query and share the dialect primitives used by `.recursive()` and `store.subgraph()`, so SQLite and PostgreSQL yield identical semantics. Node arguments accept either a raw ID string or any object with an `id` field — `Node`, `NodeRef`, and the lightweight records returned by the algorithms themselves all work. See `/graph-algorithms` for the full reference. - [#85](https://github.com/nicia-ai/typegraph/pull/85) [`12055d0`](https://github.com/nicia-ai/typegraph/commit/12055d053b22cfadd1439c9a667307fae77af6a2) Thanks [@pdlug](https://github.com/pdlug)! - Graph algorithms (`store.algorithms.*`) and `store.subgraph()` now honor the store's temporal model. **New:** Every algorithm and `store.subgraph()` accept `temporalMode` and `asOf` options, matching the shape already used by `store.query()` and collection reads. When neither is supplied, the resolved mode falls back to `graph.defaults.temporalMode` (typically `"current"`). ```typescript // Snapshot at a point in time await store.algorithms.shortestPath(alice, bob, { edges: ["knows"], temporalMode: "asOf", asOf: "2023-01-15T00:00:00.000Z", }); await store.subgraph(rootId, { edges: ["has_task"], temporalMode: "includeEnded", }); ``` The filter applies to both nodes and edges along the traversal, is orthogonal to `cyclePolicy`, and is honored by the shortest-path self-path short-circuit. **BREAKING:** `store.subgraph()` previously ignored graph temporal settings and filtered only by `deleted_at IS NULL` (equivalent to `"includeEnded"`). It now defaults to `graph.defaults.temporalMode`. Callers that relied on walking through validity-ended rows must pass `temporalMode: "includeEnded"` explicitly. Soft-delete filtering is unchanged under the default `"current"` mode, so most callers see no difference. ### Patch Changes - [#87](https://github.com/nicia-ai/typegraph/pull/87) [`f52bba6`](https://github.com/nicia-ai/typegraph/commit/f52bba63befe8111d13d04cfb9659371f7061625) Thanks [@pdlug](https://github.com/pdlug)! - Fix SQLite temporal filter timestamp format in graph algorithms and subgraph. `buildReachableCte`, `resolveTemporalFilter`, and `fetchSubgraphEdges` compiled temporal filters without passing `dialect.currentTimestamp()`, so on SQLite they fell back to raw `CURRENT_TIMESTAMP` (`YYYY-MM-DD HH:MM:SS`). Stored `valid_from` / `valid_to` use ISO-8601 (`YYYY-MM-DDTHH:MM:SS.sssZ`), and because `T` sorts above space, same-day ISO timestamps compare incorrectly against raw `CURRENT_TIMESTAMP`. Under `temporalMode: "current"` this caused `reachable` / `canReach` / `neighbors` / `shortestPath` / `degree` and the `subgraph` edge hydration to misclassify rows whose `valid_from` or `valid_to` fell on today's date, disagreeing with `store.query()` and collection reads. All three call sites now inject the dialect-specific current timestamp (`strftime('%Y-%m-%dT%H:%M:%fZ','now')` on SQLite, `NOW()` on PostgreSQL), matching the query compiler. ## 0.19.0 ### Minor Changes - [#83](https://github.com/nicia-ai/typegraph/pull/83) [`206f464`](https://github.com/nicia-ai/typegraph/commit/206f46467342eee6a060c83e057bbf1befb31c1a) Thanks [@pdlug](https://github.com/pdlug)! - **BREAKING:** `store.subgraph()` now returns an indexed result instead of flat arrays. The result shape changes from `{ nodes: Node[], edges: Edge[] }` to: ```typescript { root: Node | undefined; nodes: ReadonlyMap; adjacency: ReadonlyMap>; reverseAdjacency: ReadonlyMap>; } ``` This eliminates the indexing boilerplate every consumer had to write before traversing the subgraph. Nodes are keyed by ID for O(1) lookup, and edges are organized into forward/reverse adjacency maps keyed by `nodeId → edgeKind`. Migration: - `result.nodes` is now a `Map` — use `.size` instead of `.length`, `.values()` instead of direct iteration, `.has(id)` / `.get(id)` instead of `.find()` - `result.edges` is removed — access edges via `result.adjacency.get(fromId)?.get(edgeKind)` or `result.reverseAdjacency.get(toId)?.get(edgeKind)` - `result.root` provides the root node directly (no lookup needed) ## 0.18.0 ### Minor Changes - [#80](https://github.com/nicia-ai/typegraph/pull/80) [`0845fa9`](https://github.com/nicia-ai/typegraph/commit/0845fa92a653ed107057cf350414e13745fff8d8) Thanks [@pdlug](https://github.com/pdlug)! - Add first-class libsql backend at `@nicia-ai/typegraph/sqlite/libsql` ### New convenience export `createLibsqlBackend(client, options?)` wraps `@libsql/client` with automatic DDL execution and correct async execution profile. The caller retains ownership of the client, enabling shared-driver setups. Works with local files, in-memory databases, and remote Turso URLs. ```typescript import { createClient } from "@libsql/client"; import { createLibsqlBackend } from "@nicia-ai/typegraph/sqlite/libsql"; const client = createClient({ url: "file:app.db" }); const { backend, db } = await createLibsqlBackend(client); const store = createStore(graph, backend); ``` ### Bug fixes for async SQLite drivers - **`db.get()` crash on empty results** — switched to `db.all()[0]` to work around Drizzle's `normalizeRow` crash when libsql returns no rows ([drizzle-team/drizzle-orm#1049](https://github.com/drizzle-team/drizzle-orm/issues/1049)) - **`instanceof Promise` check fails for Drizzle thenables** — all SQLite exec helpers now use unconditional `await` since Drizzle returns `SQLiteRaw` objects that are thenable but not `Promise` instances ([drizzle-team/drizzle-orm#2275](https://github.com/drizzle-team/drizzle-orm/issues/2275)) ### Internal improvements - Extracted `wrapWithManagedClose()` helper for idempotent backend close with teardown - Shared adapter and integration test suites now accept async backend factories - libsql backend runs the full shared test suite (214 tests) ## 0.17.0 ### Minor Changes - [#77](https://github.com/nicia-ai/typegraph/pull/77) [`b9fc057`](https://github.com/nicia-ai/typegraph/commit/b9fc057e0dd62bd0f059bb78a20d18d91b1b87be) Thanks [@pdlug](https://github.com/pdlug)! - feat: support orderBy on edge properties in query builder The `orderBy` method now accepts edge aliases in addition to node aliases, allowing results to be ordered by properties on traversed edges. This eliminates the need to denormalize ordering fields onto nodes or sort in memory. ```typescript store .query() .from("Person", "p") .traverse("worksAt", "e") .to("Company", "c") .orderBy("e", "salary", "asc") // order by edge property .select((ctx) => ({ name: ctx.p.name, salary: ctx.e.salary })) .execute(); ``` Also fixes CTE alias resolution for edge aliases in `groupBy` and vector order-by compilation paths. Closes [#76](https://github.com/nicia-ai/typegraph/issues/76) ## 0.16.2 ### Patch Changes - [#73](https://github.com/nicia-ai/typegraph/pull/73) [`1c95d8e`](https://github.com/nicia-ai/typegraph/commit/1c95d8ec641442cecb38e00fab4c6d10eb162c2c) Thanks [@pdlug](https://github.com/pdlug)! - fix: dispose serialized execution queue on backend close to prevent unhandled rejections When the SQLite backend's underlying database is destroyed while operations are still queued (e.g., during Cloudflare Workers test teardown), the serialized execution queue now properly disposes pending promises. Calling `backend.close()` signals the queue to suppress errors from in-flight tasks and reject new operations with `BackendDisposedError`. Fixes [#72](https://github.com/nicia-ai/typegraph/issues/72) ## 0.16.1 ### Patch Changes - [#70](https://github.com/nicia-ai/typegraph/pull/70) [`cebf681`](https://github.com/nicia-ai/typegraph/commit/cebf681c76820db9d63c29f2eb64ed92b1eb3ad5) Thanks [@pdlug](https://github.com/pdlug)! - Widen ID parameters on `DynamicNodeCollection` and `DynamicEdgeCollection` to accept plain `string` instead of branded `NodeId`/`EdgeId` types, removing the need for casts when using the dynamic collection API with IDs from edge metadata, snapshots, or external input. ## 0.16.0 ### Minor Changes - [#66](https://github.com/nicia-ai/typegraph/pull/66) [`2f241a9`](https://github.com/nicia-ai/typegraph/commit/2f241a98fc6ec78702bcaa609e1fce9b5a1ae4f4) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.getNodeCollection(kind)` and `store.getEdgeCollection(kind)` methods for runtime string-keyed collection access. Returns the full collection API with widened generics (`DynamicNodeCollection` / `DynamicEdgeCollection`), or `undefined` if the kind is not registered. Eliminates the need for `Reflect.get(store.nodes, kind) as SomeType` patterns when iterating kinds, resolving nodes from edge metadata, or building generic graph tooling like snapshots and summaries. ## 0.15.0 ### Minor Changes - [#63](https://github.com/nicia-ai/typegraph/pull/63) [`546a7eb`](https://github.com/nicia-ai/typegraph/commit/546a7eb3693141fa8ad236c9aad3333abf635893) Thanks [@pdlug](https://github.com/pdlug)! - `createStoreWithSchema()` now auto-creates base tables on a fresh database. Previously, calling it against a database without pre-existing TypeGraph tables (e.g. a new Cloudflare Durable Object) would throw a raw "no such table" error. The function now detects missing tables and bootstraps them automatically via the new optional `bootstrapTables` method on `GraphBackend`. Both SQLite and PostgreSQL backends implement this method. `createStore()` remains unchanged for users who manage DDL manually. - [#64](https://github.com/nicia-ai/typegraph/pull/64) [`6b84b42`](https://github.com/nicia-ai/typegraph/commit/6b84b42bd9e626ca01f48d8a5bd3c18c5bfee80d) Thanks [@pdlug](https://github.com/pdlug)! - Add `StoreProjection` utility type for typing reusable helpers that work across graphs sharing a common subgraph. The type projects a store's collection surface onto a subset of node and edge keys, with node constraint names erased so that graphs registering the same node types with different unique constraints remain cross-assignable. Both `Store` and `TransactionContext` are structurally assignable to any `StoreProjection` whose keys are a subset of `G`. Also exports `GraphNodeCollections` and `GraphEdgeCollections` shared mapped types. ### Patch Changes - [#59](https://github.com/nicia-ai/typegraph/pull/59) [`36742a1`](https://github.com/nicia-ai/typegraph/commit/36742a11f47b2e1903c13ce6abce3e72285f0dbf) Thanks [@pdlug](https://github.com/pdlug)! - Reject empty `fields` arrays at the type level in `defineNodeIndex` and `defineEdgeIndex`. Previously, passing `fields: []` was accepted by TypeScript but threw at runtime. The `fields` property now requires a non-empty tuple, surfacing the error at compile time. - [#60](https://github.com/nicia-ai/typegraph/pull/60) [`dca5aba`](https://github.com/nicia-ai/typegraph/commit/dca5abad98cdb4df0ca546796f89c6470bdcf680) Thanks [@pdlug](https://github.com/pdlug)! - Export `SchemaValidationResult` and `SchemaManagerOptions` types from the root package entry point so users can type the return value of `createStoreWithSchema()` without reaching into internal subpaths. ## 0.14.0 ### Minor Changes - [#54](https://github.com/nicia-ai/typegraph/pull/54) [`bf6997a`](https://github.com/nicia-ai/typegraph/commit/bf6997afd5889556961977f45bdc9c8d38021902) Thanks [@pdlug](https://github.com/pdlug)! - ### Breaking: default recursive traversal depth lowered from 100 to 10 Unbounded `.recursive()` traversals are now capped at 10 hops instead of 100. Graphs with branching factor _B_ produce O(_B_^depth) rows before cycle detection can prune them — the previous default of 100 made exponential blowup easy to trigger accidentally. If your traversals relied on the implicit 100-hop cap, add an explicit `.maxHops(100)` call. The `MAX_EXPLICIT_RECURSIVE_DEPTH` ceiling (1000) is unchanged. ### Schema parse validation Serialized schema documents read from the database are now validated against a Zod schema at the parse boundary. Malformed, truncated, or incompatible schema documents will throw a `DatabaseOperationError` with path-level detail instead of propagating silently. Enum fields (`temporalMode`, `cardinality`, `deleteBehavior`, etc.) are validated against the known literal unions. ### Type safety improvements - Added `useUnknownInCatchVariables`, `noFallthroughCasesInSwitch`, and `noImplicitReturns` to tsconfig - Drizzle row mappers now use runtime type checks (`asString`/`asNumber`) instead of unsafe `as` casts - `NodeMeta` and `EdgeMeta` are now derived from row types via mapped types - All non-null assertions (`!`) eliminated from source code - Hardcoded constants extracted to shared `constants.ts` - Duplicate `fnv1aBase36` function consolidated into `utils/hash.ts` ## 0.13.0 ### Minor Changes - [#52](https://github.com/nicia-ai/typegraph/pull/52) [`1e3da4a`](https://github.com/nicia-ai/typegraph/commit/1e3da4aa814f3baf67a0cb54c9c753508eecf0f0) Thanks [@pdlug](https://github.com/pdlug)! - Add `batchFindFrom`, `batchFindTo`, and `batchFindByEndpoints` to edge collections for use with `store.batch()`. Edge collection lookup methods (`findFrom`, `findTo`, `findByEndpoints`) execute immediately and cannot participate in `store.batch()`. The new `batchFind*` variants return a `BatchableQuery` instead, enabling edge lookups to share a single transactional connection alongside fluent queries. ```typescript const [skills, employer, colleague] = await store.batch( store.edges.hasSkill.batchFindFrom(alice), store.edges.worksAt.batchFindFrom(alice), store.edges.knows.batchFindByEndpoints(alice, bob), ); ``` - **`batchFindFrom(from)`** — deferred variant of `findFrom` - **`batchFindTo(to)`** — deferred variant of `findTo` - **`batchFindByEndpoints(from, to, options?)`** — deferred variant of `findByEndpoints`, returns 0-or-1 element array All three preserve the same endpoint type constraints as their immediate counterparts. Closes [#51](https://github.com/nicia-ai/typegraph/issues/51). ## 0.12.0 ### Minor Changes - [#50](https://github.com/nicia-ai/typegraph/pull/50) [`a59416d`](https://github.com/nicia-ai/typegraph/commit/a59416d8cbc641fd7611ee5d5b0fb115aea59450) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.batch()` for executing multiple queries over a single connection with snapshot consistency. - **Single connection**: Acquires one connection via an implicit transaction, eliminating pool pressure from parallel `Promise.all` patterns (N connections → 1). - **Snapshot consistency**: All queries see the same database state — no interleaved writes between results. - **Typed tuple results**: Returns a mapped tuple preserving each query's independent result type, projection, filtering, sorting, and pagination. > **Correction (see #325).** The "snapshot consistency" bullet above was never > accurate and is retained only as the historical record. `batch()` opens its > implicit transaction without an isolation option, so PostgreSQL runs it at the > default read-committed isolation and a later query in the batch _can_ observe a > commit the earlier ones did not. The "single connection" bullet describes the > transactional path; connection reuse is otherwise the adapter's business, not a > consequence of `capabilities.transactions`. `batch()` also never pipelined, > despite the original issue specifying it. - **`BatchableQuery` interface**: Satisfied by both `ExecutableQuery` (from `.select()`) and `UnionableQuery` (from set operations like `.union()`, `.intersect()`). Exposes `executeOn()` for backend-delegated execution. - **Minimum 2 queries**: Enforced at the type level — single queries should use `.execute()` directly. ```typescript const [people, companies] = await store.batch( store .query() .from("Person", "p") .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })), store .query() .from("Company", "c") .select((ctx) => ({ id: ctx.c.id, name: ctx.c.name })) .orderBy("c", "name", "asc") .limit(5), ); // people: readonly { id: string; name: string }[] // companies: readonly { id: string; name: string }[] ``` Closes [#47](https://github.com/nicia-ai/typegraph/issues/47). - [#48](https://github.com/nicia-ai/typegraph/pull/48) [`753d9eb`](https://github.com/nicia-ai/typegraph/commit/753d9ebc6aa02f0f01bc52abc1de255b2d1bbd91) Thanks [@pdlug](https://github.com/pdlug)! - Add field-level projection to `store.subgraph()` via a declarative `project` option. - **Declarative field selection**: Specify which properties to keep per node/edge kind. Projected nodes always retain `kind` and `id`; projected edges always retain structural endpoint fields. Kinds omitted from `project` remain fully hydrated. - **SQL-level extraction**: Projected property fields are extracted via `json_extract()` / JSONB path expressions directly in the query, avoiding full `props` blob transfer for projected kinds. - **All-or-nothing metadata**: Include `"meta"` in the field list for the full metadata object, or omit it entirely. No partial metadata selection — the struct is small enough that subsetting adds complexity without meaningful savings. - **`defineSubgraphProject()` helper**: Curried identity function that preserves literal types for reusable projection configs. Without it, storing a projection in a variable widens field arrays to `string[]`, defeating compile-time narrowing. - **Type-safe results**: Result types narrow per-kind based on the projection — accessing omitted fields is a compile-time error. Works through both inline literals and `defineSubgraphProject()`. ```typescript const result = await store.subgraph(rootId, { edges: ["has_task", "uses_skill"], maxDepth: 2, project: { nodes: { Task: ["title", "meta"], Skill: ["name"], }, edges: { uses_skill: ["priority"], }, }, }); // result.nodes — Task has { kind, id, title, meta }; Skill has { kind, id, name } // result.edges — uses_skill has { id, kind, fromKind, fromId, toKind, toId, priority } ``` Closes [#46](https://github.com/nicia-ai/typegraph/issues/46) (alternative implementation — declarative arrays instead of callbacks). ## 0.11.1 ### Patch Changes - [#41](https://github.com/nicia-ai/typegraph/pull/41) [`68d5432`](https://github.com/nicia-ai/typegraph/commit/68d5432f830978bc05f888134ed1a69644ed97b9) Thanks [@pdlug](https://github.com/pdlug)! - Fix `.paginate()` dropping `id` from selective query results and `orderBy()` mishandling system fields. - **Fix silent data loss in `.paginate()` + `.select()`**: `FieldAccessTracker.record()` no longer allows a system field (`id`, `kind`) to be downgraded to a props field, which caused the SQL projection to extract from `props->>'id'` (nonexistent) instead of the `id` column. - **Fix `orderBy()` for system fields**: `orderBy("alias", "id")` now emits `ORDER BY cte.alias_id` instead of `ORDER BY json_extract(cte.alias_props, '$.id')`. - **Add `gt`/`gte`/`lt`/`lte` to `StringFieldAccessor`**: Enables keyset cursor pagination via `whereNode("a", (a) => a.id.lt(cursor))`. Fixes [#40](https://github.com/nicia-ai/typegraph/issues/40). ## 0.11.0 ### Minor Changes - [#38](https://github.com/nicia-ai/typegraph/pull/38) [`e26e4a5`](https://github.com/nicia-ai/typegraph/commit/e26e4a5282d9e59ab517a68dede37c38bea2a1e9) Thanks [@pdlug](https://github.com/pdlug)! - Add `createFromRecord()` and `upsertByIdFromRecord()` to `NodeCollection`. These methods accept `Record` instead of `z.input`, providing an escape hatch for dynamic-data scenarios (changesets, migrations, imports) where the data shape is determined at runtime. Runtime Zod validation is unchanged — only the compile-time type gate is relaxed. The return type remains fully typed as `Node`. Closes [#37](https://github.com/nicia-ai/typegraph/issues/37). ## 0.10.0 ### Minor Changes - [#33](https://github.com/nicia-ai/typegraph/pull/33) [`da14806`](https://github.com/nicia-ai/typegraph/commit/da14806b665418c7761b5db37641b23eb2914304) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.subgraph()` for typed BFS neighborhood extraction from a root node. Given a root node ID, traverses specified edge kinds using a recursive CTE and returns all reachable nodes and connecting edges as fully typed discriminated unions. **Options:** - `edges` — edge kinds to traverse (required) - `maxDepth` — maximum traversal depth (default: 10) - `direction` — `"out"` (default) or `"both"` for undirected traversal - `includeKinds` — filter returned nodes to specific kinds (traversal still follows all reachable nodes) - `excludeRoot` — omit the root node from results - `cyclePolicy` — cycle detection strategy (default: `"prevent"`) **Type utilities exported:** - `AnyNode` / `AnyEdge` — discriminated unions of all node/edge runtime types in a graph - `SubsetNode` / `SubsetEdge` — narrowed unions for a subset of kinds - `SubgraphOptions` / `SubgraphResult` — fully generic option and result types - [#35](https://github.com/nicia-ai/typegraph/pull/35) [`0ebc59c`](https://github.com/nicia-ai/typegraph/commit/0ebc59cf1f8d714b0d63c0759d08ed88face022c) Thanks [@pdlug](https://github.com/pdlug)! - Add runtime discriminated union types: `AnyNode`, `AnyEdge`, `SubsetNode`, `SubsetEdge`. These pure type-level utilities produce discriminated unions of runtime node/edge instances from a graph definition. Unlike `AllNodeTypes` (union of type _definitions_), `AnyNode` gives the union of runtime `Node` values — discriminated by `kind` for exhaustive `switch` narrowing. `SubsetNode` narrows the union to a specific set of kinds. ## 0.9.2 ### Patch Changes - [#27](https://github.com/nicia-ai/typegraph/pull/27) [`c2f0811`](https://github.com/nicia-ai/typegraph/commit/c2f0811863a61608c16901ce1fc61fdfbc26cb3f) Thanks [@pdlug](https://github.com/pdlug)! - Fix `count(alias, field)` and `countDistinct(alias, field)` ignoring the field argument in SQL compilation. Both functions always compiled to `COUNT(alias_id)` / `COUNT(DISTINCT alias_id)` regardless of the field argument, because: 1. The aggregate emitters in `standard-builders.ts` and `set-operations.ts` hardcoded `_id` for count/countDistinct instead of calling `compileFieldValue()` like sum/avg/min/max do. 2. `collectRequiredColumnsByAlias` in `standard-pass-pipeline.ts` explicitly skipped marking the field as required for count/countDistinct, so the CTE wouldn't include the `_props` column even if the emitter were fixed. Now `count("p", "email")` correctly compiles to `COUNT(json_extract(p_props, '$."email"'))` and `countDistinct("b", "genre")` compiles to `COUNT(DISTINCT json_extract(b_props, '$."genre"'))`. ## 0.9.1 ### Patch Changes - [#24](https://github.com/nicia-ai/typegraph/pull/24) [`733bf8a`](https://github.com/nicia-ai/typegraph/commit/733bf8abfd7b0fa9901a08ff67ce1c9343a2e961) Thanks [@pdlug](https://github.com/pdlug)! - Fix `checkUniqueBatch` exceeding SQL bind parameter limit on SQLite/D1/Durable Objects. Bulk constraint operations (`bulkGetOrCreateByConstraint`, `bulkFindByConstraint`) passed all keys in a single `IN (...)` clause. With hundreds of unique keys, this exceeded SQLite's 999 bind parameter limit, causing `SQLITE_ERROR: too many SQL variables`. The fix chunks the keys array in `checkUniqueBatch` using the same pattern already used by `getNodes`, `insertNodesBatch`, and other batch operations. SQLite chunks at 996 keys per query (999 max − 3 fixed params), PostgreSQL at 65,532. ## 0.9.0 ### Minor Changes - [#21](https://github.com/nicia-ai/typegraph/pull/21) [`88beee4`](https://github.com/nicia-ai/typegraph/commit/88beee42ce0ecfe2064b0b3889653e889b0c74aa) Thanks [@pdlug](https://github.com/pdlug)! - Add `transactionMode` to SQLite execution profile, fixing Cloudflare Durable Object compatibility. `createSqliteBackend` previously used raw `BEGIN`/`COMMIT`/`ROLLBACK` SQL for all sync SQLite drivers. This crashes on Cloudflare Durable Object SQLite (via `drizzle-orm/durable-sqlite`) because the driver does not support raw transaction SQL through `db.run()`. The new `transactionMode` option (`"sql"` | `"drizzle"` | `"none"`) controls how transactions are managed: - `"sql"` — TypeGraph issues `BEGIN`/`COMMIT`/`ROLLBACK` directly (default for better-sqlite3, bun:sqlite) - `"drizzle"` — delegates to Drizzle's `db.transaction()` (default for async drivers) - `"none"` — transactions disabled (default for D1 and Durable Objects) D1 and Durable Object sessions are auto-detected by Drizzle session name. Users can override via `executionProfile: { transactionMode: "..." }`. **Breaking:** `isD1` removed from `SqliteExecutionProfileHints` and `SqliteExecutionProfile`. Use `transactionMode: "none"` instead. `D1_CAPABILITIES` removed — capabilities are now derived from `transactionMode`. ## 0.8.0 ### Minor Changes - [#19](https://github.com/nicia-ai/typegraph/pull/19) [`5b1dec6`](https://github.com/nicia-ai/typegraph/commit/5b1dec64f280a2ec638c69b6fa5a1bc08ba92e88) Thanks [@pdlug](https://github.com/pdlug)! - Support unconstrained edges in `defineGraph`. Edges defined without `from`/`to` constraints (e.g., `defineEdge("sameAs")`) can now be passed directly to `defineGraph` without an `EdgeRegistration` wrapper. They are automatically allowed to connect any node type in the graph to any other. - **`EdgeEntry` widened** — accepts any `EdgeType`, not just those with endpoints - **`NormalizedEdges`** — falls back to all graph node types when `from`/`to` are undefined - Constrained edges, `EdgeRegistration` wrappers, and narrowing validation are unchanged ## 0.7.0 ### Minor Changes - [#16](https://github.com/nicia-ai/typegraph/pull/16) [`0a2f08f`](https://github.com/nicia-ai/typegraph/commit/0a2f08fa7d755ee6adb59db4d34a26a3863c0c79) Thanks [@pdlug](https://github.com/pdlug)! - Tighten type safety across store and collection APIs. **Breaking:** `TypedNodeRef` has been renamed to `NodeRef` and the old untyped `NodeRef` has been removed. Replace `TypedNodeRef` with `NodeRef` — the type is structurally identical. Unparameterized `NodeRef` (with the new default) covers the old untyped usage. - **`EdgeId`** — branded edge ID type, mirroring `NodeId`. Prevents mixing IDs from different edge types at compile time. - **`Edge`** — edge instances now carry endpoint node types. `edge.fromId` is `NodeId`, `edge.toId` is `NodeId`, and `edge.id` is `EdgeId`. - **`getNodeKinds` / `getEdgeKinds`** — return `readonly (keyof G["nodes"] & string)[]` instead of `readonly string[]`. - **`constraintName` literal unions** — `findByConstraint`, `getOrCreateByConstraint`, and their bulk variants now only accept constraint names that exist on the node registration, catching typos at compile time. ## 0.6.0 ### Minor Changes - [#14](https://github.com/nicia-ai/typegraph/pull/14) [`45624e0`](https://github.com/nicia-ai/typegraph/commit/45624e0ef5caf28c5a7bf8931f0ae96ce542c20d) Thanks [@pdlug](https://github.com/pdlug)! - Restructure SQLite/Postgres entry points to decouple DDL generation from native dependencies. **Breaking changes:** - `./drizzle`, `./drizzle/sqlite`, `./drizzle/postgres`, `./drizzle/schema/sqlite`, `./drizzle/schema/postgres` entry points are removed. Import backend factories, schema tables/factories, and DDL helpers from `./sqlite` and `./postgres`. - `createLocalSqliteBackend` moves from `./sqlite` to `./sqlite/local`. The `./sqlite` entry point no longer depends on `better-sqlite3`. - `getSqliteMigrationSQL` is renamed to `generateSqliteMigrationSQL`. - `getPostgresMigrationSQL` is renamed to `generatePostgresMigrationSQL`. - Individual table type aliases (`NodesTable`, `EdgesTable`, `UniquesTable`, `SchemaVersionsTable`, `EmbeddingsTable`) are removed from both schema modules. Use `SqliteTables["nodes"]` or `PostgresTables["edges"]` instead. **Migration guide:** | Before | After | | ------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------- | | `import { ... } from "@nicia-ai/typegraph/drizzle/sqlite"` | `import { ... } from "@nicia-ai/typegraph/sqlite"` | | `import { ... } from "@nicia-ai/typegraph/drizzle/postgres"` | `import { ... } from "@nicia-ai/typegraph/postgres"` | | `import { ... } from "@nicia-ai/typegraph/drizzle/schema/sqlite"` | `import { ... } from "@nicia-ai/typegraph/sqlite"` | | `import { ... } from "@nicia-ai/typegraph/drizzle/schema/postgres"` | `import { ... } from "@nicia-ai/typegraph/postgres"` | | `import { createLocalSqliteBackend } from "@nicia-ai/typegraph/sqlite"` | `import { createLocalSqliteBackend } from "@nicia-ai/typegraph/sqlite/local"` | | `getSqliteMigrationSQL()` | `generateSqliteMigrationSQL()` | | `getPostgresMigrationSQL()` | `generatePostgresMigrationSQL()` | | `NodesTable`, `EdgesTable`, `UniquesTable`, `SchemaVersionsTable`, `EmbeddingsTable` | `SqliteTables["nodes"]` / `PostgresTables["nodes"]` (and corresponding table keys) | ## 0.5.0 ### Minor Changes - [#12](https://github.com/nicia-ai/typegraph/pull/12) [`c40b8a4`](https://github.com/nicia-ai/typegraph/commit/c40b8a4c99f5ccddaf1bceea8c927f6aeb0300f4) Thanks [@pdlug](https://github.com/pdlug)! - Add read-only lookup methods and store-level clear for graph data management. **New APIs:** - `findByConstraint` / `bulkFindByConstraint` — look up nodes by a named uniqueness constraint without creating. Returns `Node | undefined` (or `(Node | undefined)[]` for bulk). Soft-deleted nodes are excluded. - `findByEndpoints` — look up an edge by `(from, to)` with optional `matchOn` property fields without creating. Returns `Edge | undefined`. Soft-deleted edges are excluded. - `store.clear()` — hard-delete all data for the current graph (nodes, edges, uniques, embeddings, schema versions). Resets collection caches so the store is immediately reusable with raw, unversioned semantics; reopen it through a managed factory before relying on schema-version fencing. ## 0.4.0 ### Minor Changes - [#10](https://github.com/nicia-ai/typegraph/pull/10) [`550eec6`](https://github.com/nicia-ai/typegraph/commit/550eec6bbe34427be9095fe59571b55f75c68792) Thanks [@pdlug](https://github.com/pdlug)! - Add node and edge get-or-create operations with explicit API naming. **New APIs:** - `getOrCreateByConstraint` / `bulkGetOrCreateByConstraint` — deduplicate nodes by a named uniqueness constraint - `getOrCreateByEndpoints` / `bulkGetOrCreateByEndpoints` — deduplicate edges by `(from, to)` with optional `matchOn` property fields - `hardDelete` for node and edge collections - `action: "created" | "found" | "updated" | "resurrected"` result discriminant **Breaking changes:** - `upsert` → `upsertById`, `bulkUpsert` → `bulkUpsertById` - `onConflict: "skip" | "update"` → `ifExists: "return" | "update"` - `ConstraintNotFoundError` → `NodeConstraintNotFoundError` - Removed generic `FindOrCreate*` type exports in favor of explicit `NodeGetOrCreateByConstraint*` and `EdgeGetOrCreateByEndpoints*` types ## 0.3.1 ### Patch Changes - [#8](https://github.com/nicia-ai/typegraph/pull/8) [`4732792`](https://github.com/nicia-ai/typegraph/commit/4732792a9ff7ed665f55bb314029c06024f5b62e) Thanks [@pdlug](https://github.com/pdlug)! - Fix `AnyPgDatabase` type to accept standard Drizzle instances created without an explicit schema ## 0.3.0 ### Minor Changes - [#6](https://github.com/nicia-ai/typegraph/pull/6) [`4553aed`](https://github.com/nicia-ai/typegraph/commit/4553aedf3cd7390acb7509e1c321a42bed225f1e) Thanks [@pdlug](https://github.com/pdlug)! - Big performance increases, cleaner APIs, prepared queries, and batch collection APIs. ### Breaking Changes **Renamed APIs:** - `selectAggregate()` is now `aggregate()` - `EdgeTypeNames` / `NodeTypeNames` are now `EdgeKinds` / `NodeKinds` (including getter functions) **Traversal expansion:** `includeImplyingEdges` replaced with `expand` option supporting four modes: `"none"`, `"implying"`, `"inverse"`, and `"all"` (default: `"inverse"`) **Recursive traversal:** The chained methods `.maxHops()`, `.minHops()`, `.collectPath()`, and `.withDepth()` are consolidated into a single `recursive()` call with an options object: ```ts // Before .traverse("p", "knows", "friend").recursive().maxHops(5).collectPath() // After .traverse("p", "knows", "friend").recursive({ maxHops: 5, path: true }) ``` New `cyclePolicy: "prevent" | "allow"` option (default: `"prevent"`). Unbounded recursion capped at depth 100; explicit `maxHops` validated up to 1,000. **Store:** `Store` class is now a type-only export — use `createStore()`. `StoreConfig` replaced by `StoreOptions`. **Moved to `@nicia-ai/typegraph/schema`:** All schema management APIs (`serializeSchema`, `deserializeSchema`, `initializeSchema`, `ensureSchema`, `migrateSchema`, `computeSchemaDiff`, `getMigrationActions`, `isBackwardsCompatible`, and related types) are now imported from the new `@nicia-ai/typegraph/schema` entry point. **Removed from main entry:** `KindRegistry`, Result utilities (`ok`/`err`/`isOk`/`isErr`/`unwrap`/`unwrapOr`), date helpers (`encodeDate`/`decodeDate`), validation utilities, and compiler/profiler internals. ### New Features **Prepared queries** — precompile queries once and execute repeatedly with different bindings at zero recompilation cost: ```ts const prepared = store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq(param("name"))) .select((ctx) => ctx.p) .prepare(); const alice = await prepared.execute({ name: "Alice" }); const bob = await prepared.execute({ name: "Bob" }); ``` **Batch collection APIs:** - `getByIds(ids)` — batched lookup preserving input order, returns `undefined` for missing IDs - `bulkInsert` — void-returning fire-and-forget ingestion - `bulkCreate` — multi-row `INSERT ... RETURNING` instead of per-item inserts - `bulkUpsert` (edges) — batch lookup instead of N+1 sequential calls **Node `find({ where })`** — filter nodes using the full query predicate system directly from collections. ### Performance - SQL compiler restructured into plan/passes/emitter pipeline with predicate pre-indexing, column pruning, and single-hop recursive lowering - Drizzle backend split into modular operations with dialect-driven strategy dispatch - SQLite prepared statement caching with LRU eviction - Compilation caching on immutable query builder instances - Bind-limit-aware batch chunking (SQLite: 999 params, PostgreSQL: 65,535 params) - Benchmark regression guardrails added to CI for both SQLite and PostgreSQL ## 0.2.0 ### Minor Changes - [`bdd5f34`](https://github.com/nicia-ai/typegraph/commit/bdd5f349453b19e9616f00d7591b436195feb925) Thanks [@pdlug](https://github.com/pdlug)! - Improve support for custom table names and use web crypto to support both node and edge runtimes. ## 0.1.1 ### Patch Changes - [`6f16bf9`](https://github.com/nicia-ai/typegraph/commit/6f16bf93ebd0811f386df63b80b8b80a3ee26c2f) Thanks [@pdlug](https://github.com/pdlug)! - Verify npmjs trusted publishing ## 0.1.0 ### Minor Changes - [`3d78324`](https://github.com/nicia-ai/typegraph/commit/3d78324472ac4cb4ac929b52c7501c08a5e7b6ca) Thanks [@pdlug](https://github.com/pdlug)! - Initial public release # Data Sync Patterns > Strategies for synchronizing external data with your TypeGraph store When adding TypeGraph as a graph overlay to an existing application, you need to keep your graph data in sync with your source of truth. This guide covers practical patterns for syncing external data into TypeGraph using the bulk operations API. ## Overview Most applications adding TypeGraph will have existing data in relational tables, external APIs, or document stores. Rather than migrating this data, TypeGraph works alongside it as an overlay that provides: - Graph traversals across your existing entities - Semantic search via vector embeddings - Relationship discovery and inference The key challenge is keeping the graph in sync with your source data. We cover three approaches: | Approach | Best For | Complexity | |----------|----------|------------| | [On-demand sync](#on-demand-sync) | Low-volume, real-time needs | Low | | [Batch sync](#batch-sync) | Bulk imports, periodic refresh | Medium | | [Event-driven sync](#event-driven-sync) | High-volume, near-real-time | Higher | ## Bulk Operations API TypeGraph provides bulk operations for efficient sync workflows: ```typescript // Create or update a single node await store.nodes.Document.upsertById(id, props); // Create many nodes at once await store.nodes.Document.bulkCreate(items); // Insert many nodes without returning results (dedicated fast path) await store.nodes.Document.bulkInsert(items); // Create or update many nodes at once await store.nodes.Document.bulkUpsertById(items); // Delete many nodes at once await store.nodes.Document.bulkDelete(ids); ``` ### upsertById Creates a node if it doesn't exist, or updates it if it does. This includes "un-deleting" soft-deleted nodes: ```typescript // First call creates the node const doc1 = await store.nodes.Document.upsertById("doc_123", { title: "Original Title", content: "...", }); // Second call updates the existing node const doc2 = await store.nodes.Document.upsertById("doc_123", { title: "Updated Title", content: "...", }); // doc1.id === doc2.id - same node, updated in place ``` To reopen a previously ended fact without changing its identity, use the explicit clear operation: ```typescript await store.nodes.Document.upsertById("doc_123", currentProps, { clearValidTo: true, }); ``` Omitting both end fields preserves the current end; `validTo` sets it; `clearValidTo: true` removes it. With `coalesceUnchangedUpserts`, replaying a clear against an already-open row skips the write, but capability validation still runs first. ### bulkCreate Efficiently creates multiple nodes in a single operation. Uses a single multi-row INSERT with RETURNING when the backend supports it: ```typescript const documents = await store.nodes.Document.bulkCreate([ { props: { title: "Doc 1", content: "..." } }, { props: { title: "Doc 2", content: "..." } }, { props: { title: "Doc 3", content: "..." }, id: "custom_id" }, ]); ``` If you only need write side effects and not created payloads, use `bulkInsert`. ### bulkInsert Inserts multiple nodes without returning results. This is the dedicated fast path for bulk ingestion — automatically wrapped in a transaction: ```typescript await store.nodes.Document.bulkInsert([ { props: { title: "Doc 1", content: "..." } }, { props: { title: "Doc 2", content: "..." } }, { props: { title: "Doc 3", content: "..." }, id: "custom_id" }, ]); ``` Prefer `bulkInsert` over `bulkCreate` when you don't need results. ### bulkUpsertById Creates or updates multiple nodes. Ideal for sync workflows where you don't know which records already exist: ```typescript // Sync a batch of external records const externalRecords = await fetchExternalData(); const synced = await store.nodes.Document.bulkUpsertById( externalRecords.map((record) => ({ id: record.id, // Use external ID as graph node ID props: { title: record.title, content: record.body, source: { table: "documents", id: record.id }, }, })) ); ``` A feed can deliver the same key twice in one page, so items are applied in order: the first item for an id creates or updates the row, and every later copy is an update over the value the earlier item wrote. The batch ends in the same state the equivalent sequence of `upsertById` calls would leave, whether or not the row existed before the batch. Because a later copy is an update, it merges over the earlier value — a field the last copy omits keeps what an earlier copy set. Edges follow the same rule, except that an update never repoints an edge: the endpoints of the first write stand. #### One batch cannot hand a unique value from one row to another Item order settles which value each id ends up with, but the writes themselves are grouped: every create in the batch runs before every update. A batch where one item **releases** a `unique` constraint value and a later item **claims** it therefore fails, even though the same operations applied one at a time succeed: ```typescript // "alice@example.com" is currently held by person_a. await store.nodes.Person.bulkUpsertById([ { id: "person_a", props: { email: "moved@example.com" } }, // releases it { id: "person_b", props: { email: "alice@example.com" } }, // claims it ]); // ✗ UniquenessError: person_b's create is checked while person_a still holds // the value. On a backend with transactions nothing is written at all. ``` Edges have the same limitation for a `cardinality` slot: a batch that ends the lone `oneActive` edge from a source while creating its replacement fails with a `CardinalityError`, because the replacement's create is checked before the update that frees the slot. This is a stated property of the bulk APIs rather than a bug to work around blindly. A batch describes the **set** of rows you want, not a script to reach them by, and grouping the creates apart from the updates is what makes a batch a couple of statements instead of one per item. The failure is loud and typed — never a silently dropped write. Two workarounds, both exact: - **Split the handoff across two batches**: one carrying the items that release the value, then one carrying the items that claim it. Reordering items *within* a single batch does not help, since the grouping ignores that order. This keeps bulk throughput for everything else. - **Apply the conflicting items as sequential `upsertById` calls** (for edges, as `update` then `create`), which frees the value before the claim is checked. If a sync feed can legitimately swap unique values between records, the two-batch shape is the reliable one: send one `bulkUpsertById` for the ids that already exist, then a second for the new ones. The interchange importer (`importGraph` with `onConflict: "update"`) is the other option that handles it directly — it applies each row in document order regardless of batch size, at the cost of the per-row validation and reporting that a bulk write skips. ### bulkReplaceById Use node `bulkReplaceById` when each source record is authoritative and should replace the stored document rather than merge with it: ```typescript await store.nodes.Document.bulkReplaceById( externalRecords.map((record) => ({ id: record.id, props: { title: record.title, content: record.body, }, })) ); ``` Omitted optional fields are removed. IDs must be distinct within the call; replacement is a set of final documents, not an ordered patch stream. On an eligible serverless backend this is the read-free sync path: TypeGraph submits creation, replacement, resurrection, claims, and search sidecars as one atomic program instead of reading every stored document before writing. ### bulkDelete Deletes multiple nodes by ID. Silently ignores IDs that don't exist: ```typescript // Remove nodes that no longer exist in source const deletedIds = await findDeletedRecords(); await store.nodes.Document.bulkDelete(deletedIds); ``` ### getOrCreate APIs Use get-or-create methods when your dedupe key is not a direct ID: ```typescript // Match by a named uniqueness constraint const byEmail = await store.nodes.User.getOrCreateByConstraint( "user_email", { email: "alice@example.com", name: "Alice" }, { ifExists: "update" } ); // byEmail.action: "created" | "found" | "updated" | "resurrected" // Match edges by endpoints (+ optional matchOn fields) const membership = await store.edges.memberOf.getOrCreateByEndpoints( user, org, { role: "admin", source: "sync" }, { matchOn: ["role"], ifExists: "update", validFrom: sourceMembership.startedAt, validTo: sourceMembership.endedAt, onImmutableLowerBound: "preserve", } ); // membership.action: "created" | "found" | "updated" | "resurrected" ``` For endpoint writes, `validFrom` applies when a new edge is created and on the `"resurrected"` branch, where it restates the revived row's whole window. With `onImmutableLowerBound: "preserve"`, an `"updated"` live edge keeps its stored lower bound while still applying props and `validTo`; the default `"refuse"` policy instead refuses a different stated start. A `"found"` result performs no write and preserves the existing validity window. To reopen an ended live edge, pass `clearValidTo: true` together with `ifExists: "update"`; the default return mode refuses that combination rather than ignoring the clear. An end that precedes the row's effective start is refused — see [Inverted validity windows](/errors/#inverted_validity_window). ### Edge Bulk Operations Edges also support bulk operations: ```typescript // Create many edges at once (returns created edges) const edges = await store.edges.relatedTo.bulkCreate([ { from: doc1, to: doc2, props: { confidence: 0.9 } }, { from: doc1, to: doc3, props: { confidence: 0.7 } }, { from: doc2, to: doc3, props: { confidence: 0.8 } }, ]); // Insert many edges without returning results (fast path) await store.edges.relatedTo.bulkInsert([ { from: doc1, to: doc2, props: { confidence: 0.9 } }, { from: doc1, to: doc3, props: { confidence: 0.7 } }, { from: doc2, to: doc3, props: { confidence: 0.8 } }, ]); // Delete many edges at once await store.edges.relatedTo.bulkDelete(edgeIds); ``` ## On-Demand Sync Sync individual records when they're accessed or modified. Best for low-volume scenarios where you want real-time consistency. ```typescript import { type Store } from "@nicia-ai/typegraph"; import { db, documents } from "./drizzle-schema"; interface AppDocument { id: string; title: string; content: string; updatedAt: Date; } async function syncDocument(store: Store, doc: AppDocument) { // Generate embedding for semantic search const embedding = await generateEmbedding(doc.content); // Upsert ensures we create or update as needed return store.nodes.Document.upsertById(doc.id, { title: doc.title, content: doc.content, embedding, source: { table: "documents", id: doc.id }, }); } // Sync on read - ensure graph is current before querying async function getRelatedDocuments(documentId: string) { // First, ensure the source document is synced const appDoc = await db.select().from(documents).where(eq(documents.id, documentId)).get(); if (!appDoc) throw new Error("Document not found"); await syncDocument(store, appDoc); // Now query the graph for relationships return store .query() .from("Document", "d") .whereNode("d", (d) => d.id.eq(documentId)) .traverse("relatedTo", "r") .to("Document", "related") .select((ctx) => ({ id: ctx.related.id, title: ctx.related.title, confidence: ctx.r.confidence, })) .execute(); } // Sync on write - update graph when source changes async function updateDocument(documentId: string, updates: Partial) { // Update source of truth first const [updated] = await db .update(documents) .set({ ...updates, updatedAt: new Date() }) .where(eq(documents.id, documentId)) .returning(); // Then sync to graph await syncDocument(store, updated); return updated; } ``` ## Batch Sync Process records in batches for bulk imports or periodic refresh. Best for large datasets or when you need to backfill data. ### Basic Batch Sync ```typescript interface SyncOptions { batchSize?: number; onProgress?: (processed: number, total: number) => void; } async function syncAllDocuments(store: Store, options: SyncOptions = {}) { const { batchSize = 100, onProgress } = options; // Get total count for progress reporting const [{ count }] = await db.select({ count: sql`count(*)` }).from(documents); let processed = 0; let offset = 0; while (offset < count) { // Fetch a batch from source const batch = await db.select().from(documents).limit(batchSize).offset(offset); if (batch.length === 0) break; // Generate embeddings in parallel (respecting API rate limits) const embeddings = await batchGenerateEmbeddings(batch.map((d) => d.content)); // Bulk upsert the batch await store.nodes.Document.bulkUpsertById( batch.map((doc, i) => ({ id: doc.id, props: { title: doc.title, content: doc.content, embedding: embeddings[i], source: { table: "documents", id: doc.id }, }, })) ); processed += batch.length; offset += batchSize; onProgress?.(processed, count); } return { processed, total: count }; } // Usage await syncAllDocuments(store, { batchSize: 50, onProgress: (processed, total) => { console.log(`Synced ${processed}/${total} documents`); }, }); ``` ### Incremental Sync Only sync records that have changed since the last sync: ```typescript interface SyncState { lastSyncAt: Date; } async function incrementalSync(store: Store, state: SyncState): Promise { const since = state.lastSyncAt; const now = new Date(); // Fetch only changed records const changed = await db .select() .from(documents) .where(gt(documents.updatedAt, since)) .orderBy(documents.updatedAt); if (changed.length > 0) { const embeddings = await batchGenerateEmbeddings(changed.map((d) => d.content)); await store.nodes.Document.bulkUpsertById( changed.map((doc, i) => ({ id: doc.id, props: { title: doc.title, content: doc.content, embedding: embeddings[i], source: { table: "documents", id: doc.id }, }, })) ); console.log(`Synced ${changed.length} changed documents`); } // Handle deletions (if your source tracks them) const deleted = await db .select({ id: documents.id }) .from(documents) .where(and(gt(documents.deletedAt, since), isNotNull(documents.deletedAt))); if (deleted.length > 0) { await store.nodes.Document.bulkDelete(deleted.map((d) => d.id)); console.log(`Removed ${deleted.length} deleted documents`); } return { lastSyncAt: now }; } ``` ### Scheduled Sync Job Run incremental sync on a schedule: ```typescript import { CronJob } from "cron"; // Store sync state (in production, persist this to a database) let syncState: SyncState = { lastSyncAt: new Date(0) }; // Run every 5 minutes const syncJob = new CronJob("*/5 * * * *", async () => { try { syncState = await incrementalSync(store, syncState); } catch (error) { console.error("Sync failed:", error); // Alert, retry, etc. } }); syncJob.start(); ``` ## Event-Driven Sync React to changes in your source data via events, webhooks, or database triggers. Best for high-volume scenarios requiring near-real-time sync. ### Message Queue Pattern ```typescript import { Queue, Worker } from "bullmq"; // Define sync job types interface SyncJob { type: "upsert" | "delete"; entityType: "Document" | "User"; entityId: string; } // Producer: Enqueue sync jobs when source data changes const syncQueue = new Queue("sync"); async function onDocumentCreated(doc: AppDocument) { await syncQueue.add("sync", { type: "upsert", entityType: "Document", entityId: doc.id, }); } async function onDocumentUpdated(doc: AppDocument) { await syncQueue.add("sync", { type: "upsert", entityType: "Document", entityId: doc.id, }); } async function onDocumentDeleted(docId: string) { await syncQueue.add("sync", { type: "delete", entityType: "Document", entityId: docId, }); } // Consumer: Process sync jobs const syncWorker = new Worker( "sync", async (job) => { const { type, entityType, entityId } = job.data; if (type === "delete") { await store.nodes[entityType].delete(entityId); return; } // Fetch current state from source const record = await fetchEntity(entityType, entityId); if (!record) { // Record was deleted between enqueue and processing await store.nodes[entityType].delete(entityId); return; } // Generate embedding if needed const embedding = await generateEmbedding(record.content); // Upsert to graph await store.nodes[entityType].upsertById(entityId, { ...record, embedding, source: { table: entityType.toLowerCase() + "s", id: entityId }, }); }, { concurrency: 10, connection: redis, } ); ``` ### Webhook Handler Process webhooks from external systems: ```typescript import { Hono } from "hono"; const app = new Hono(); app.post("/webhooks/documents", async (c) => { const event = await c.req.json<{ type: "created" | "updated" | "deleted"; data: AppDocument; }>(); switch (event.type) { case "created": case "updated": { const embedding = await generateEmbedding(event.data.content); await store.nodes.Document.upsertById(event.data.id, { title: event.data.title, content: event.data.content, embedding, source: { table: "documents", id: event.data.id }, }); break; } case "deleted": { await store.nodes.Document.delete(event.data.id); break; } } return c.json({ ok: true }); }); ``` ### Database Triggers (PostgreSQL) Use LISTEN/NOTIFY for real-time sync from PostgreSQL: ```typescript import { Client } from "pg"; // Set up listener const listener = new Client({ connectionString: process.env.DATABASE_URL }); await listener.connect(); await listener.query("LISTEN document_changes"); listener.on("notification", async (msg) => { if (msg.channel !== "document_changes") return; const payload = JSON.parse(msg.payload!); const { operation, id } = payload; if (operation === "DELETE") { await store.nodes.Document.delete(id); return; } // Fetch and sync the changed document const doc = await db.select().from(documents).where(eq(documents.id, id)).get(); if (doc) { const embedding = await generateEmbedding(doc.content); await store.nodes.Document.upsertById(id, { title: doc.title, content: doc.content, embedding, source: { table: "documents", id }, }); } }); ``` Corresponding PostgreSQL trigger: ```sql CREATE OR REPLACE FUNCTION notify_document_changes() RETURNS TRIGGER AS $$ BEGIN PERFORM pg_notify( 'document_changes', json_build_object( 'operation', TG_OP, 'id', COALESCE(NEW.id, OLD.id) )::text ); RETURN COALESCE(NEW, OLD); END; $$ LANGUAGE plpgsql; CREATE TRIGGER document_changes_trigger AFTER INSERT OR UPDATE OR DELETE ON documents FOR EACH ROW EXECUTE FUNCTION notify_document_changes(); ``` ## Syncing Relationships When syncing data that includes relationships, sync nodes first, then edges: ```typescript interface ExternalUser { id: string; name: string; email: string; managerId?: string; } async function syncUsers(users: ExternalUser[]) { // Step 1: Sync all user nodes first await store.nodes.User.bulkUpsertById( users.map((u) => ({ id: u.id, props: { name: u.name, email: u.email, source: { table: "users", id: u.id }, }, })) ); // Step 2: Sync manager relationships // First, remove all existing manages edges (clean slate approach) const existingEdges = await store.edges.manages.find(); if (existingEdges.length > 0) { await store.edges.manages.bulkDelete(existingEdges.map((e) => e.id)); } // Then create edges for users with managers const usersWithManagers = users.filter((u) => u.managerId); await store.edges.manages.bulkInsert( usersWithManagers.map((u) => ({ from: { kind: "User" as const, id: u.managerId! }, to: { kind: "User" as const, id: u.id }, })), ); } ``` ## Handling Sync Failures ### Retry with Exponential Backoff ```typescript async function syncWithRetry( fn: () => Promise, options: { maxRetries?: number; baseDelay?: number } = {} ): Promise { const { maxRetries = 3, baseDelay = 1000 } = options; let lastError: Error | undefined; for (let attempt = 0; attempt <= maxRetries; attempt++) { try { return await fn(); } catch (error) { lastError = error as Error; if (attempt < maxRetries) { const delay = baseDelay * Math.pow(2, attempt); console.warn(`Sync attempt ${attempt + 1} failed, retrying in ${delay}ms`); await new Promise((resolve) => setTimeout(resolve, delay)); } } } throw lastError; } // Usage await syncWithRetry(() => store.nodes.Document.bulkUpsertById(items)); ``` ### Dead Letter Queue Track failed syncs for manual intervention: ```typescript interface FailedSync { entityType: string; entityId: string; error: string; failedAt: Date; attempts: number; } const failedSyncs: FailedSync[] = []; async function syncWithDLQ(entityType: string, entityId: string, syncFn: () => Promise) { try { await syncWithRetry(syncFn); } catch (error) { failedSyncs.push({ entityType, entityId, error: (error as Error).message, failedAt: new Date(), attempts: 3, }); console.error(`Sync failed after retries: ${entityType}:${entityId}`); } } // Periodically retry or alert on failed syncs async function processFailedSyncs() { for (const failed of failedSyncs) { console.log(`Failed sync: ${failed.entityType}:${failed.entityId} - ${failed.error}`); // Retry, alert, or log for manual intervention } } ``` ## Best Practices ### Use Consistent IDs Map external IDs to graph node IDs consistently: ```typescript // Good: Use external ID directly when it's unique and stable await store.nodes.Document.upsertById(externalDoc.id, { ... }); // Good: Namespace if IDs might collide across sources await store.nodes.Document.upsertById(`notion:${notionPage.id}`, { ... }); await store.nodes.Document.upsertById(`gdrive:${driveFile.id}`, { ... }); ``` ### Track Sync Metadata Store sync information for debugging and auditing: ```typescript const Document = defineNode("Document", { schema: z.object({ title: z.string(), content: z.string(), embedding: embedding(1536).optional(), source: externalRef("documents"), // Sync metadata lastSyncedAt: z.string().datetime().optional(), syncVersion: z.number().optional(), }), }); await store.nodes.Document.upsertById(doc.id, { ...props, lastSyncedAt: new Date().toISOString(), syncVersion: (existingNode?.syncVersion ?? 0) + 1, }); ``` ### Validate Before Sync Validate external data before syncing to avoid corrupting your graph: ```typescript const ExternalDocumentSchema = z.object({ id: z.string().min(1), title: z.string().min(1), content: z.string(), }); async function syncDocument(rawDoc: unknown) { const result = ExternalDocumentSchema.safeParse(rawDoc); if (!result.success) { console.error("Invalid document data:", result.error); return; } await store.nodes.Document.upsertById(result.data.id, { title: result.data.title, content: result.data.content, }); } ``` ### Monitor Sync Health Track sync metrics for observability: ```typescript const syncMetrics = { successful: 0, failed: 0, lastSyncDuration: 0, lastSyncAt: null as Date | null, }; async function monitoredSync(fn: () => Promise) { const start = Date.now(); try { await fn(); syncMetrics.successful++; } catch (error) { syncMetrics.failed++; throw error; } finally { syncMetrics.lastSyncDuration = Date.now() - start; syncMetrics.lastSyncAt = new Date(); } } // Expose metrics endpoint app.get("/metrics/sync", (c) => c.json(syncMetrics)); ``` ## Next Steps - [Integration Patterns](/integration) - Database setup and deployment patterns - [Semantic Search](/semantic-search) - Add vector embeddings during sync - [Query Builder](/queries/overview) - Query your synced graph data # Fulltext Search > BM25-style fulltext search with hybrid retrieval for RAG applications TypeGraph supports fulltext search directly in your SQLite or PostgreSQL database — no external search service required. Combine it with semantic search to get **hybrid retrieval**: the gold-standard pattern for RAG applications. ## Overview Vector search is great at finding *semantically* similar content, but it misses exact matches: proper nouns, SKU numbers, code identifiers, rare technical terms. Fulltext search handles those. Running both and fusing the results with Reciprocal Rank Fusion typically beats either approach alone. **Key capabilities:** - Declare `searchable()` string fields in your Zod schema - Native BM25 ranking (SQLite FTS5) and `ts_rank_cd` (PostgreSQL tsvector) - Google-style query syntax: quoted phrases, `-excluded`, `OR` - `n.$fulltext.matches()` predicate composes with metadata filters and graph traversal - Hybrid search via `$fulltext.matches()` + `.similarTo()` in one query, fused with RRF - Tunable RRF via `.fuseWith({ k, weights })` on the query builder, or `store.search.hybrid({ fusion })` ## Use Cases ### Hybrid RAG Combine exact-match retrieval with semantic similarity: ```typescript const hits = await store.search.hybrid("Document", { limit: 10, vector: { fieldPath: "embedding", queryEmbedding: await embed(question), }, fulltext: { query: question }, }); const context = hits.map((h) => h.node.content).join("\n\n"); ``` ### Multi-tenant fulltext with metadata filters The most important composition — `$fulltext.matches()` in the same query as any other predicate: ```typescript const results = await store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext.matches("climate change", 20).and(d.tenantId.eq(tenant.id)) ) .select((ctx) => ctx.d) .execute(); ``` ### Authorised search via graph traversal Only return documents the user is allowed to read: ```typescript const results = await store .query() .from("User", "u") .whereNode("u", (u) => u.id.eq(currentUserId)) .traverse("canRead", "e") .to("Document", "d") .whereNode("d", (d) => d.$fulltext.matches(userQuery, 10)) .select((ctx) => ctx.d) .execute(); ``` ## Schema Design ### Declaring Searchable Fields Use `searchable()` to mark string fields for fulltext indexing: ```typescript import { defineNode, searchable } from "@nicia-ai/typegraph"; import { z } from "zod"; const Document = defineNode("Document", { schema: z.object({ title: searchable({ language: "english" }), body: searchable({ language: "english" }), tenantId: z.string(), published: z.boolean(), }), }); ``` ### How Indexing Works TypeGraph stores one fulltext row per node. When you create or update a node, the values of every `searchable()` field are concatenated and indexed as a single document. This single-document-per-node design lets a single query find matches that span multiple source fields — a title hit plus a body hit both contribute to the same score. - **PostgreSQL**: the `typegraph_node_fulltext` table carries a `tsvector` column populated at INSERT time, with a GIN index. - **SQLite**: the same shape is backed by an FTS5 virtual table with BM25 ranking. Sync is automatic — the fulltext index stays in sync with node data through every `create`, `update`, `upsert`, and `delete` (soft and hard). ### Searchable Options ```typescript searchable({ language: "english", // Postgres regconfig / SQLite FTS5 tokenizer }) ``` - **`language`**: Postgres uses this as the `regconfig` for stemming (`english`, `spanish`, `french`, etc.). SQLite FTS5 tokenizer is fixed at table creation time, so the language is stored but treated as metadata. ### Adding `searchable()` to an Existing Graph When you add `searchable()` to a field on a node kind that already has rows in production, those pre-existing rows are not indexed until you backfill the index: ```typescript const stats = await store.search.rebuildFulltext(); console.log( `Upserted ${stats.upserted}, cleared ${stats.cleared}, ` + `skipped ${stats.skipped} across ${stats.kinds.length} kinds`, ); if (stats.skippedIds && stats.skippedIds.length > 0) { console.warn("Nodes with corrupt props were skipped:", stats.skippedIds); } // For systemic corruption, raise the cap to collect the full list: const forensic = await store.search.rebuildFulltext(undefined, { maxSkippedIds: 1_000_000, }); ``` `store.search.rebuildFulltext()` iterates nodes with keyset pagination on `id` (stable under shared timestamps and light concurrent writes), transacts per page, and cleans up stale fulltext rows for soft-deleted nodes. Rebuild is a maintenance operation: concurrent hard-deletes between page fetches can be missed by a single pass. Run during a maintenance window for full consistency. Scope to a single kind with `store.search.rebuildFulltext("Document")` to avoid scanning unrelated data. Also useful for: - Recovering after a `DROP TABLE` / `TRUNCATE` of the fulltext table. - Re-tokenizing after changing `language` on a `searchable()` field. - Recovering from bulk inserts that bypassed the store layer. ### Checking Whether Search Is Ready `store.search.rebuildFulltext()` fixes *content*. It cannot fix storage that is missing, unattested, or provisioned at the wrong shape — and it throws `StoreNotInitializedError` when it is, because the hot-path gate refuses fulltext writes until both the deployment marker attests the shared table and the graph-local activation marker admits this graph. To find out which situation you are in without writing anything: ```typescript const health = await store.probeContributions(); const fulltext = health.entries.find( (entry) => entry.contribution === "fulltext", ); if (fulltext?.state !== "ready") { console.warn(`fulltext search is ${fulltext?.state}`, fulltext?.detail); } ``` The probe writes nothing, so it is safe to call from a health check, on a read path, or on a replica — which is the point: the alternative was to run `store.repairContributions()`, a write with repair side effects, and hope. When it reports `degraded`, escalate through the contribution health ladder — probe, then `repairContributions()`, then `rebuildContribution("fulltext")`, which is scoped to the calling graph: it deletes and refills only that graph's rows in the shared fulltext table, and drops and recreates the table itself only when no other graph has rows in it. The three rungs, what each one writes, and when to stop are in [Contribution health: probe, repair, rebuild](/troubleshooting#contribution-health-probe-repair-rebuild). ## Database Setup ### Initialization is required (boot via `createStoreWithSchema`) Fulltext storage is **durably materialized once, at application boot**, by `createStoreWithSchema`: ```typescript // Run this once at startup — outside request handlers and transactions. const [store] = await createStoreWithSchema(graph, backend); ``` `createStore(graph, backend)` is a synchronous, zero-I/O *attach*: it does not create tables, repair DDL, or record that fulltext storage is materialized. A fulltext read or write — a `searchable()` field write, `store.search.fulltext()`, `store.search.hybrid()`, `n.$fulltext.matches()`, `store.search.rebuildFulltext()`, or a transaction that touches fulltext — against a database that was never initialized throws `StoreNotInitializedError`. Use `createStore()` only to attach to a database a prior `createStoreWithSchema` boot already initialized. Graphs with no `searchable()` fields are unaffected. ### PostgreSQL No extensions required. The built-in `tsvector` type and GIN indexes work on every managed Postgres (RDS, Supabase, Neon, Cloud SQL, Aiven). The fulltext table's DDL ships in `bootstrapTables()` and the migration SQL. The first privileged `createStoreWithSchema` boot records one deployment-scoped physical marker for that shared table plus a graph-local activation marker. Later graphs reuse the physical attestation and need only the DML activation write; fulltext operations require both markers (see above): ```typescript import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; // Includes the fulltext table, tsvector column, GIN index, and pgvector const migrationSQL = generatePostgresMigrationSQL(); ``` ### SQLite No extensions required. FTS5 is compiled into the standard SQLite distribution shipped with `better-sqlite3`, `libsql`, `bun:sqlite`, and most other drivers. The FTS5 virtual table uses the `porter unicode61 remove_diacritics 2` tokenizer. ## Querying ### `n.$fulltext.matches()` — The Query Predicate `n.$fulltext.matches(query, k?, options?)` is a node-level fulltext predicate. It's exposed on every `NodeAccessor`; at runtime it throws a clear `UnsupportedPredicateError` if the node kind has no `searchable()` fields, with a suggestion for how to fix the schema. > **Visible in types, guarded at runtime.** `$fulltext` is present on every > `NodeAccessor` at the TypeScript level for simplicity — a type-level > brand would not survive modifiers like `.min(1).optional()`. The runtime > check is the single source of truth: adding a `searchable()` field is > what makes `.matches()` actually work. A query that type-checks can still > throw `UnsupportedPredicateError` the first time it runs if no field > was declared searchable. ```typescript store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext.matches("climate change")) .select((ctx) => ctx.d) .execute(); ``` It compiles to a JOIN against the fulltext index, adds an ORDER BY on relevance rank, and applies the top-k limit — all in a single SQL statement that composes with every other query-builder feature. **`k` vs `limit`**: `k` (the second positional arg) is the top-k cap applied **inside the fulltext CTE** — how many candidates to pull from the index before outer filtering and fusion. It defaults to `50`, which is fine for single-predicate use. `.limit()` on the query controls the **final result count**. When feeding into RRF (`store.search.hybrid` or `.fuseWith()`), pass a larger `k` per predicate (e.g. 200) so there are enough candidates for the fused ranking to be meaningful. Traversal happens after candidate generation. A candidate can therefore fan out into several match rows. Use query-level `.where((ctx) => ...)` to filter those completed rows, then apply an explicit `.orderBy()` and `.limit()` for the final result. The final order does not change which nodes entered the top-k candidate set, and the final limit counts match rows rather than distinct source nodes. ### Query Modes ```typescript d.$fulltext.matches("climate -warming", 10, { mode: "websearch" }) // Google-style: quoted phrases, -excluded terms, OR operator d.$fulltext.matches("climate change", 10, { mode: "phrase" }) // Exact phrase match d.$fulltext.matches("climate change", 10, { mode: "plain" }) // All terms must appear (implicit AND), no special syntax d.$fulltext.matches("climate & !warming", 10, { mode: "raw" }) // Dialect-native syntax (tsquery on Postgres, FTS5 on SQLite) ``` **When to use each:** | Mode | Best For | Example | |------|----------|---------| | `websearch` (default) | User-facing search boxes | `"climate change" -hoax OR warming` | | `phrase` | Proper nouns, exact quotes | `"New York Times"` | | `plain` | Programmatic queries | `climate change` | | `raw` | Advanced users who know the dialect syntax | `climate<->change` | ### Composing with Filters Fulltext is just another predicate — combine with `.and()`: ```typescript store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext .matches("machine learning", 20, { mode: "websearch" }) .and(d.published.eq(true)) .and(d.publishedAt.gte("2024-01-01")) .and(d.tenantId.eq(tenant)) ) .select((ctx) => ctx.d) .execute(); ``` ### Composing with Graph Traversal `$fulltext.matches()` works inside any traversal: ```typescript // Find documents matching "climate" that were written by someone I follow const results = await store .query() .from("Person", "me") .whereNode("me", (p) => p.id.eq(currentUserId)) .traverse("follows", "f") .to("Person", "author") .traverse("authored", "a", { direction: "in" }) .to("Document", "d") .whereNode("d", (d) => d.$fulltext.matches("climate", 10)) .select((ctx) => ({ title: ctx.d.title, author: ctx.author.name, })) .execute(); ``` ### Hybrid Search (Query Builder) Use `$fulltext.matches()` and `.similarTo()` in the same query and TypeGraph automatically fuses the two signals with Reciprocal Rank Fusion at the SQL layer: ```typescript const results = await store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext .matches("renewable energy", 50) .and(d.embedding.similarTo(queryVector, 50)) .and(d.tenantId.eq(tenant)) ) .select((ctx) => ctx.d) .limit(10) .execute(); ``` The compiled SQL builds two CTEs (one for the vector side, one for the fulltext side), orders each by relevance, and the outer query sorts by `1/(60 + rank_vector) + 1/(60 + rank_fulltext)`. One round-trip, fully composable with any other predicate. ### Tuning RRF Defaults (k=60, equal weights) suit most workloads. Bias toward fulltext for exact-match queries, toward vectors for conceptual queries: ```typescript store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext .matches("renewable energy", 50) .and(d.embedding.similarTo(queryVector, 50)) .and(d.tenantId.eq(tenant)) ) .fuseWith({ k: 60, weights: { vector: 1.0, fulltext: 1.5 } }) .limit(10) .execute(); ``` `.fuseWith()` throws at compile time if the query lacks either a `.similarTo()` or a `$fulltext.matches()`. Validation rejects non-finite or negative `k`/weights. The same shape is accepted by `store.search.hybrid({ fusion })` and validated by the same function on both paths. ### Hybrid Search (Store API) For tunable RRF parameters, use `store.search.hybrid()`. On the built-in backends this runs as a **single SQL statement** — both sources, RRF fusion, and node hydration composed together — so a hybrid query costs one round trip instead of three. The saving scales with per-statement cost: decisive on serverless HTTP drivers, Cloudflare D1 / Durable Objects, and remote databases; on a local low-latency connection the two paths are within a few milliseconds of each other. (Kind expansions via `includeSubClasses`, and custom backends without the composed statement, transparently use a multi-statement path with identical results.) ```typescript const results = await store.search.hybrid("Document", { limit: 10, vector: { fieldPath: "embedding", queryEmbedding: await embed(question), metric: "cosine", k: 50, // Candidates to retrieve from vector }, fulltext: { query: question, k: 50, // Candidates to retrieve from fulltext includeSnippets: true, }, fusion: { method: "rrf", k: 60, // RRF constant (classic default) weights: { vector: 1.0, fulltext: 1.5, // Weight fulltext higher for exact-match workloads }, }, }); ``` Each hit carries sub-scores from both halves so you can debug ranking: ```typescript for (const hit of results) { console.log(hit.node.title, hit.score); console.log(" vector rank:", hit.vector?.rank); console.log(" fulltext rank:", hit.fulltext?.rank); console.log(" snippet:", hit.fulltext?.snippet); } ``` ### Fulltext-Only Store API For quick fulltext lookups that don't need the query builder: ```typescript const hits = await store.search.fulltext("Document", { query: "quarterly earnings", limit: 10, mode: "websearch", includeSnippets: true, }); for (const hit of hits) { console.log(hit.node.title, hit.score, hit.snippet); } ``` #### Options reference | Option | Type | Default | Description | |--------|------|---------|-------------| | `query` | `string` | — *(required)* | User-supplied query string. Parsed according to `mode`. | | `limit` | `number` | — *(required)* | Max rows to return. Must be a positive integer. | | `mode` | `"websearch" \| "phrase" \| "plain" \| "raw"` | `"websearch"` | Parser for `query`. See [Query Modes](#query-modes). | | `language` | `string` | the kind's declared language | Override the stemming/tokenization language for this query. By default the query is parsed with the language the kind's `searchable()` fields declare — a plan-time constant, which is what lets PostgreSQL serve the match from the `tsv` GIN index (a per-row language reference makes the tsquery non-constant and forces a scan). Postgres only — SQLite FTS5's tokenizer is fixed at table-create time and a per-query override throws. | | `minScore` | `number` | *(none)* | Drop hits whose backend-native score is below this threshold. Score units depend on the strategy. | | `includeSnippets` | `boolean` | `false` | Return a highlighted `…` snippet per hit. Noticeably slower than plain search — request only for final-page results. | | `where` | `(accessor) => Predicate` | *(none)* | Property predicate compiled into the search statement's candidate set — the engine ranks only matching rows, so a filter never shrinks results below `limit` when enough matches exist (libSQL DiskANN: bounded by its 4× over-fetch). Same accessor and semantics as `store.nodes..find({ where })`. | | `offset` | `number` | `0` | Rank-relative pagination: skip the first `offset` ranked hits. | | `includeSubClasses` | `boolean` | `false` | Also search `subClassOf` descendant kinds and merge their scores into one ranking. | The same three options are available on `store.search.vector` and `store.search.hybrid` (where `where` and `includeSubClasses` apply to both halves). Search always follows current-read semantics: tombstoned nodes and nodes outside their validity window never rank. Returned hits are `FulltextSearchHit>` with `node`, `score` (higher = more relevant), `rank` (1-based), and `snippet` (when requested). ## Reciprocal Rank Fusion RRF is the de facto standard for combining ranked lists from multiple retrievers. The formula: ```text score(doc) = Σ weight_source / (k + rank_source) ``` Where `k` is the RRF constant (classic default: 60), `rank_source` is the document's 1-based ordinal rank in each source, and `weight_source` lets you bias toward one retriever. **Why it works:** RRF is rank-based, not score-based. It doesn't care that vector distances are in `[0, 2]` while BM25 scores are unbounded — it only cares about ordinal position. That makes it robust to heterogeneous score distributions across retrievers. **Tuning tips:** - Over-fetch from each side (default: 4× the requested limit). More candidates per source = better recall. - Bump `weights.fulltext` higher when exact matches matter (names, IDs, proper nouns). Bump `weights.vector` for conceptual queries. - Leave `k = 60` alone unless benchmarks show otherwise. ## Best Practices ### All Searchable Fields Share One Index TypeGraph indexes all `searchable()` fields on a node as one document (see [How Indexing Works](#how-indexing-works)). There's a single `n.$fulltext` accessor per node — `searchable()` declarations on individual fields are what bring it into existence and what determine which text gets indexed. ### Use `includeSnippets` Sparingly Highlighting (`ts_headline` on Postgres, `snippet()` on SQLite) is noticeably slower than plain search. Request it only for final-page results, not for large over-fetch pools. ### Pair with a Reranker for Top Quality RRF is a strong baseline, but production RAG systems typically add a cross-encoder reranker (Cohere Rerank, `bge-reranker`, etc.) as a final stage. TypeGraph gives you the candidate set — the reranker picks the winning order: ```typescript const candidates = await store.search.hybrid("Document", { limit: 50, // Over-fetch for reranker vector: { fieldPath: "embedding", queryEmbedding }, fulltext: { query }, }); const reranked = await cohere.rerank({ query, documents: candidates.map((c) => c.node.content), top_n: 10, }); ``` ### Filter Before You Fuse Applying predicates via `.and()` shrinks the candidate pool before the fusion ORDER BY, which improves both latency and ranking quality — there are fewer irrelevant candidates competing for top positions: ```typescript // Fast: tenant filter applied inside each CTE .whereNode("d", (d) => d.$fulltext.matches(query, 50) .and(d.embedding.similarTo(queryVec, 50)) .and(d.tenantId.eq(tenant)) ) // Slow: tenant filter applied AFTER fusion .whereNode("d", (d) => d.$fulltext.matches(query, 5000) .and(d.embedding.similarTo(queryVec, 5000)) ) // ...then filter results in JS ``` ## Limitations ### One Fulltext Predicate Per Query A single query can contain at most one `$fulltext.matches()` predicate. This mirrors the constraint on `.similarTo()` and keeps the RRF fusion model well-defined. If you need to search multiple terms, combine them into one query string using websearch mode: ```typescript // Good d.$fulltext.matches("climate change OR global warming", 20) // Rejected (at query-build time, not by the type checker) d.$fulltext.matches("climate", 10).and(d.$fulltext.matches("warming", 10)) ``` This invariant is enforced when the query is compiled (`UnsupportedPredicateError`), not by TypeScript — so a surprising second `.matches()` call surfaces as a runtime error the first time the query runs. ### No `.matches()` Under OR or NOT Fulltext predicates must appear at top level or inside AND groups. They rewrite query structure (adding a CTE and ORDER BY) in a way that isn't compatible with disjunction or negation semantics. ### Tokenizer Is Fixed on SQLite FTS5 tokenizer options are set at CREATE VIRTUAL TABLE time. TypeGraph ships with `porter unicode61 remove_diacritics 2` — a solid default for English and accented Latin-script languages. For CJK or other tokenizers, create the fulltext table manually with your preferred options. ### No Per-Field Weighting All searchable fields on a node contribute equally to the combined document. Postgres `setweight()`-style per-field bias is a planned extension; today, structure your fields to put the most important text first or split high-weight content into a dedicated kind. This limitation applies even when you [swap in a custom `FulltextStrategy`](#custom-fulltext-strategies) — TypeGraph concatenates searchable fields into one `content` string before handing it to the strategy. ## Custom Fulltext Strategies `createPostgresBackend(db, { fulltext })` and `createSqliteBackend(db, { fulltext })` accept a `FulltextStrategy` that owns the **entire** fulltext pipeline — DDL, INSERT/UPSERT (single + batch), DELETE (single + batch), MATCH condition, rank expression, and snippet generation. The same strategy flows through `store.search.fulltext()`, `store.search.hybrid()`, `$fulltext.matches()` in the query builder, `store.search.rebuildFulltext()`, and `bootstrapTables()` DDL. Use this when the built-in `tsvector` isn't the right fit — for example, BM25 inside Postgres (ParadeDB / `pg_search`), trigram similarity (`pg_trgm`), or fulltext optimized for CJK languages (`pgroonga`). Most SQLite users should leave the default `fts5Strategy` in place. ### Constraints on alternate strategies Before implementing a strategy, know what the abstraction does **not** let you change today: - **Side table is mandatory.** Every strategy writes one row per `(graph_id, node_kind, node_id)` to a dedicated fulltext table. A strategy cannot skip the side table and index a column on the main nodes table directly (e.g. a GIN trigram index on `typegraph_nodes.props`). Strategies *can* choose the column layout, index type, and any computed projection inside that side table. - **Content is pre-concatenated.** TypeGraph joins every `searchable()` field value with `\n` before the strategy sees it — `UpsertFulltextParams.content` is a single string. Per-field indexing (`setweight`, per-column BM25 boosts, pgroonga per-column weights) is not plumbed through today; a richer per-field payload is planned but not yet part of the public strategy contract. - **One language per row.** When a node has multiple `searchable()` fields with different `language` values, the first field's language wins and is recorded on the row. TypeGraph emits a one-time warning per conflicting schema; true multilingual indexing needs a dedicated node kind per language. ### Strategy skeleton The `FulltextStrategy` contract and every type referenced by it are exported from the Drizzle-free backend-authoring entrypoint. Fields below are the minimum surface; see `src/query/dialect/fulltext-strategy.ts` in the TypeGraph source for `tsvectorStrategy` and `fts5Strategy` as full references. ```typescript import { sql, type FulltextStrategy, type SqlFragment, } from "@nicia-ai/typegraph/backend"; /** * Example: a trigram-based strategy on top of pg_trgm. Illustrative — * not production code. pg_trgm supports plain-term matching only, so * `supportedModes` advertises `"plain"` and rejects everything else at * compile time. */ export const pgTrgmStrategy: FulltextStrategy = { name: "pg_trgm", supportedModes: ["plain"], supportsSnippets: false, // no native highlight; emit NULL snippet supportsPrefix: false, // trigram similarity, not prefix supportsLanguageOverride: false, languages: ["simple"], matchCondition(table, query) { return sql`${sql.identifier(table)}."content" % ${query}`; }, rankExpression(table, query) { return sql`similarity(${sql.identifier(table)}."content", ${query})`; }, snippetExpression() { // `supportsSnippets: false` — callers get NULL and skip the field. return sql`NULL`; }, // Declares the table(s) this strategy owns as authoritative // TableContributions. pg_trgm brings its own table (not the typed // Drizzle `tables.fulltext`), so it is emitted verbatim from // `createDdl` and is invisible to drizzle-kit unless you export your // own table object. `runtimeEnsure: true` because no // drizzle-kit-managed setup can create it. ownedTables(primaryTableName) { return [ { logicalName: "fulltext", owner: "pg_trgm", tableName: primaryTableName, createDdl: [ `CREATE EXTENSION IF NOT EXISTS pg_trgm;`, `CREATE TABLE IF NOT EXISTS "${primaryTableName}" ( "graph_id" TEXT NOT NULL, "node_kind" TEXT NOT NULL, "node_id" TEXT NOT NULL, "content" TEXT NOT NULL, "language" TEXT NOT NULL, "updated_at" TIMESTAMPTZ NOT NULL, PRIMARY KEY ("graph_id", "node_kind", "node_id") );`, `CREATE INDEX IF NOT EXISTS "${primaryTableName}_trgm_idx" ON "${primaryTableName}" USING GIN ("content" gin_trgm_ops);`, ], runtimeEnsure: true, }, ]; }, buildUpsert(table, params, timestamp) { return [ sql` INSERT INTO ${sql.identifier(table)} ("graph_id", "node_kind", "node_id", "content", "language", "updated_at") VALUES (${params.graphId}, ${params.nodeKind}, ${params.nodeId}, ${params.content}, ${params.language}, ${timestamp}) ON CONFLICT ("graph_id", "node_kind", "node_id") DO UPDATE SET "content" = EXCLUDED."content", "language" = EXCLUDED."language", "updated_at" = EXCLUDED."updated_at" `, ]; }, buildBatchUpsert(table, params, timestamp) { if (params.rows.length === 0) return []; // Dedup last-write-wins by nodeId, then emit a single multi-VALUES INSERT. // Postgres ON CONFLICT rejects repeated conflict keys inside one statement. // (The shipped helpers in fulltext-strategy.ts show this pattern.) return [/* … */]; }, buildDelete(table, params) { return [ sql` DELETE FROM ${sql.identifier(table)} WHERE "graph_id" = ${params.graphId} AND "node_kind" = ${params.nodeKind} AND "node_id" = ${params.nodeId} `, ]; }, buildBatchDelete(table, params) { if (params.nodeIds.length === 0) return []; return [/* DELETE … WHERE node_id IN (…) */]; }, }; ``` Wire it in at backend construction: ```typescript import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; const backend = createPostgresBackend(db, { fulltext: pgTrgmStrategy }); ``` Capabilities (`phraseQueries`, `prefixQueries`, `highlighting`, `languages`) are derived automatically from the strategy, so `store.search.fulltext({ mode: "websearch" })` now throws `ConfigurationError` before any SQL is generated — the strategy's `supportedModes` is the source of truth. ## Troubleshooting ### `StoreNotInitializedError: fulltext storage … is not initialized` The database was never booted through `createStoreWithSchema`, so the deployment-scoped physical marker or this graph's activation marker is missing (or it is `stale` — the strategy/DDL changed since it was recorded, or `failed` — the last boot-time attempt errored). Bare `createStore()` deliberately does **not** self-heal this on the hot path. Fix: call `createStoreWithSchema(graph, backend)` once at application startup — outside request handlers and adopted transactions — before any fulltext operation. If you previously relied on fulltext tables being created lazily on first write via `createStore()`, that path was removed; move the initialization to an explicit boot step. A `stale` reason means the recorded shape no longer matches the active strategy/DDL: migrate or drop the fulltext table and re-run the boot, or restore the original strategy. `ContributionUnavailableError` with `state: "physical-storage-missing"` is different: the physical fulltext table disappeared after initialization. Gated fulltext searches and searchable node writes raise this typed error; compiled query-builder predicates can still surface the engine's missing-relation error. Run `store.rebuildContribution("fulltext")`; rerunning ordinary initialization cannot reconstruct the missing indexed content. Use `probeContributions()` at startup when an application must detect this out-of-band catalog damage before the first dependent read or write. ### `Cannot call .$fulltext.matches() on alias "x"` `$fulltext` is exposed on every node accessor at the type level, but calling `.matches()` requires the node kind to have at least one `searchable()` field — otherwise there's no indexed content to search. The runtime guard throws a clear error pointing at the alias: ```text Cannot call .$fulltext.matches() on alias "d" — its node kind has no fields declared with searchable(). Add at least one: `title: searchable({ language: "english" })`. ``` Fix by adding a searchable field to the schema: ```typescript // Before: title: z.string(), // After: title: searchable({ language: "english" }), ``` Refinements like `.min(1)` and `.trim()` are preserved — you can write `searchable({ language: "english" }).min(1)` and the field is still indexed. ### Empty fulltext results after bulk insert TypeGraph syncs the fulltext index inline with each node write. If you bulk-inserted via raw SQL that bypassed the store layer, the fulltext table won't have entries. Re-run the inserts through `store.nodes.X.create()` / `.bulkCreate()`, run `store.search.rebuildFulltext()` to populate the index from existing rows, or issue `backend.upsertFulltext()` / `backend.upsertFulltextBatch()` calls directly. ### After adding `searchable()` to existing data See [Adding `searchable()` to an Existing Graph](#adding-searchable-to-an-existing-graph) above for the rebuild recipe and the caveats that apply to concurrent workloads. ### `"Fulltext match predicates cannot be nested under OR or NOT"` See [No `.matches()` Under OR or NOT](#no-matches-under-or-or-not) above. Move the `$fulltext.matches()` to the top level or inside an `.and()`. ### Hybrid results miss obvious matches Increase the per-source `k` (over-fetch). The default is 4× the final `limit`, which is tuned for small result pages. Large corpora benefit from `k: 200` or higher on each side. ### Postgres: `text search configuration "xyz" does not exist` The `language` you passed to `searchable({ language })` must be an installed `regconfig` on your Postgres server. Every stock install ships `simple`, `english`, `french`, `german`, `italian`, `portuguese`, `russian`, `spanish`, and `swedish`; anything else requires an extension (`zhparser` for Chinese, `pg_trgm` for trigram-based matching, or a custom dictionary). TypeGraph emits a `console.warn` at query time when you pass a language outside the backend-advertised list, but a typo or missing extension only fails when Postgres tries to build the `tsvector`. To diagnose: ```sql SELECT cfgname FROM pg_ts_config; ``` Pick a name from that list, or install the extension that provides the one you want. If you're running a managed Postgres (RDS, Supabase, Neon, Cloud SQL, Aiven), check the provider's docs for which language extensions are enabled — some require a restart or explicit enabling. ### Swapping to a custom fulltext strategy See [Custom Fulltext Strategies](#custom-fulltext-strategies) for the full interface, constraints, and a skeleton implementation. ## API Reference - **Schema**: [`searchable()`](/queries/predicates#searchable) - **Predicate**: [`n.$fulltext.matches()`](/queries/predicates#searchable) - **Tunable fusion**: `QueryBuilder.fuseWith({ k, weights })` - **Rebuild**: `store.search.rebuildFulltext(nodeKind?, { pageSize? })` - **Store API**: `store.search.fulltext()` and `store.search.hybrid()` — see the [Schemas & Stores reference](/schemas-stores). See also: - [Semantic Search](/semantic-search) — vector embeddings and `.similarTo()` - [Predicates reference](/queries/predicates) — complete predicate catalog - [Knowledge Graph for RAG](/examples/knowledge-graph-rag) — end-to-end example combining fulltext, vector, and graph traversal # Graph Algorithms > Traversal, connectivity, label propagation, and PageRank on store.algorithms Graph queries like "are Alice and Bob connected?" or "who is within two hops of this node?" are common enough that writing them as recursive CTEs by hand gets repetitive. TypeGraph exposes the high-utility algorithms as a small facade on the store: ```typescript store.algorithms.shortestPath(alice, bob, { edges: ["knows"] }); store.algorithms.weightedShortestPath(alice, bob, { edges: ["knows"], weightProperty: "strength", }); store.algorithms.reachable(alice, { edges: ["knows"] }); store.algorithms.canReach(alice, bob, { edges: ["knows"] }); store.algorithms.neighbors(alice, { edges: ["knows"], depth: 2 }); store.algorithms.degree(alice, { edges: ["knows"] }); store.algorithms.weaklyConnectedComponents({ edges: ["knows"], nodeKinds: ["Person"], }); store.algorithms.labelPropagation({ edges: ["knows"], nodeKinds: ["Person"], }); store.algorithms.pageRank({ edges: ["knows"], nodeKinds: ["Person"] }); store.algorithms.personalizedPageRank({ edges: ["knows"], seeds: [{ id: alice.id, kind: "Person" }], }); ``` Traversal calls use a set-based breadth-first frontier. Each level expands the current `(id, kind)` set with a bind-limit-aware SQL query. Transactional backends keep the visited set in a temporary working table; backends without a pinned transaction use a chunked inline working relation. Each node is admitted only at its minimum depth. Reachability rounds deduplicate edge targets before checking target-node visibility and do not compute unused predecessors. `shortestPath` and `canReach` search from both endpoints, retain predecessors for path reconstruction, and stop when the frontiers meet. `degree` remains a single `COUNT` query. Exact algorithms behave identically on SQLite and PostgreSQL; PageRank follows the numerical-tolerance contract below. The inline fallback is intentionally limited to bounded traversals. Whole-graph iterative algorithms such as WCC, label propagation, and PageRank require a pinned transactional connection and fail through their typed capability gate when one is unavailable; they never silently switch to a second full algorithm implementation. ## When to Reach for Algorithms | You want to... | Use | |----------------|-----| | Find the fewest-hop route between two nodes | `shortestPath` | | Find the cheapest route by a numeric edge property | `weightedShortestPath` | | List every node reachable from a source | `reachable` | | Check whether a node is reachable at all | `canReach` | | Get the k-hop neighborhood of a node | `neighbors` | | Count incident edges (in, out, or both) | `degree` | | Partition all visible nodes by undirected connectivity | `weaklyConnectedComponents` | | Find deterministic communities by neighbor-label voting | `labelPropagation` | | Rank nodes by global structural importance | `pageRank` | | Rank nodes relative to weighted seed nodes | `personalizedPageRank` | | Filter, sort, or project over traversal results | `.query().traverse()` / `.recursive()` | | Hydrate an entity plus all its relationships | `store.subgraph()` | Traversal algorithms return lightweight `{ id, kind, depth }` records rather than fully hydrated nodes. WCC returns one `{ id, kind, componentId, componentKind, size }` membership per visible node. Label propagation returns `{ id, kind, labelId, labelKind }` memberships. PageRank returns `{ id, kind, score }` records ordered by descending score. Use `store.nodes..getByIds(...)` when you need the full node data. ## Shared Options Every traversal algorithm takes the same base options: | Option | Type | Default | Description | |--------|------|---------|-------------| | `edges` | `readonly EdgeKinds[]` | *(required)* | Edge kinds to follow | | `maxHops` | `number` | `10` | Maximum traversal depth (1 – 1000) | | `direction` | `"out" \| "in" \| "both"` | `"out"` | Edge direction | | `cyclePolicy` | `"prevent" \| "allow"` | `"prevent"` | Compatibility option; both values use set-based node de-duplication | | `temporalMode` | `TemporalMode` | `graph.defaults.temporalMode` | Filter applied to nodes and edges along the traversal — see [Temporal Behavior](#temporal-behavior) | | `asOf` | `string` (ISO-8601) | *(none)* | Snapshot timestamp, required when `temporalMode: "asOf"` | | `workingMemory` | `string` | *(inherits server `work_mem`)* | Opt-in, transaction-scoped `work_mem` override for iterative rounds on PostgreSQL (`SET LOCAL` semantics); validated as `kB\|MB\|GB` within 64kB–2147483647kB, ignored by SQLite | `direction: "both"` treats edges as undirected. `cyclePolicy` remains accepted for compatibility with recursive query-builder traversals, but it does not change these algorithms: their node-set results and shortest paths never need to revisit a node. ## shortestPath Finds the fewest-hop path from `from` to `to`. Returns `undefined` when no path exists within `maxHops`. ```typescript const path = await store.algorithms.shortestPath(alice, bob, { edges: ["knows"], maxHops: 6, }); if (path) { console.log(`${path.depth} hops:`, path.nodes.map((n) => n.id)); } ``` The result contains the ordered sequence of nodes (endpoints inclusive) and the hop count: ```typescript type ShortestPathResult = Readonly<{ nodes: readonly Readonly<{ id: string; kind: string }>[]; depth: number; }>; ``` Source equal to target returns a zero-length path containing just that node. Endpoints that don't pass the resolved temporal filter return `undefined` — see [Temporal Behavior](#temporal-behavior). When several minimum-hop paths exist, predecessor and meeting-node ties use the smallest `(id, kind)` identity under portable binary ordering, so the selected path is identical across backends and execution strategies. ## weightedShortestPath Finds the minimum-total-weight path from `from` to `to`, weighting each traversed edge by a numeric property stored on it. This is the shape of LDBC's "trusted connection paths" query (Interactive IC14): the cheapest route where each hop's cost reflects, say, interaction strength — which is usually not the fewest-hop route. ```typescript const path = await store.algorithms.weightedShortestPath(alice, bob, { edges: ["knows"], weightProperty: "interactionCost", direction: "both", }); if (path) { console.log(`weight ${path.totalWeight} over ${path.depth} hops`); } ``` ```typescript type WeightedShortestPathResult = Readonly<{ nodes: readonly Readonly<{ id: string; kind: string }>[]; depth: number; // hop count, nodes.length - 1 totalWeight: number; // sum of traversed edge weights }>; ``` Options differ from the unweighted traversals: | Option | Type | Default | Description | |--------|------|---------|-------------| | `edges` | `readonly EdgeKinds[]` | *(required)* | Edge kinds to follow | | `weightProperty` | `string` | *(required)* | Top-level edge property holding each edge's non-negative numeric weight | | `defaultWeight` | `number` | *(none)* | Substituted for edges missing the property; must be non-negative and within the audit's upper bound (~9.7e289) | | `direction` | `"out" \| "in" \| "both"` | `"out"` | Edge direction | | `maxIterations` | `number` | `1000` | Relaxation-round backstop; exceeding it throws `GraphAlgorithmConvergenceError` | | `temporalMode` / `asOf` / `workingMemory` | | | Same as the shared options above | There is no `maxHops`: cost-ordered search does not settle nodes in hop order, so a hop bound is not a natural stopping rule here. The algorithm relaxes frontier nodes round by round (with parallel edges collapsing to their cheapest member), prunes any candidate costing strictly more than a known path to the target — equal-cost candidates stay in play so every equal-cost route to the target is considered — and stops when no distance improves. **Weights are validated up front.** Before any traversal rounds run, every visible edge of the selected kinds is audited; the call throws a typed `InvalidEdgeWeightError` naming the offending edge when a weight is: - **negative** — the pruning that makes the search terminate early assumes non-negative weights, so they are rejected rather than silently mis-answered; - **non-numeric** — a JSON string like `"5"` does not count; the property must be stored as a JSON number (a JSON `null` counts as missing, not non-numeric); - **out of range** — a magnitude above ~9.7e289 (bounded so path sums can never overflow the double range, on either backend) or a nonzero magnitude below the smallest IEEE 754 double. One engine caveat: SQLite's JSON parser rounds sub-denormal text like `1e-400` to `0` before SQL can observe it, so only PostgreSQL can reject that case; - **missing** without a configured `defaultWeight`. The audit covers the selected edge kinds globally (not just edges the traversal happens to reach), so a data problem fails deterministically no matter which endpoints you query. Weight arithmetic uses IEEE 754 double precision on both backends: total weights are always backend-identical, and — unless a single call's `edges` list exceeds the backend's bind-parameter budget (hundreds of kinds, where equal-weight predecessor ties can resolve differently) — so is the returned node sequence. Path extraction honors the backend's [`recursiveTraversal`](/backend-setup#recursive-traversal-capability) declaration. A supported backend reconstructs the result with one recursive statement. A backend declaring `{ supported: false, reason }` still runs the full weighted search when it supports the temporary working-table operations the algorithm requires; TypeGraph reconstructs the same result by walking predecessors instead, using `path.depth + 1` extraction statements. This affects round trips, not the selected path or its weight. The unweighted traversal algorithms do not emit a recursive CTE and are unaffected by this capability. ## reachable Returns every node reachable from `from` within `maxHops`, annotated with the shortest depth at which it was discovered. ```typescript const reachable = await store.algorithms.reachable(alice, { edges: ["knows"], maxHops: 3, }); // [{ id, kind, depth: 0 }, { id, kind, depth: 1 }, ...] ``` Pass `excludeSource: true` to drop the zero-depth source entry. Results are sorted by ascending depth. ## canReach Fast boolean check that uses the same balanced bidirectional BFS as `shortestPath` and stops as soon as the two visited sets meet. ```typescript const connected = await store.algorithms.canReach(alice, bob, { edges: ["knows"], maxHops: 6, }); ``` Use this when you only care whether a path exists. It shares the same search cost as `shortestPath` but does not return the reconstructed path. ## neighbors The k-hop neighborhood of a node, with the source always excluded. `depth` defaults to `1`, so `neighbors(alice)` returns Alice's immediate connections. ```typescript const immediate = await store.algorithms.neighbors(alice, { edges: ["knows"], }); const twoHop = await store.algorithms.neighbors(alice, { edges: ["knows"], depth: 2, }); ``` Semantically equivalent to `reachable({ maxHops: depth, excludeSource: true })`, but the name reads more naturally for neighborhood queries. ## degree Counts edges incident to a node under the resolved temporal filter. Self-loops contribute once when `direction` is `"both"`. ```typescript const friends = await store.algorithms.degree(alice, { edges: ["knows"], direction: "out", }); const connections = await store.algorithms.degree(alice, { edges: ["knows"], // direction: "both" is the default }); const everything = await store.algorithms.degree(alice); // No `edges` option counts all edge kinds in the graph ``` | Option | Type | Default | Description | |--------|------|---------|-------------| | `edges` | `readonly EdgeKinds[]` | all kinds | Edge kinds to count (empty array returns 0) | | `direction` | `"out" \| "in" \| "both"` | `"both"` | Count outgoing, incoming, or either | | `temporalMode` | `TemporalMode` | `graph.defaults.temporalMode` | Filter applied to the counted edges | | `asOf` | `string` (ISO-8601) | *(none)* | Snapshot timestamp, required when `temporalMode: "asOf"` | `degree` runs a single `COUNT` query, not a recursive CTE, so it's efficient even for hub nodes with thousands of edges. ## PageRank and Personalized PageRank `pageRank` scores every visible node by the stationary probability of a random walk over the selected edges. `personalizedPageRank` uses the same power iteration but teleports to weighted seed nodes instead of uniformly across the graph. Both methods operate on the visible induced graph; `nodeKinds` can narrow that graph without allowing transitions through excluded nodes. ```typescript const globalScores = await store.algorithms.pageRank({ edges: ["knows", "cites"], nodeKinds: ["Person", "Paper"], direction: "out", topK: 20, }); const relatedToAlice = await store.algorithms.personalizedPageRank({ edges: ["knows"], direction: "both", seeds: [ { id: alice.id, kind: "Person", weight: 3 }, { id: bob.id, kind: "Person" }, // weight defaults to 1 ], }); // [{ id, kind, score }, ...] — highest score first ``` | Option | Type | Default | Description | |--------|------|---------|-------------| | `edges` | `readonly EdgeKinds[]` | *(required)* | Edge kinds defining transitions | | `nodeKinds` | `readonly NodeKinds[]` | all visible kinds | Induced node subgraph to rank | | `direction` | `"out" \| "in" \| "both"` | `"out"` | Follow stored, reversed, or undirected transitions | | `dampingFactor` | `number` in `[0, 1)` | `0.85` | Probability of following an edge rather than teleporting | | `tolerance` | positive finite `number` | `1e-8` | Maximum per-node score change accepted as convergence | | `maxIterations` | positive integer | `1000` | Power-iteration backstop; exceeding it throws | | `topK` | positive safe integer | all scores | Return the first K scores after deterministic ordering | | `workingMemory` | PostgreSQL memory string | inherited | Same transaction-scoped override described above | Personalized seeds are qualified by the full `(kind, id)` identity so same-ID nodes of different kinds remain distinct. Seed weights must be finite and positive; duplicates are combined and the vector is normalized. A seed outside the temporal/node-kind scope throws `ConfigurationError` instead of silently losing teleport mass. Dangling-node mass is redistributed through the same teleport vector, keeping the scores normalized to approximately one. Parallel physical edges retain their multiplicity. With `direction: "both"`, a physical self-loop contributes once rather than once per expansion direction. PageRank uses double-precision arithmetic on both backends. Repeated runs on SQLite are deterministic; on PostgreSQL a plan change between runs can reorder floating-point summation and shift scores by a few last bits, which can reorder near-tied rows. SQLite and PostgreSQL scores are expected to agree within the requested numerical tolerance rather than bit-for-bit. Exact score ties use portable binary `(id, kind)` ordering. As with WCC, an exhausted iteration budget throws `GraphAlgorithmConvergenceError`—partial scores are never returned. `topK` bounds only result extraction: TypeGraph applies the limit in SQL after ordering by descending score and the deterministic tie-breaker, so rows beyond K do not reach the driver. PageRank still initializes and iterates over the entire visible induced graph; `topK` does not make the graph computation itself partial. Convergence needs roughly `ln(1/tolerance) / ln(1/dampingFactor)` rounds in the worst case. With the default `tolerance` and `maxIterations`, damping factors above roughly `0.985` cannot converge in time — every round still runs against the full working table before the budget-exhaustion error is thrown — so raise `maxIterations` (or loosen `tolerance`) alongside a high damping factor. The tolerance is also absolute: typical scores are on the order of `1/N`, so reliably ranking the low-score tail of a large graph calls for a proportionally smaller tolerance. ## labelPropagation Runs deterministic synchronous Community Detection using Label Propagation (CDLP) over the undirected projection of selected edge kinds. Each visible node starts with its own `(id, kind)` identity as its label. A round adopts the most frequent label among the node's visible neighbors from the previous round; equal vote counts resolve to the minimum label under portable binary ordering. ```typescript const memberships = await store.algorithms.labelPropagation({ edges: ["knows", "worksAt"], nodeKinds: ["Person", "Company"], // optional; all kinds by default maxIterations: 1000, // default onMaxIterations: "throw", // default; "return" accepts the fixed-round labeling }); // [{ id, kind, labelId, labelKind }, ...] ``` The graph is a neighbor set for voting: parallel edges and repeated selected edge kinds do not multiply a vote, and self-loops do not make a node its own neighbor. An isolated in-scope node therefore retains its initial label. Labels and vote counts are exact integers, and edge-kind chunks are accumulated before a label is staged, so results do not depend on bind limits or chunk order. Changed-node frontiers restrict each later round to nodes whose neighbor labels may have changed. Synchronous voting has no self-vote, so structures whose neighborhoods mirror each other never converge — they alternate between two labelings forever. This covers every tree-shaped component (an isolated edge pair, a path, a star, an org-chart hierarchy), every even-length cycle, and complete bipartite blocks; empirically most sparse random graphs contain at least one such component. One oscillating component anywhere in the selection prevents global convergence, and raising `maxIterations` cannot help. Dense neighborhoods built on odd cycles — triangles, cliques, and communities of them — converge. `onMaxIterations` selects the completion contract: - `"throw"` (default) returns only a converged labeling. A detected period-two oscillation throws `GraphAlgorithmConvergenceError` immediately rather than burning the remaining budget, and exhausting `maxIterations` throws the same typed error. Partial labelings are never returned. - `"return"` yields the labeling after exactly `maxIterations` synchronous rounds (or at convergence, whichever comes first) — the fixed-round contract of the LDBC Graphalytics CDLP benchmark. Synchronous rounds are deterministic and chunk-independent, so this labeling is exact and identical on SQLite and PostgreSQL. Use it for tree-shaped or mixed data where a converged labeling need not exist; a detected oscillation fast-forwards to the parity-exact final labeling instead of running every remaining round. Like WCC and PageRank, label propagation requires `backend.capabilities.graphAnalytics?.supported === true`, runs in one repeatable snapshot, honors temporal and recorded-time views, and accepts the transaction-scoped `workingMemory` option. ## weaklyConnectedComponents Computes an exact partition over the undirected projection of the selected edge kinds. By default every visible node is returned. `nodeKinds` restricts the operation to the induced subgraph over those kinds; an in-scope node with no selected incident edge forms a singleton component. ```typescript const memberships = await store.algorithms.weaklyConnectedComponents({ edges: ["knows", "worksAt"], nodeKinds: ["Person", "Company"], // optional; all kinds by default minComponentSize: 2, // optional; singleton components are omitted maxIterations: 1000, // default }); // [{ id, kind, componentId, componentKind, size }, ...] ``` The component identity is the smallest `(id, kind)` member under portable binary ordering. Results and representatives are therefore deterministic on SQLite and PostgreSQL even when the PostgreSQL database uses a linguistic default collation. Edges are always treated as undirected—WCC has no `direction` option. `minComponentSize` is an inclusive positive safe-integer filter: a value of `2` returns every member of components containing two or more nodes, while a value of `1` preserves the default result. TypeGraph computes component sizes and filters memberships in extraction SQL, before rows reach the driver. The full visible induced graph is still processed to convergence, so the option bounds result materialization rather than WCC computation. WCC is iterative and runs multiple SQL rounds in one repeatable snapshot. It requires `backend.capabilities.graphAnalytics?.supported === true`; built-in SQLite and PostgreSQL connections that permit temporary tables advertise support. Cloudflare D1, Durable Objects SQLite, `neon-http`, and other restricted backends throw `UnsupportedBackendCapabilityError` before temporary state is created. PostgreSQL is checked again at execution time, because a read replica or a role without `TEMP` can reject the working-table transaction even when the backend's static capability is true: a replica refuses the read-write transaction itself, and a role without `TEMP` refuses the `CREATE TEMP TABLE` inside it. Both are reported as `UnsupportedBackendCapabilityError`; its `cause` preserves the original database error. If propagation has not converged after `maxIterations`, TypeGraph throws `GraphAlgorithmConvergenceError` rather than returning a partial partition. Each round expands only the indexed frontier of nodes whose label changed in the previous round. Candidate labels remain staged until every edge-kind chunk has run, preserving synchronous and bind-limit-independent iteration semantics; only changed rows are written back. Late convergence rounds therefore avoid rescanning the full edge set and do not churn unchanged working rows. PostgreSQL refreshes planner statistics for a sufficiently large temporary working table and refreshes them again after multiplicative growth. This avoids plans based on PostgreSQL's initial one-row estimate for a new temporary table. The policy is automatic, applies to WCC, label propagation, PageRank, and growing traversal frontiers, and is a no-op on SQLite. Iterative operations (WCC, label propagation, PageRank, and the working-table traversals) accept an opt-in `workingMemory` override of the session's `work_mem` for their rounds. When set, it is applied with `SET LOCAL work_mem` semantics inside the operation's own transaction — the session and server settings are never modified, and the override ends with the transaction. When omitted (the default), operations inherit the server's configured `work_mem`. Note that `work_mem` is a threshold each sort/hash operator (and each parallel worker) may allocate up to, **not** a per-operation budget: a single round can allocate several multiples of it, and concurrent algorithm calls multiply again. Raise it deliberately — for example `workingMemory: "64MB"` keeps whole-graph rounds from spilling their sorts to disk on large single-tenant analytical runs (such as LDBC SNB SF1-scale benchmarks) — rather than as a blanket setting on a shared cluster. The value must be a plain integer with a `kB`, `MB`, or `GB` suffix within PostgreSQL's accepted `work_mem` range (64kB to 2147483647kB) — both backends reject malformed or out-of-range values with the same error. SQLite validates and otherwise ignores it. Without `nodeKinds`, WCC seeds every visible node, so unrelated nodes still appear as singleton components. On heterogeneous graphs, pass the kinds that define the graph being analyzed. For example, `{ edges: ["knows"], nodeKinds: ["Person"] }` avoids seeding posts and comments while retaining isolated people. Temporal views expose the same facade with the coordinate sealed: ```typescript const historical = await store .asOf("2024-01-01T00:00:00.000Z") .algorithms.weaklyConnectedComponents({ edges: ["knows"] }); ``` ## Passing Nodes or IDs Every node-oriented algorithm accepts either a raw ID string or any object with an `id: string` field — `Node`, `NodeRef`, the lightweight records returned by traversals, and `store.subgraph()` results all work. WCC is a whole-graph operation and does not take a node identifier. Global PageRank is also whole-graph; personalized PageRank instead takes explicit `{ id, kind, weight? }` seed identities. ```typescript const alice = await store.nodes.Person.getById(aliceId); // All equivalent store.algorithms.canReach(alice, bobId, { edges: ["knows"] }); store.algorithms.canReach(alice.id, bobId, { edges: ["knows"] }); store.algorithms.canReach(aliceId, { kind: "Person", id: bobId }, { edges: ["knows"], }); ``` ## Direction and Cycles `direction: "both"` lets you treat a directed edge kind as undirected for reachability questions: ```typescript // Did Dave ever know Alice, regardless of who "added" whom first? const knew = await store.algorithms.canReach(dave, alice, { edges: ["knows"], direction: "both", }); ``` Cycles cannot multiply work: a node enters the visited set once, at its minimum depth. This differs from `.recursive()`, whose path-returning semantics use the per-path behavior described in [Cycle Detection](/queries/recursive#cycle-detection). The algorithm `cyclePolicy` option is therefore compatibility-only. ## Depth Limits `maxHops` is capped at `MAX_EXPLICIT_RECURSIVE_DEPTH` (1000). The default of 10 covers typical connectivity questions on most graphs. Within that bound, single-source traversal examines each reached node once and each incident edge at most once per direction, rather than enumerating paths. Each expanded level costs a database round trip (inline working relations are split to respect the backend bind limit). Transactional backends keep those statements in one snapshot and drop their temporary working table in a `finally` cleanup; non-transactional backends use their normal best-effort multi-statement behavior. The application-clock valid-time instant is pinned once for the whole traversal, and recorded-time views remain pinned to their requested recorded coordinate. ## Temporal Behavior Algorithms honor the same temporal model as the rest of the store. The default temporal mode is `graph.defaults.temporalMode` (typically `"current"`), and every algorithm accepts per-call `temporalMode` and `asOf` options. ```typescript // Default: uses graph.defaults.temporalMode — typically "current". await store.algorithms.shortestPath(alice, bob, { edges: ["knows"] }); // Snapshot at a specific point in time. Both nodes and edges must have // been valid at that timestamp to participate in the traversal. await store.algorithms.shortestPath(alice, bob, { edges: ["knows"], temporalMode: "asOf", asOf: "2023-01-15T00:00:00.000Z", }); // Include validity-ended (but not soft-deleted) rows — useful for // historical traversal without needing a specific timestamp. await store.algorithms.reachable(alice, { edges: ["knows"], temporalMode: "includeEnded", }); // Include soft-deleted rows too. Traversal can cross through tombstones. await store.algorithms.canReach(alice, ghost, { edges: ["knows"], temporalMode: "includeTombstones", }); ``` **Semantic rules:** - The temporal filter applies to **both nodes and edges** along the traversal. An edge can only be traversed if it passes the filter *and* its endpoint node passes too. - `asOf` is required when `temporalMode: "asOf"` and rejected (throws `ValidationError`) in every other mode — pinning an instant is only meaningful in `"asOf"` mode. - Temporal filtering is orthogonal to `cyclePolicy` — cycle detection operates on path membership, not on time. A node that was valid → ended → re-valid is not treated as "visited" just because it appears in two validity periods. - The shortest-path self-path short-circuit also respects the resolved mode: calling `shortestPath(a, a, ...)` returns `undefined` if node `a` does not pass the temporal filter, and a zero-hop result otherwise. ## End-to-End Example The runnable example [`examples/14-research-copilot.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/14-research-copilot.ts) combines every algorithm with semantic search and ontology-expanded topic matching over a corpus of landmark ML papers. It produces an explainable literature-review digest in one run against a single SQLite file — a good starting point for your own RAG + graph workloads. ## What's Not Included These algorithms cover shortest path (weighted and unweighted), reachability, neighborhoods, degree, weakly connected components, deterministic label propagation, and global and personalized PageRank. They do **not** cover: - Strongly connected components - Topological sort - Centrality measures beyond degree (betweenness, closeness, eigenvector) - Modularity-optimizing community detection such as Leiden or Louvain For those, export edges via `.query().traverse()` or `store.subgraph()` and use a specialized library such as [graphology](https://graphology.github.io/) in memory. See [Limitations](/limitations#graph-analytics-limits) for the full list of excluded analytics. # Graph Extensions > Extend a TypeGraph schema at runtime — durable, multi-process safe, with full Zod validation and unique-constraint enforcement. Graph extensions let your application declare new node and edge kinds **at runtime** — durable across restarts, with semantic parity to compile-time `defineNode` / `defineEdge`. The motivating use case: **agent-driven schema induction**, where an LLM proposes a typed schema from a corpus, an operator approves it, and the live graph immediately ingests under the new schema with no code change or restart. :::note[See it end-to-end] For a runnable scenario with an operator-approved agent in TypeScript, see [Agent-Driven Schema](/examples/agent-driven-schema). For the same loop driven by an open-weight LLM against public-record clinical data — with a repair loop and smoke-test pattern — see [`pdlug/typegraph-clinical-demo`](https://github.com/pdlug/typegraph-clinical-demo). ::: This guide covers the core verbs: | Verb | Purpose | | ------------------------------------------------ | ----------------------------------------------------------------- | | `defineGraphExtension` | Build a typed extension (pure value, no I/O) | | `store.evolve(extension)` | Atomically commit a new schema version with the extension applied | | `store.introspect()` | Snapshot the merged schema, persisted extension, version, and hash | | `store.materializeIndexes()` | Run declared `CREATE INDEX` DDL against the live database | | `store.deprecateKinds(...)` / `undeprecateKinds` | Soft-deprecate kinds for codegen / lint signaling | | `store.removeKinds(...)` | Remove graph-extension-declared kinds from the active schema | | `store.materializeRemovals()` | Delete rows queued by graph-extension-kind removal | For the schema-management primitives that graph extensions ride on top of, see [Schema Migrations](/schema-management) and [Evolving Schemas](/schema-evolution). ## When to use graph extensions Use them when **the kind set is not known at code time**: - Agent / LLM proposes a new typed schema from observed data. - Multi-tenant deployments where each tenant defines their own kinds. - ETL pipelines that ingest sources with shifting structure. - Plugins / extensions that contribute kinds at install time. For everything else — kinds you can declare in TypeScript at deploy time — use the compile-time DSL (`defineNode`, `defineEdge`, `defineGraph`). The compile-time path is type-safe end-to-end; graph extensions trade some of that type-safety for the ability to evolve without redeploying. ## A complete example ```ts import { z } from "zod"; import { createStoreWithSchema, defineGraph, defineNode, defineGraphExtension, } from "@nicia-ai/typegraph"; import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; // 1. Boot with a compile-time kind. const Document = defineNode("Document", { schema: z.object({ title: z.string(), body: z.string() }), }); const baseGraph = defineGraph({ id: "research_corpus", nodes: { Document: { type: Document } }, edges: {}, }); const { backend } = createLocalSqliteBackend(); const [store] = await createStoreWithSchema(baseGraph, backend); // 2. An agent proposes a new kind at runtime. const proposal = defineGraphExtension({ nodes: { Paper: { description: "An academic paper inferred from the corpus", properties: { title: { type: "string", minLength: 1 }, doi: { type: "string", minLength: 1 }, year: { type: "number", int: true, min: 1900, max: 2100 }, }, unique: [{ name: "paper_doi_unique", fields: ["doi"] }], }, }, indexes: [ { entity: "node", kind: "Paper", name: "paper_by_doi", fields: ["doi"], unique: true, }, ], }); // 3. Operator approves; commit atomically. const evolved = await store.evolve(proposal); // 4. Use the dynamic-collection accessor (the type system does not // widen for extension kinds — see "Reaching extension kinds" below). const papers = evolved.getNodeCollection("Paper")!; await papers.create({ title: "Attention is all you need", doi: "10.5555/3295222.3295349", year: 2017, }); ``` A complete runnable version is in [`examples/16-graph-extensions.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/16-graph-extensions.ts). ## Instantiate a graph template For many tenant graphs with the same final shape, register a verified schema once and instantiate each target as schema version 1. The template document stays in the database; instantiation sends only its id, the target graph id, and a client-computed content hash. ```ts const [source] = await createAdapterStoreWithSchema(baseGraph, adminBackend); const template = await registerGraphTemplate(adminBackend, { templateId: "research-corpus-v1", reconciled: source.reconciledSchema, }); const target = await instantiateGraph(adminBackend, { template, graphId: "tenant-123", }); // The returned target snapshot can create a request-scoped Store without // another schema read. `tenantGraph` is the same compile-time shape with id // "tenant-123". const store = createAdapterStore(tenantGraph, runtimeBackend, { reconciled: target.reconciled, }); ``` Instantiation is idempotent for the same template and target. A target already initialized with a different schema is refused. On PostgreSQL the clone statement takes the same graph advisory-lock key as schema commits; SQLite's schema and marker writes are serialized by its writer lock. Templates clone the source graph's graph-local runtime-contribution activation markers along with its schema row. Deployment-scoped physical attestations, such as the shared fulltext table marker, already belong to the database and are not duplicated per template target. A target can therefore be reopened with `createVerifiedStore` or `createVerifiedAdapterStore` from a later serverless isolate without another schema reconciliation or provisioning DDL step. Registration and instantiation are DML-only. They assert the deployment-wide base-schema marker and throw `BaseSchemaMigrationError` rather than attempting DDL when adoption is missing, stale, or newer than the running library. Run the normal privileged `createStoreWithSchema` boot or the published external migration before handing runtime requests to a DML-only role. Template eligibility follows the backend's capabilities. Embedding declarations are allowed when the backend is configured with `vector: false`, because that backend has no graph-scoped vector storage to clone. A vector-enabled backend continues to refuse embedding-bearing schema-only templates. ## The graph extension `defineGraphExtension` accepts a structured value describing the new kinds. The extension is JSON-serializable — that's load-bearing for durability (see [Restart parity](#restart-parity-the-load-bearing-invariant)). ### Document format versioning Every document carries a `version` field (currently `1`). The validator stamps the version automatically when you call `defineGraphExtension`, so consumer code never has to set it explicitly. Stored documents from before this field existed are treated as `version: 1` (the legacy default). The forward-compat policy: - **Additive minor changes** (new optional property modifier, new `format` value, new top-level slice within the same major) ride forward without bumping `version`. The validator does not reject unknown top-level keys, and the persistence-side zod is `.loose()` on every nested object — an older runtime reading a newer extension silently ignores unknown fields and continues working. - **Breaking changes** bump `version` to a higher major. An older runtime reading a higher-version extension fails with `GRAPH_EXTENSION_VERSION_UNSUPPORTED` and an actionable error pointing the operator at upgrading the library — there is no automatic downgrade path. The current major is exported as `CURRENT_GRAPH_EXTENSION_VERSION` for tooling that wants to pre-flight check. - **Legacy extensions** (committed before `version` existed) and extensions that explicitly omit `version` are interpreted as `LEGACY_GRAPH_EXTENSION_VERSION`, pinned permanently to `1`. This is deliberately distinct from `CURRENT`: when a future v2 ships, legacy v1 extensions continue parsing as v1 (so the version-mismatch path can route them through migration) rather than being silently re-classified as v2 by a default-equals-current rule. ```ts import { CURRENT_GRAPH_EXTENSION_VERSION, LEGACY_GRAPH_EXTENSION_VERSION, } from "@nicia-ai/typegraph"; console.log(CURRENT_GRAPH_EXTENSION_VERSION); // 1 (today; bumps with breaking changes) console.log(LEGACY_GRAPH_EXTENSION_VERSION); // 1 (always; the pre-versioning default) ``` ### Property types (the v1 subset) The following types are supported. The set is deliberately small so that LLM-induced schemas can be audited at a glance and so the persistence layer never has to reconstruct opaque Zod refinements from JSON. | Type | JSON shape | | --------- | ------------------------------------------------------------------------- | | `string` | `{ type: "string", minLength?, maxLength?, pattern?, format? }` | | `number` | `{ type: "number", int?, min?, max? }` | | `boolean` | `{ type: "boolean" }` | | `enum` | `{ type: "enum", values: ["a", "b", ...] }` | | `array` | `{ type: "array", items: }` | | `object` | `{ type: "object", properties: { foo: , ... } }` (one nesting only) | Supported string formats: `"datetime"`, `"uri"`, `"email"`, `"uuid"`, `"date"`. These route to the corresponding Zod factories (`z.iso.datetime()`, `z.url()`, `z.email()`, `z.uuid()`, `z.iso.date()`). Modifiers available on every property: - `optional?: true` — omits from the required set. - `description?: string` — surfaces in tooling. - `searchable?: SearchableModifier` — string only; routes through the `searchable()` brand for fulltext indexing. - `embedding?: { dimensions: number }` — array-of-number only; routes through the `embedding()` brand for vector search. ### Unique constraints Pass `unique: [{ name, fields, ... }]` per kind. The `name` is required (used as the diffing identity key) and must be unique within the kind. Supports `scope`, `collation`, and a restricted `where` clause limited to `isNull` / `isNotNull` (the only operations round-trippable through the persisted form). ### Relational indexes Pass `indexes: [...]` at the document top level to declare relational indexes for graph-extension or compile-time host kinds: ```ts const proposal = defineGraphExtension({ nodes: { Paper: { properties: { doi: { type: "string" }, title: { type: "string", searchable: { language: "english" } }, }, }, }, indexes: [ { entity: "node", kind: "Paper", name: "paper_by_doi", fields: ["doi"], unique: true, }, ], }); ``` Index `name`s are unique across the merged graph. A graph-extension index that reuses a compile-time index name, or a later extension that reuses an earlier graph-extension index name for a different declaration, is rejected. ### Edges Edges follow the same shape as nodes and declare endpoints using kind names. An array-valued `to` permits every combination of `from` and `to` kinds. A map-valued `to` restricts targets by source kind: ```ts const proposal = defineGraphExtension({ edges: { assignedTo: { from: ["Employee", "Student"], to: { Employee: ["Department"], Student: ["Course"], }, properties: {}, }, }, }); const evolved = await store.evolve(proposal); ``` All endpoint kinds must resolve in the merged graph when `evolve()` runs. Map keys must exactly cover `from`, and each target array must be nonempty. The map is persisted and restored on restart, preserving the same runtime pair validation as [compile-time declarations](/core-concepts#source-dependent-targets). Adding allowed pairs broadens an extension edge. Removing a pair tightens it, even if the overall source and target kind sets stay the same. Tightening currently requires the **entire edge kind** to be empty, not just the removed pair. ### Ontology Pass `ontology: [{ metaEdge, from, to }, ...]` to declare ontology relations between kinds (subClassOf, partOf, etc.). The meta-edge name must match a meta-edge known to the merged graph. ## `store.evolve(extension, options?)` ```ts const evolved = await store.evolve(extension); const evolved = await store.evolve(extension, { ref }); const evolved = await store.evolve(extension, { eager: {} }); ``` `evolve` is the consumer-facing primitive that drives extension. It: 1. **Catches up to persisted state** — folds any persisted extension and deprecation set into the local baseline so a stale store doesn't trample another writer's progress. 2. **Merges** the new extension into the baseline graph. Re-declaring an existing extension kind with the same shape is a no-op; with a non-additive change against existing rows it throws `IncompatibleChangeError` (code `INCOMPATIBLE_CHANGE`). 3. **Atomically commits** a new schema version via `commitSchemaVersion` (CAS on the active version). 4. **Returns the resulting `Store`** carrying the extended graph. The type parameter `G` does NOT widen — see [Reaching extension kinds](#reaching-extension-kinds-from-the-type-system) below. :::caution[Use the returned Store] `Store` instances are immutable schema snapshots. After `evolve()` resolves, use its returned Store for **every subsequent Store operation in the same request**, not just operations involving newly added kinds. A Store captured before the commit remains pinned to the previous schema version, so its next managed write fails the schema-version fence. ::: ### Plan outside and apply inside a caller-owned transaction `planEvolution(extension)` prepares a named schema change before the caller opens its write transaction. It returns an immutable `"noop"` or `"change"` plan with `graphId`, `baseline: { version, hash }`, and `result: { version, hash }`. Change plans expose an ordered `requirements` array whose entries name new-kind additions, empty-kind checks, vector slots, and identity work. A `new-kind` entry describes a graph delta; it does not indicate that a removal is queued or direct callers to run `materializeRemovals()`. The plan is opaque and bound to the loaded TypeGraph module: it cannot be serialized, cloned, or reconstructed. It can be passed between compatible Stores for the same graph that use the same loaded module; apply still checks the active graph and fenced baseline version/hash. The default `{ source: "database" }` reloads the active schema. `{ source: "cached" }` uses a previously loaded planning snapshot on the same Store; a cached plan is not fresh database evidence. A stale baseline is refused during apply, so retry by rolling back the whole caller transaction and replanning outside it. Use `withEvolvedTransaction(nativeTx, plan, callback, { waitBudgetMs })` when the schema change, TypeGraph writes, and application SQL must share the caller's commit. Enter this boundary before other TypeGraph callbacks on the same native transaction. The callback receives an evolved transaction context, including its reads, collections, and supported composition operations. It does not receive a replacement root Store. ```ts const cachedStore = store; const ref = { current: cachedStore }; const plan = await cachedStore.planEvolution(proposal); const writerStore = cachedStore.withBackend(writerBackend); const provisional = await db.transaction(async (nativeTx) => { const outcome = await writerStore.withEvolvedTransaction( nativeTx, plan, async (tx) => { const person = await tx.nodes.Person.create({ name: "Ada" }); return person.id; }, plan.status === "change" ? { waitBudgetMs: 5000 } : undefined, ); await nativeTx.insert(applicationEvents).values({ personId: outcome.result, schemaVersion: outcome.receipt.schema.version, }); return outcome; }); const refreshed = await cachedStore.refreshSchema({ minVersion: provisional.receipt.schema.version, ref, }); ``` The callback and receipt finish before the outer SQL transaction commits. The receipt's schema version/hash and recorded anchor are provisional until that commit succeeds. A callback failure must reject the outer transaction; catching it and committing does not prove rollback. Callback contexts and queries built from them expire when the callback returns, including after a failure. Use `refreshSchema()` only after awaiting a successful outer commit. When the cached Store already matches `minVersion`, refresh returns it without a read; that shortcut does not check for a newer database version. Otherwise refresh reads the active schema, accepts a newer committed version, and refuses a missing or older one. It applies no extension or storage provisioning. The adopted apply path supports metadata-only changes, required-empty checks, and transactional identity and vector provisioning. Configure the adapter with `schemaProvisioning: "transactional"` on a privileged connection when the plan names new vector slots or identity work. The adapter's default DML-only policy refuses those requirements before the schema fence, DDL, callback, or mutation. The privileged path rechecks storage on the caller's fenced session, provisions the required relations and vector contribution markers there, and rolls them back with the outer transaction. Missing bootstrap tables still refuse; run the normal bootstrap before serving adopted evolution requests. The apply path uses a bounded schema fence, baseline validation, and a version CAS; it can issue multiple statements. `waitBudgetMs` bounds fence acquisition for change plans; no-op plans refuse an explicit `waitBudgetMs`. On timeout, roll back and retry the complete native transaction. Do not pass `ref` or eager-index options to `withEvolvedTransaction`; they cannot be honored before the outer commit and are refused. Generic and concurrent eager indexes remain explicit maintenance after commit: call `materializeIndexes()` on the refreshed Store when they are needed. When a wiring pass produces a no-op, ordinary recorded adoption is sufficient and never takes the exclusive evolution fence. Reconcile only if the Store is behind the named snapshot; matching versions make this refresh a cached read: ```typescript if (plan.status === "noop") { const current = await store.refreshSchema({ minVersion: plan.baseline.version }); await db.transaction((nativeTx) => current.withRecordedTransaction(nativeTx, async (tx) => { await tx.nodes.Person.create({ name: "Ada" }); }), ); } ``` ### The `ref` pattern `Store` is immutable by construction — `evolve()` returns the Store for the resulting schema, using a fresh instance when the schema advances. Long-lived consumer code that holds the Store in a singleton needs a way to re-bind it before the operation completes. Pass `options.ref: { current: store }` (a `StoreRef>`): ```ts const ref: StoreRef> = { current: store }; const evolved = await ref.current.evolve(extension, { ref }); // `ref.current === evolved`; use either reference from here on. const papers = ref.current.getNodeCollectionOrThrow("Paper"); await papers.create({ title: "...", doi: "...", year: 2024 }); ``` Capture `ref.current` once at request entry only when that request will not change the schema. If it does call `evolve()`, switch to the returned Store or dereference `ref.current` again after the call. `StoreRef` is structurally just `{ current: T }`; the library doesn't provide a factory because the consumer composes the handle themselves (it could be a Vue ref, MobX observable, Zustand atom, etc.). The ref covers schema changes made by calls that receive it. If another process or isolate can advance the schema, probe the committed version before reusing a cached Store and perform a verified open when it changes. See [Per-request connections: cache the verified Store](/integration#per-request-connections-cache-the-verified-store) for the complete `getCommittedSchemaVersion()` recipe. ### Eager materialization Pass `eager: {}` to materialize indexes immediately after the schema commit: ```ts const evolved = await store.evolve(extension, { eager: {} }); ``` Or pass options for finer control: ```ts // Restrict to the extension kind whose index was declared in the // proposal above. const evolved = await store.evolve(extension, { eager: { kinds: ["Paper"], stopOnError: true }, }); ``` Omit `eager` to skip materialization and run `materializeIndexes()` later. Per-index failures throw `EagerMaterializationError` AFTER the new `Store` is constructed and `ref.current` is updated, so the caller can recover via the ref handle. The schema commit is **not** rolled back if materialization fails — eager is convenience, not a transaction. ```ts const ref = { current: store }; try { await store.evolve(extension, { ref, eager: {} }); } catch (error) { if (error instanceof EagerMaterializationError) { // Schema is committed; ref.current is the new store. log.warn( { failed: error.failedIndexNames }, "indexes did not materialize; will retry", ); await ref.current.materializeIndexes(); } else { throw error; } } ``` ## Reaching extension kinds from the type system TypeScript can't see kinds that don't exist at compile time. The `Store` returned by `evolve()` keeps the same generic parameter as the original — `evolved.nodes.Paper` would not type-check. The escape hatch is `store.getNodeCollection(kind)` and `store.getEdgeCollection(kind)`, which return a typed `DynamicNodeCollection` / `DynamicEdgeCollection`: ```ts const papers = evolved.getNodeCollection("Paper"); if (papers === undefined) { throw new Error("Paper kind not registered on this store"); } await papers.create({ title: "...", doi: "...", year: 2024 }); const all = await papers.find({}); ``` The throwing variants `getNodeCollectionOrThrow(kind)` / `getEdgeCollectionOrThrow(kind)` are the right call when the caller already knows the kind has been evolved onto the store — they raise `KindNotFoundError` with the offending `kindName`, `entity`, and host `graphId` instead of returning `undefined`, so a typo fails loudly at the call site rather than crashing later on `papers!.create(...)`. `DynamicNodeCollection` exposes the same CRUD surface as `store.nodes.X` — `create`, `getById`, `find`, `update`, `delete`, etc. — but with `DynamicNode` element types since the specific Zod schema isn't visible to TypeScript at the call site. In TypeScript, nodes returned by a dynamic collection carry the nominal `DynamicNode` type, preserving the requested kind literal. That proof lets runtime kinds participate directly in Operational Identity without weakening compile-time references to arbitrary string kinds: ```ts const paper = await evolved .getNodeCollectionOrThrow("Paper") .create({ title: "Runtime schemas" }); await evolved.identity.assertSame(document, paper); ``` Identity reads can consequently return `IdentityNodeReference` values for either compile-time or runtime kinds. See [Operational Identity](/identity). For consumers that need the live Zod schema itself — MCP tool wrappers that validate inputs before forwarding to `collection.create`, or agent prompts that want richer JSON Schema than `introspect()` exposes — `store.getNodePropsSchema(kind)` / `getNodePropsSchemaOrThrow(kind)` (and the edge counterparts) return the exact `z.ZodObject` the store uses internally. Identity holds: `evolved.getNodePropsSchema("Paper")` is the same instance the store parses against on `papers.create(...)`. ```ts import { z } from "zod"; const schema = evolved.getNodePropsSchemaOrThrow("Paper"); const parsed = schema.parse(input); // same Zod issues as papers.create surfaces const jsonSchema = z.toJSONSchema(schema); // for MCP tool descriptions ``` These accessors return only the props validator. Operation-level checks — uniqueness, endpoint resolution, temporal validity, backend constraints — still run only through `collection.create` / `update`. See [Dynamic Props Schema Access](/schemas-stores#dynamic-props-schema-access) for the full reference. For codegen consumers, the kind set is reachable by iterating the registry's `nodeKinds` and `edgeKinds` maps: ```ts const allNodeKinds = [...store.registry.nodeKinds.keys()]; const allEdgeKinds = [...store.registry.edgeKinds.keys()]; const personType = store.registry.getNodeType("Person"); // NodeType | undefined ``` `KindRegistry` also exposes `hasNodeType(name)` / `hasEdgeType(name)` for existence checks. ### Querying extension kinds `store.query()` requires every `from` / `traverse` / `to` kind to be a compile-time literal in `Store`. The string-keyed siblings `fromDynamic` / `traverseDynamic` / `optionalTraverseDynamic` / `toDynamic` admit kinds added via `evolve()` so an MCP server (or any caller working from kind names in a string variable) can build typed multi-hop traversals without `as any`: ```ts const rows = await store.query() .fromDynamic("Paper", "p") .traverseDynamic("authoredBy", "a") .toDynamic("Author", "u") .whereNode("p", (p) => p.field("year").number().gte(2020)) .select((ctx) => ({ paper: ctx.p, author: ctx.u, edge: ctx.a })) .execute(); ``` Each method runtime-validates against the registry: a typo throws `KindNotFoundError`, and a `toDynamic` target that isn't a valid endpoint for the current edge / direction throws `EndpointError`. Compile-time `from` / `traverse` / `to` are unchanged. When the extension document is available in typed code, mint Store-bound runtime-kind evidence from the exact persisted definition. The token narrows collections and query aliases without asking callers to restate the schema in Zod: ```ts const tagKind = store.runtimeNodeKind("Tag", extension.nodes.Tag); const taggedWithKind = store.runtimeEdgeKind( "taggedWith", extension.edges.taggedWith, ); const tags = store.getNodeCollectionOrThrow(tagKind); const rows = await store.query() .fromDynamic(tagKind, "tag") // ctx.tag has the declared Tag fields .traverseDynamic(taggedWithKind, "edge") // ctx.edge is narrowed too .toDynamic("Document", "document") .select((ctx) => ({ label: ctx.tag.label, weight: ctx.edge.weight })) .execute(); ``` Definitions, rather than hand-authored Zod schemas, are the type evidence: TypeGraph compares the complete graph-extension declaration that TypeScript infers against the definition persisted for that kind. This keeps refinements such as enums, optionality, arrays, and numeric constraints on one authoritative surface. Tokens are bound to the issuing Store and active schema hash; use a fresh token after reopening or evolving a Store. Token lookups intentionally use the throwing `getNodeCollectionOrThrow` / `getEdgeCollectionOrThrow` variants because valid Store-issued evidence already proves that the kind exists. Predicate accessors on dynamic aliases use a `.field(name)` discriminator: - `BaseFieldAccessor` methods (`eq`, `isNull`, `in`, `notIn`) are available directly on `field("name")`. - Type-specific predicates sit behind a discriminator method that asserts the field's type — `.string()` / `.number()` / `.date()` / `.array()` / `.object()` / `.embedding()`. Each validates against the registered Zod schema and throws `TypeError` on mismatch, so `field("year").string()` against a number field is caught at query-build time, not as a silent "method is undefined" later. - `.field("missing")` throws when the property isn't on the schema. #### Mixed typed and dynamic aliases Typed and dynamic aliases interleave freely in one query. The predicate accessor is resolved per alias — a typed alias keeps its narrow `StringFieldAccessor` etc., while a dynamic alias gets `.field()`: ```ts const rows = await store.query() .from("Document", "d") // typed compile-time kind .traverseDynamic("taggedWith", "e") // runtime edge .toDynamic("Tag", "n") // runtime target .whereNode("d", (d) => d.title.eq("the doc")) // typed: direct .whereNode("n", (n) => n.field("label").string().eq("research")) // dynamic: discriminator .select((ctx) => ({ doc: ctx.d, tag: ctx.n })) .execute(); ``` A typed `traverse("typedEdge", "e")` followed by `.toDynamic(target, "n")` keeps the edge alias `e` typed — `e.role.eq(...)` works directly, no discriminator needed. Only the dynamic-declared aliases use `.field()`. #### Optional dynamic traversal `optionalTraverseDynamic` is the LEFT-JOIN sibling — papers without authors still surface, with the edge and target aliases as `undefined`: ```ts const rows = await store.query() .fromDynamic("Paper", "p") .optionalTraverseDynamic("authoredBy", "a") .toDynamic("Author", "u") .select((ctx) => ({ paper: ctx.p, author: ctx.u, edge: ctx.a })) .execute(); // row.author and row.edge are undefined for papers without an authoredBy edge. ``` ### Search facade The `store.search` facade — `fulltext`, `vector`, `hybrid`, and `rebuildFulltext` — accepts any registered kind, compile-time or runtime, with no type cast. The hit's `node` type narrows to the concrete typed node only when the kind literal is statically known in `Store`; extension kinds widen to the base `Node`. Misspelled kind names throw `KindNotFoundError` at the call site instead of returning empty results. ```ts // Compile-time kind: hit.node.title is narrowed. const compileTimeHits = await store.search.fulltext("Document", { query: "climate", limit: 10, }); // Extension kind: same call shape, no cast. hit.node is the base // `Node` shape since "Paper" isn't in the static `G`. const runtimeHits = await store.search.fulltext("Paper", { query: "attention transformer", limit: 10, }); ``` ## `store.introspect()` `introspect()` returns a frozen snapshot of the merged schema and the durable-state metadata the store has loaded so far. Its shape: | Field | Type | Notes | | ---------------------- | ------------------------------------- | ---------------------------------------------------------------------------------------------- | | `graphId` | `string` | The graph's stable id. | | `annotations` | `GraphAnnotations \| undefined` | Merged graph-scoped metadata from `defineGraph` and runtime extensions. | | `kinds` | `readonly KindIntrospection[]` | Merged node kinds with `origin: "compile-time" \| "runtime"`, description, annotations, etc. | | `edges` | `readonly EdgeIntrospection[]` | Merged edge kinds with the same origin discriminator and endpoint information. | | `ontology` | `readonly OntologyIntrospection[]` | Ontology relations declared on either tier. | | `deprecatedKinds` | `ReadonlySet` | Kinds flagged via `deprecateKinds(...)`. Informational, not a gate. | | `extension` | `GraphExtension \| undefined` | The persisted graph-extension document, or `undefined` when no extensions have been committed. | | `schemaVersion` | `number \| undefined` | Active schema version on the backend. `undefined` until the first commit. | | `schemaHash` | `string \| undefined` | Hash of the active schema document. `undefined` under the same condition. | ```ts const intro = store.introspect(); console.log(intro.schemaVersion); // e.g. 2 console.log(intro.annotations?.displayName); console.log(intro.extension?.nodes?.Paper); // ExtensionNodeDef or undefined console.log([...intro.deprecatedKinds]); // ["LegacyDocument"] ``` The `extension` field round-trips: passing it back through `defineGraphExtension(intro.extension!)` and `evolve()` against an empty graph reconstructs the same extension kinds. For schema tooling that has an extension document but no Store, call `introspectGraphExtension(extension)`. It compiles the extension through the same TypeGraph compiler used by `evolve()` and returns `kinds` and `edges` with JSON Schema `properties`, descriptions, annotations, and endpoint names. The result describes only the supplied document; it has no graph ID, committed schema version, or schema hash. ```ts import { introspectGraphExtension } from "@nicia-ai/typegraph"; const declaration = introspectGraphExtension(extension); console.log(declaration.kinds[0]?.properties); ``` Graph extensions may also carry graph-scoped annotations: ```ts const extension = defineGraphExtension({ annotations: { displayName: "Customer knowledge", capabilities: { semanticSearch: true }, }, }); ``` Annotation keys are shallow-merged. A later extension replaces the complete value of each key it supplies; it does not recursively merge nested objects. ## Population statistics and stored-data validation `await store.describe()` pairs the merged schema introspection with current population statistics. It returns node and edge counts for every declared kind plus present, explicit-null, and non-null counts (and non-null coverage) for each directly addressable declared property: ```ts const description = await store.describe(); const people = description.statistics.nodes.find( (entry) => entry.kind === "Person", ); console.log(people?.count); console.log( people?.properties.find((property) => property.path === "/email")?.coverage, ); ``` The schema coordinate includes the active schema version and hash when present and a `schemaFence`. TypeGraph reads that coordinate before and after the bounded, sequential SQL aggregate statements and refuses the result if it changed. Node and edge properties are queried separately, and wide schemas are split into fixed-width path batches. The database, rather than TypeGraph's JavaScript process, computes all counts. Coverage follows ordinary nested JSON Schema `properties`; TypeGraph intentionally does not invent population semantics through `$ref`, unions, intersections, arrays, or conditionals. `validateStore()` remains authoritative for those schemas. Concurrent writes can affect different `describe()` path batches differently; the schema fence detects schema changes, not data changes. Use `validateStore()` to find rows that no longer satisfy a kind's current declared Zod schema, for example after tightening a rule around existing data: ```ts let cursor: string | undefined; do { const page = await store.validateStore({ entity: "node", kind: "Person", pageSize: 250, ...(cursor === undefined ? {} : { cursor }), }); for (const failure of page.violations) { console.log(failure.id, failure.path, failure.reason); } cursor = page.nextCursor; } while (cursor !== undefined); ``` Undeclared properties are healthy semi-structured state and are never reported as violations, including when the authored Zod object is strict. Each failure names the record id, JSON-pointer path, top-level property when applicable, Zod issue code, and reason. `pageSize` is the number of records scanned, not a cap on violations: one record can contribute several Zod issues. Each request performs a bounded SQL keyset scan (`LIMIT pageSize + 1`) and reports `scannedCount`; it never materializes or rescans the complete kind just to continue. Cursors bind the entity, kind, schema fence, and last scanned id. A schema change throws `StoreAnalysisCursorStaleError`. Data pages are deliberately live rather than a claimed cross-request snapshot, so concurrent inserts, updates, and deletes can affect later pages. Each page still reads the schema coordinate before and after its data statement and refuses a concurrent schema flip. Both analysis methods are current-only. They are absent from `StoreView`; recorded/as-of population analysis is deferred until it can be backed by an equally explicit temporal contract. Root-store calls use sequential SQL statements plus schema bracketing, so they also work on non-interactive transactional adapters. Transaction callbacks expose the same methods through their pinned session. Choose `repeatable_read` or `serializable` when every `describe()` aggregate or every `validateStore()` page must observe one database snapshot, and consume all validation pages before the callback returns: ```ts await store.transaction( async (tx) => { const description = await tx.describe(); let cursor: string | undefined; do { const page = await tx.validateStore({ entity: "node", kind: "Person", ...(cursor === undefined ? {} : { cursor }), }); cursor = page.nextCursor; } while (cursor !== undefined); return description; }, { isolationLevel: "repeatable_read", accessMode: "read_only" }, ); ``` Calling root `store.describe()` from inside a transaction callback still does not join that transaction; use `tx.describe()` or `tx.validateStore()` for the bound-session behavior. ## `store.materializeIndexes(options?)` ```ts const result = await store.materializeIndexes(); // Restrict to specific compile-time or extension kinds. const result = await store.materializeIndexes({ kinds: ["Paper"] }); const result = await store.materializeIndexes({ stopOnError: true }); ``` `materializeIndexes` runs `CREATE INDEX` DDL for the indexes declared on the merged graph and tracks per-deployment status in `typegraph_index_materializations`. It's a separate verb from `evolve()` because: - DDL is **per-database**, not per-graph (two replicas of the same `schema_doc` are still two databases — DDL has to run on each). - Postgres uses `CREATE INDEX CONCURRENTLY` so live tables never take an `AccessExclusiveLock`. CIC cannot run inside a transaction, which is why `materializeIndexes` runs at the top-level backend, never inside `transaction()`. - Best-effort by default: per-index failures land in the result with the captured `Error` and the loop continues. Pass `stopOnError: true` to halt on the first failure. The returned `MaterializeIndexesResult` has one entry per declared index with `status: "created" | "alreadyMaterialized" | "failed" | "skipped"`. The `skipped` status surfaces when the backend recognizes the declaration but has no separate ANN index to build for it — e.g. sqlite-vec (KNN lives in the `vec0` virtual table), SQLite without a vector engine, or `embedding(dims, { indexType: "none" })` opting out of automatic materialization. Graph-extension-declared relational indexes use the same declaration shape as compile-time `defineNodeIndex` / `defineEdgeIndex`, but in a JSON-serializable form. They are persisted in `schema_doc.extension`, re-derived on restart, and surface in `store.graph.indexes` with `origin: "runtime"`. ### Vector indexes Vector indexes are **auto-derived** from `embedding()` brands on both compile-time and extension node kinds. Every top-level node field declared with `embedding(dims, opts?)` produces one `VectorIndexDeclaration` that flows through `materializeIndexes()` like any relational index. No extra wiring required. ```ts const Document = defineNode("Document", { schema: z.object({ title: z.string(), // Auto-derives a cosine HNSW vector index with pgvector // defaults (m=16, ef_construction=64). embedding: embedding(384), }), }); // Customize the auto-derived index by passing options at the brand. const Image = defineNode("Image", { schema: z.object({ embedding: embedding(512, { metric: "l2", m: 32, efConstruction: 100 }), }), }); // Opt out of automatic materialization while keeping the embedding. const Manual = defineNode("Manual", { schema: z.object({ embedding: embedding(384, { indexType: "none" }), }), }); ``` Embeddings live in per-`(graphId, nodeKind, fieldPath)` typed tables named `tg_vec___` (each carrying the field's fixed dimension) — there is no single shared embeddings table. The privileged migrator provisions each table plus a durable marker: at boot via `createStoreWithSchema`, and for a field a runtime `evolve()` introduces, by that `evolve()` call. The runtime hot path then asserts the marker (never DDL), so a least-privilege role can read/write embeddings. On `materializeIndexes()`: - Postgres with pgvector: emits `CREATE INDEX ... USING hnsw ...` (or `ivfflat`) on the field's per-`(graphId, kind, field)` vector table and reports `created`. - SQLite with `sqlite-vec`: KNN lives in the `vec0` virtual table, so there's no separate ANN index to build; declarations report `skipped` (with a reason), not `failed`. - libSQL / Turso: the DiskANN index is created via the strategy's own DDL (`libsql_vector_idx` + `vector_top_k`). - SQLite without a vector engine: declarations report `skipped` with a reason indicating the backend lacks vector support. The vector declaration's identity key within a single graph is `(kind, fieldPath)` — v1 allows at most one vector index per (kind, field) pair. The auto-derived deterministic declaration name is `tg_vec_{kind}_{field}_{metric}` — clean and scannable for inspection in `pg_indexes` and result entries. Changing the metric requires a different declaration name and explicit re-materialization. Cross-graph disambiguation lives at the materialization boundary, not in the declaration name. Vector status rows in `typegraph_index_materializations` are keyed on the compound `{graphId}::{declaration.name}` for both auto-derived and explicit `VectorIndexDeclaration` entries — so two graphs reusing the same declaration name (whether auto-derived from the same kind/field or constructed explicitly via `defineGraph({ indexes: [...] })`) don't collide in the status table. Each graph's `materializeIndexes()` call creates its own physical pgvector index on that graph's per-`(graphId, kind, field)` vector table and records its own status row. ### Fulltext indexes (out of scope for v1) Fulltext indexes are NOT in the unified declaration channel for v1. The fulltext table's canonical index (Postgres GIN on `tsv`, SQLite FTS5 virtual table) is created with the table itself by `bootstrapTables` per the active `FulltextStrategy`. Per-kind fulltext indexes are an "advanced strategy" surface that doesn't fit the relational-style declaration model and is reserved for future work. ### Caveats (Postgres) - `IF NOT EXISTS` does not validate shape — only that something with that name exists. Drift detection uses TypeGraph's recorded signature, not PG metadata. Signature mismatch surfaces as `failed` with a `different signature` message. - Failed `CONCURRENTLY` builds leave invalid indexes (`pg_index.indisvalid = false`). v1 surfaces this as a `failed` result; the operator drops the invalid index manually before retry. ## `store.deprecateKinds(...)` / `undeprecateKinds(...)` ```ts await store.deprecateKinds(["LegacyDocument"]); console.log([...store.introspect().deprecatedKinds]); // ["LegacyDocument"] await store.undeprecateKinds(["LegacyDocument"]); ``` Soft-deprecation surfaces in `store.introspect().deprecatedKinds: ReadonlySet` for introspection (codegen, UI tooling, lints) but does not gate reads, writes, or queries. Bumps the schema version like any other change; idempotent — re-deprecating an already-deprecated kind is a no-op. Use cases: - Codegen routes around deprecated kinds when generating new client code. - Lint rules flag new code that touches deprecated kinds. - UI tooling hides deprecated kinds from picker menus. ## `store.removeKinds(...)` / `materializeRemovals()` `removeKinds()` removes graph-extension-declared kinds from the active schema. It is intentionally two-phase: 1. **Schema commit.** `removeKinds(names)` rewrites the persisted graph extension without the named graph-extension kinds, cascades extension edges and ontology relations that can no longer resolve, and commits a new schema version with CAS. 2. **Data cleanup.** `materializeRemovals()` deletes rows for removed node and edge kinds on the current deployment. ```ts const withoutPaper = await evolved.removeKinds(["Paper"]); await withoutPaper.materializeRemovals(); ``` Pass `{ eager: {} }` to run cleanup inline after the schema commit: ```ts const withoutPaper = await evolved.removeKinds(["Paper"], { eager: {} }); ``` Removing an embedding field from a surviving kind orphans its per-`(graphId, kind, field)` `tg_vec_*` table; `materializeRemovals()` reclaims it and reports the count in `MaterializeRemovalsResult.reclaimedVectorFields`. For source-dependent edges, removing a node kind removes its source entry and any target references to it. A source entry whose targets are exhausted is also removed. The edge kind survives while another valid pair remains; it is cascaded only when no pairs remain. For example, removing `Course` from the `assignedTo` extension above preserves `Employee → Department`. Removal only applies to graph-extension-declared kinds. Removing a compile-time kind throws `RemoveCompileTimeKindError`; deploy new TypeScript code for compile-time schema removal. Removing a graph-extension kind that is still referenced by a compile-time edge or ontology relation throws `KindHasReferentsError`, because TypeGraph cannot rewrite your compiled graph for you. ## Restart parity (the load-bearing invariant) The graph extension is the **durable source of truth**. Every call to `evolve()` persists the merged document into `schema_doc.extension`. On startup, `createStoreWithSchema()` reads it back, runs the same compiler, and reconstructs identical Zod-bearing `GraphDef`. Net: an extension kind defined via `evolve()` is indistinguishable from a compile-time kind after restart. Verify this in your own tests: ```ts const [store] = await createStoreWithSchema(baseGraph, backend); const evolved = await store.evolve(proposal); await evolved.getNodeCollection("Paper")!.create({ title: "...", doi: "...", year: 2024 }); // Different process / different deployment / fresh store... const [restored] = await createStoreWithSchema(baseGraph, backend); expect(restored.registry.hasNodeType("Paper")).toBe(true); const all = await restored.getNodeCollection("Paper")!.find({}); expect(all).toHaveLength(1); ``` ## Multi-process safety Concurrent writers compete on the `commitSchemaVersion` CAS. One wins; the loser sees one of two errors with very different recovery semantics: - **`StaleVersionError`** — the local view of the active version is out of date. Routine race signal: refetch and retry. - **`SchemaContentConflictError`** — a different writer wrote a row at the same version with a different content hash. NOT a routine race. Two writers tried to commit semantically different schemas at the same version, which means one of them is operating on an inconsistent view of the world. Surface to the operator; do not blindly retry. Retry recipe (only catches `StaleVersionError`): ```ts async function evolveWithRetry( ref: StoreRef>, extension: GraphExtension, attempts = 3, ): Promise> { for (let attempt = 0; attempt < attempts; attempt++) { try { return await ref.current.evolve(extension, { ref }); } catch (error) { if (error instanceof StaleVersionError) { // Refetch happens implicitly inside evolve()'s next call — // catch-up auto-merges the persisted state into the local // baseline, so the next attempt diffs against fresh state. continue; } // SchemaContentConflictError, GraphExtensionValidationError, // EagerMaterializationError, etc. all surface to the caller — // they require operator intervention or different handling, not // blind retry. throw error; } } throw new Error(`Failed to evolve after ${attempts} attempts`); } ``` The internal `#catchUpToStored` step inside `evolve()` (and `deprecateKinds`, `materializeIndexes`) folds the persisted graph-extension document and deprecation set into the local baseline before computing the next state, so a stale store applying an extension on top of an out-of-date baseline doesn't trample another writer's progress. ## Trust boundary When the graph extension originates from an **untrusted source** — an LLM completion, user input, an external API — treat it as untrusted data. Specifically: - **Validation runs at the boundary.** `defineGraphExtension(doc)` rejects any input that doesn't match the v1 subset (`GraphExtensionValidationError` with per-issue paths). Don't skip this step. If you want Result-style handling for untrusted JSON, call `validateGraphExtension(raw, { strict: true })` and surface the structured issues before calling `evolve()`. - **Property types are deliberately small.** The supported set excludes things like `bigint`, `Date`, custom Zod refinements, and arbitrary functions. An LLM cannot inject executable code by proposing an extension document. - **Operator approval is your gate.** The library doesn't enforce human-in-the-loop — your application does. Show the diff to a human before calling `evolve()`. - **Persisted documents are part of your data.** They're stored in `schema_doc` along with every other schema artifact; back them up, audit them, version-control them. ## Out of scope for v1 - **Fulltext index unification.** Vector indexes flow through the unified channel (auto-derived from `embedding()` brands). Fulltext is still per-strategy: the GIN / FTS5 index is created with the fulltext table at `bootstrapTables` time. Per-kind fulltext indexes are reserved for future work. - **Multiple vector indexes per (kind, field).** v1 allows at most one. To use a different metric for the same field, use a different field name or wait for v2. - **Hard-blocking reads/writes on deprecated kinds.** Deprecation is informational. If you want strict enforcement, wrap collection access yourself. - **Auto drop+recreate on signature drift.** `materializeIndexes` surfaces drift as a `failed` result; manual remediation is required to avoid risky lock semantics. ## See also - [Schema Migrations](/schema-management) — the lower-level primitives `evolve()` rides on. - [Evolving Schemas](/schema-evolution) — recipes for compile-time schema changes. - [Errors](/errors) — `EagerMaterializationError`, `GraphExtensionValidationError`, `StaleVersionError`, `SchemaContentConflictError`. # Graph Merge > Branch a TypeGraph store, let many writers edit it independently, and fold their work back into one canonical graph with deterministic entity resolution, conflict reporting, edge repointing, and provenance. Graph Merge turns a TypeGraph store into something you can **fork, edit in parallel, and reconcile** — the way you already fork, branch, and merge code. Several writers (agents, importers, reviewers, background workers) each build graph changes in isolation, and a single deterministic step folds them back into one canonical graph: duplicate entities are resolved, edges are repointed onto the survivors, disagreements are surfaced (never silently overwritten), and you get a full report of what happened and who contributed it. It ships as a core package subpath: ```typescript import { branch, merge } from "@nicia-ai/typegraph/graph-merge"; ``` Everything here is defined over ordinary TypeGraph stores, schemas, indexes, backends, and ontology semantics — there is no separate service to run. ## What you can build Graph Merge exists because "append everything" is the wrong default for graphs: it produces duplicate entities and dangling relationships. With a real merge primitive you can build: - **Multi-agent knowledge-graph construction.** Run N extraction agents in parallel, each on its own branch, then merge. The same real-world entity discovered by three agents collapses to one canonical node; every agent's edges follow it; disagreements come back as conflicts to adjudicate. - **Parallel ETL / import reconciliation.** Ingest an EHR export, a claims feed, and a lab feed as independent branches and reconcile them into one patient-care graph — by exact identifier, blocking key, or fuzzy name match. - **Master-data / entity dedup (CRM, FHIR, catalogs).** Use declared `unique` constraints as definitional identity and similarity scoring for the rest. - **Human-in-the-loop review queues.** `planMerge()` returns the exact proposed write set, conflicts, and entity-resolution evidence without changing the target. Persist that JSON artifact, review it in another process, and apply the reviewed bytes later with `applyMergePlan()`. - **Incremental ingestion against a live graph.** `mergeIncremental()` lets new batches land on a target that has *advanced* since the branch was taken, re-discovering already-committed entities instead of duplicating them. - **Semantic deduplication.** Plug in an embedder for `vector` or `hybrid` similarity to collapse near-duplicates that exact and trigram matching miss. The throughline: **isolation while writing, determinism while merging, and a report you can act on.** ## How it works The mental model is a three-act lifecycle: 1. **`branch()`** stamps the base store's `base@V` and materializes an isolated, independently-mutable working copy. With `revisionTracking: true` (or `history: true`), `base@V` uses the store's durable revision anchor: a per-graph random origin plus a monotonic clock. Validation therefore does not fingerprint every live row or mistake a coincident revision in a separately created store for the branch's base. Existing stores retain the schema-and-content-fingerprint fallback. Writers edit the working copy with the normal store API; the base is never touched. 2. Writers do whatever they want — create nodes/edges, modify inherited rows, delete inherited rows. 3. **Plan, then apply.** `planMerge()` diffs every branch against the base and runs a fixed planning pipeline. `applyMergePlan()` validates the serialized artifact and its digest, checks its revision fence inside the write transaction, then mechanically applies the already-resolved writes: ```text stage (diff every branch) → generate candidates (exact unique · blocking key · similarity) → cluster (group nodes that are the same entity) → canonicalize (pick a survivor, union properties, resolve conflicts) → repoint + dedupe edges onto survivors → reconcile delete/modify and types → emit a revision-fenced JSON plan → validate + commit transactionally + build the report ``` `merge()` remains the one-call convenience wrapper over this same lifecycle; it plans and immediately applies. If the target Store carries a reconciled schema version, its commit acquires and validates the normal schema-write fence before row DML. A raw target remains outside that guarantee. PostgreSQL serialization failures are retried automatically around the complete merge commit. The pipeline is **deterministic by construction**: candidate sets are sorted, clusters resolve by stable keys, and every conflict is decided on an explicit `branchOrder` (or lexicographic branch id) — *never* wall-clock arrival. Merging the same branches in any order yields the same committed graph and the same normalized report. That property is what makes a merge safe to retry, cache, and reason about. ## Quick start Create a base store, fork one branch per writer, write to the branch stores, then merge them back into the target. ```typescript import { createStoreWithSchema } from "@nicia-ai/typegraph"; import { asBranchId, branch, isOk, merge, unwrap } from "@nicia-ai/typegraph/graph-merge"; const [base] = await createStoreWithSchema(graph, baseBackend, { // Recommended for graphs that branch repeatedly or stay live while agents work. revisionTracking: true, }); // branch() is backend-agnostic: you supply a factory for each branch's backend. const makeBranchBackend = async () => createFreshBackend(); const sourceA = unwrap(await branch(base, makeBranchBackend, { id: asBranchId("source-a") })); const sourceB = unwrap(await branch(base, makeBranchBackend, { id: asBranchId("source-b") })); await sourceA.store.nodes.Patient.create({ name: "Anna Rivera", birthDate: "1974-03-09", mrn: "MRN-001" }); await sourceB.store.nodes.Patient.create({ name: "Ana Rivera", birthDate: "1974-03-09", mrn: "MRN-001" }); const result = await merge(base, [sourceA, sourceB], { resolve: { Patient: { block: (node) => node.mrn ?? node.birthDate, similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.78, }, }, onPropertyConflict: "flag", branchOrder: [sourceA.id, sourceB.id], }); if (!isOk(result)) throw result.error; console.log(result.data.resolutions); // the two patients collapsed to one console.log(result.data.conflicts); // the "Anna" vs "Ana" spelling disagreement ``` `branch()` returns a `Result`; `unwrap` throws on failure (or branch on `isOk`). The default working-copy strategy clones the base through TypeGraph's streaming interchange, so each branch gets a fresh backend from your factory without building a graph-sized export document in memory. ## Reviewable plan/apply lifecycle Use the two-step API when approval must happen before accepted graph truth changes. The target must have `revisionTracking: true` or `history: true` so the plan can carry a durable, store-specific revision fence. ```typescript import { applyMergePlan, applyMergePlanInTransaction, isOk, planMerge, } from "@nicia-ai/typegraph/graph-merge"; const planned = await planMerge(base, [sourceA, sourceB], { resolve: { Patient: { block: (node) => node.mrn ?? node.birthDate, similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.78, }, }, }); if (!isOk(planned)) throw planned.error; // Persist outside the target graph, or send to a separate review process. const stored = JSON.stringify(planned.data); const reviewed = JSON.parse(stored); // Later, against the same unchanged target: const applied = await applyMergePlan(base, reviewed); if (!isOk(applied)) throw applied.error; console.log(applied.data.merged); // actual committed effects ``` Planning does not mutate the target. The public plan contains only JSON-safe, deterministically ordered data: its target/schema/revision fence, resolved write set, review information, match evidence, and a stable content digest. It never contains a Store, backend, `Map`, `Set`, callback, or embedder. Applying it does not re-run blocking, candidate generation, similarity scoring, embeddings, canonical selection, or conflict callbacks. The final report copies the reviewed entity-resolution evidence unchanged. The envelope is deliberately explicit: `formatVersion` selects the wire schema; `digest` identifies its canonical content; `mode`, `target`, and `anchors` state what was observed; `proposed` summarizes the review; `writes` is the complete mechanical write set; and `review` holds the conflicts, resolutions, evidence, diagnostics, warnings, and other report material known before apply. ### Apply a plan with application writes Use `applyMergePlanInTransaction(target, tx, artifact)` when the merge, a graph receipt or anchor, and application SQL must share one caller-owned commit. Build `tx` by passing the native transaction to the **same target Store's** `withRecordedTransaction()` callback. Apply the plan before any other write to the target graph in that transaction; after it returns, the callback may make more graph writes and the caller may run more SQL on the native handle. ```typescript await db.transaction(async (nativeTx) => { const { result: report, receipt } = await target.withRecordedTransaction( nativeTx, async (tx) => { const applied = await applyMergePlanInTransaction(target, tx, reviewed); await tx.nodes.MergeReceipt.create({ planDigest: reviewed.digest.value, mergedNodes: applied.merged.nodes, }); return applied; }, ); await nativeTx.insert(mergeRuns).values({ planDigest: reviewed.digest.value, recordedAt: receipt.recorded, mergedNodes: report.merged.nodes, }); }); // await this outer commit before reporting success ``` The adopted applier returns `Promise` and throws a typed `MergeError` on refusal or failure. It does not open, commit, roll back, or retry a transaction. Let the exception reject the outer callback so all merge and application writes roll back together. Never catch it inside the transaction and then commit. When the driver reports a retryable transaction failure, retry the entire outer transaction, including the application writes; do not add an inner retry or nested transaction around the merge. PostgreSQL requires the transaction's observed isolation to be `READ COMMITTED`. SQLite acquires its serialized writer slot before checking the plan fence. The plan must explicitly have `persistProvenance: false`: atomic sidecar provenance persistence is refused on this path. `includeInReport` remains supported, so the returned report can still contain the in-memory provenance index. On a history store, `receipt.recorded` is allocated after the `withRecordedTransaction()` capture callback returns. The caller can persist that anchor with application SQL on `nativeTx` before the outer commit, as the example above does. Await the outer commit before treating the report or receipt as durable. The plan's `proposed` summary describes **proposed changes**. It deliberately does not call them “merged”: `MergeReport.merged` is reserved for the actual effects returned after a successful transaction. Coalescing and idempotent identity operations can make actual counts differ from the proposal. ### Merge after schema evolution in one caller transaction Prepare the evolution first, then call `planMergeForEvolution(target, evolutionPlan, branches, options?)` outside the write transaction. This route resolves writes against the graph produced by the evolution plan while checking the current target's durable data and revision fence. The serialized merge plan names the resulting schema version/hash. If the target schema or revision changes during planning, the planner refuses the artifact; replan outside the transaction. Branches forked from the original baseline can merge existing kinds. To include a newly added kind, call `branchForEvolution(target, evolutionPlan, makeBackend)` before the caller transaction (on PostgreSQL, pass the working-copy manager's `makeBackend`; see [PostgreSQL table-backed working copies](#postgresql-table-backed-working-copies)), then add data on that isolated branch. The planner accepts branches from either one matching baseline; a mixed set of old-schema and resulting-schema forks is refused. Pass `{ revisionJournal: false }` as the fourth `branchForEvolution()` argument when its working copy does not need journal-backed changed-key lineage. The branch remains revision-tracked, and merge planning uses the portable diff when no other lineage source is available. ```typescript const evolutionPlan = await target.planEvolution(extension); const futureBranch = unwrap( await branchForEvolution(target, evolutionPlan, makeIsolatedBackend), ); try { await futureBranch.store.getNodeCollectionOrThrow("Tag").create({ label: "New" }); const mergePlan = unwrap( await planMergeForEvolution(target, evolutionPlan, [futureBranch]), ); await db.transaction(async (nativeTx) => { const { result: report, receipt } = await target.withEvolvedTransaction( nativeTx, evolutionPlan, (tx) => applyMergePlanInTransaction(target, tx, mergePlan), ); await nativeTx.insert(mergeRuns).values({ mergedNodes: report.merged.nodes, schemaVersion: receipt.schema.version, }); }); } finally { await futureBranch.close(); } ``` Apply the merge before other graph writes in the evolved callback. The applier uses the evolved graph and checks the plan's resulting schema and revision fences on the same caller session. Passing a merge plan for the old schema refuses before merge mutation. Evolution's schema CAS is not treated as a prior callback entity write. Roll back the entire native transaction on any refusal; the schema change, merge, recorded capture, and application SQL then roll back together. The report and receipt are provisional until the outer commit succeeds. An adapter configured with `schemaProvisioning: "transactional"` can provision required identity or vector storage on the same native session before the merge callback. The default DML-only policy refuses such requirements before the schema fence or merge mutation. Bootstrap base storage before adopting either route; run generic eager index maintenance separately after the outer commit. ### Candidate write sets for a planned schema `planCandidateWriteSetForEvolution()` is the branch-free counterpart for a bounded candidate batch. First use `captureCandidateWriteSetTargetForEvolution(target, evolutionPlan)` when authoring the JSON document; it records the evolution plan's resulting schema identity rather than the currently active one. The planner stages the candidate against that resulting graph and returns the same resulting-schema merge artifact accepted by `withEvolvedTransaction()`. Candidate resolution still includes the committed target as an accepted source. Existing unique matches and property conflicts are therefore visible in the reviewed plan before the evolution transaction begins, rather than surfacing as late write-time failures. ```typescript const evolutionPlan = await target.planEvolution(extension); const writeSet = { formatVersion: 1 as const, sourceId: "import-batch-42", target: captureCandidateWriteSetTargetForEvolution(target, evolutionPlan), nodes: [{ kind: "Tag", id: "import-batch-42:tag-1", properties: { label: "Research" }, validFrom: "2026-01-01T00:00:00.000Z", }], edges: [], }; const mergePlan = unwrap(await planCandidateWriteSetForEvolution({ target, evolutionPlan, makeBackend: makeIsolatedBackend, writeSet, })); await db.transaction(async (nativeTx) => target.withEvolvedTransaction(nativeTx, evolutionPlan, (tx) => applyMergePlanInTransaction(target, tx, mergePlan), ), ); ``` The schema change and accepted candidate writes share the caller's one transaction and recorded revision. If another writer changes the target while planning, `MergePlanningStaleError` is an expected concurrency result: discard the candidate plan, recapture the target for a new evolution plan, and replan. For a frozen ancestor and a live destination, use the named incremental planner: ```typescript const planned = await planMergeIncremental({ forkPoint, target, branches, options, }); if (!isOk(planned)) throw planned.error; const applied = await applyMergePlan(target, planned.data); ``` When `target` records history, a durable branch can use its sealed recorded fork point without keeping a second frozen Store: ```typescript const forkPoint = created.branch.recordedForkPoint; if (forkPoint === undefined) throw new Error("History was not captured at fork"); const planned = await planMergeIncremental({ forkPoint, target, branches: [created.branch], options: { onBasePropertyConflict: "flag" }, }); ``` `recordedForkPoint` is available when the source captured history at fork time; it contains both the recorded instant and the branch's `base@V` token. The planner reads ancestor rows from the target's recorded relations, validates the origin, schema, and revision anchor, and enumerates only changed keys when lineage can prove a complete delta. A missing or incompatible anchor is refused before planning. The direct `mergeIncremental()` wrapper accepts the same fork point. Keep the durable descriptor with the branch: reopening restores the recorded fork point from the sealed origin. The same target revision must still be current when the reviewed plan is applied. If it moved during planning, planning returns `MergePlanningStaleError` and no artifact. This is an expected retry-and-replan outcome under concurrency: recapture the target, create a new plan, and review its new digest before retrying. If it moved afterwards, `applyMergePlan()` returns `StaleMergePlanError` before plan writes. Re-plan, review the new digest and proposal, then apply the new artifact; never edit an old plan or retry it as though it still represented the target. A successful plan is single-use: a second or concurrent application is stale. Persisting a plan or approval in the target graph also advances this revision. For exact-plan approval, use external storage or a separate graph ID; writes to that graph do not advance this target's revision. This does not provide atomic writes across graphs, and any intervening target write still requires a fresh plan. For candidate batches whose review records belong in the target itself, use the durable review protocol below. `merge()` and `mergeIncremental()` remain convenient compatibility wrappers. They invoke the same planner and applier contiguously and return the same `MergeReport` shape as before, now with match evidence on each resolution. Use the wrappers when no external approval boundary is needed. :::caution[Sensitive plans and trust] A plan contains the complete resolved writes and may therefore contain personal, regulated, or otherwise sensitive application data. Protect it like the source graph: encrypt it where appropriate, restrict access, and avoid logging it. The digest identifies the exact canonical artifact and detects accidental or unrecorded changes. It is **not** a signature, proof of origin, authentication, or authorization. Authenticate untrusted storage and authorize the caller before passing a plan to `applyMergePlan()`. ::: ### Durable candidate review in the target graph `planCandidateWriteSetReview()` separates immutable review evidence from a revision-bound execution plan. Its `MergeReviewArtifact` retains the original candidate write set, reviewed plan, normalized merge options, explicit policy identity/context, and target baseline. You can persist this artifact and later approval records in the target before calling `revalidateCandidateWriteSetReview()` to compute a fresh execution plan. Both review versions support candidate write sets only. They do not rebase arbitrary artifacts from `planMerge()` or `planMergeIncremental()`. Candidate planning on revision-tracked graphs reads existing candidate ids and edge endpoints by key, then seeds only those rows in the transient working copy. On identity-enabled graphs, it also follows live same-id peers and current identity assertions from those references to a fixed point. The planner reads peers of a candidate edge with `one` cardinality by source, peers of a `unique` edge by its endpoint pair, and the active peer of a `oneActive` edge by source. The active-only read checks an open `validTo` even when `validFrom` is in the future, and does not return ended history. These reads let the transient copy enforce the same cardinality rule as a complete clone. On graphs with ontology relations, it also reads live nodes sharing each candidate reference's id across kinds, so disjointness sees the same peers as a complete clone. Ontology subtype relationships remain graph metadata. The candidate diff and its target baseline are bounded to that dependency set and any committed rows recalled by configured unique or index sources. Planning still fences the target revision before and after these reads. With edge match-identity constraints, a backend offering `findEdgesByMatchIdentity` seeds the exact durable owners named by the candidate. A missing keyed read, an owner excluded from the clone projection, or a target without revision tracking uses the complete clone path. A custom backend lacking the optional `findActiveEdgesBySourceV1` read also uses that path for `oneActive` graphs. On the complete clone path, when the copy and target really share one serialized connection, clone export is materialized before import, but its snapshot still holds the connection's exclusive stream lease while it is collected. Concurrent review calls on that resource can therefore return a merge error caused by a `ConfigurationError` with `details.code: INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` or `INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS`. Await the whole review call before starting another on the same serialized resource. A `pg.Pool` with more than one connection is not one serialized resource; do not declare its pool object as `{ mode: "shared" }` just because the working copies use that pool. See [Serialized connections](/backend-setup#serialized-connections). The following continues the [candidate write set example](#constraint-aware-ingestion-branches). `Artifact`, `Decision`, and `evidence` are application-defined node/edge kinds; `proposal` is an existing node. The target enables `history` or `revisionTracking`. ```typescript import { applyMergePlan, planCandidateWriteSetReview, revalidateCandidateWriteSetReview, unwrap, } from "@nicia-ai/typegraph/graph-merge"; const policy = { id: "acceptance-policy-v1", context: { requiredApprovals: 1, resolverVersion: "2026-09-01" }, }; const review = unwrap( await planCandidateWriteSetReview({ target: store, makeBackend, writeSet, policy, }), ); const artifact = await store.nodes.Artifact.create( { content: JSON.stringify(review) }, { id: review.digest.value }, ); await store.edges.evidence.create(proposal, artifact, { note: "review" }); // After an authenticated reviewer approves under the application's policy: const decision = await store.nodes.Decision.create({ approved: true, reviewDigest: review.digest.value, }); await store.edges.evidence.create(decision, artifact, { note: "approval" }); // Later: authenticate the stored artifact and decision, then check current // authorization, approval validity, and policy before reusing that approval. const persisted = await store.nodes.Artifact.getById(artifact.id); if (persisted === undefined) throw new Error("Missing review artifact"); const checked = unwrap( await revalidateCandidateWriteSetReview({ target: store, makeBackend, review: JSON.parse(persisted.content), policy, }), ); if (checked.status !== "compatible") { console.log(checked.differences); throw new Error("Create a new review and obtain a new approval"); } if (checked.reviewDigest.value !== decision.reviewDigest) { throw new Error("Approval does not identify the validated review"); } // Keep this fresh execution plan ephemeral: another target write makes it stale. const applied = unwrap(await applyMergePlan(store, checked.plan)); // A separate commit AFTER successful apply; see the recovery boundary below. await store.nodes.Artifact.create({ content: JSON.stringify({ reviewDigest: checked.reviewDigest, approvalId: decision.id, executionPlanDigest: checked.plan.digest, executionTarget: checked.plan.target, report: applied, }), }); ``` All approval records and links must be committed before final revalidation. Do not persist each replacement execution plan in the target: that repeats the staleness cycle. Retain the original review, and use the returned `reviewDigest` plus the fresh plan's `digest` and `target` fence to relate approval to execution. Both review APIs return `Result<..., MergeError>`. Revalidation accepts the persisted artifact as `unknown` and replans its retained candidate input once target, policy, and baseline checks pass. Supply current merge `options` and `policy` again; callbacks are never restored from serialized data. | Revalidation status | Meaning and next step | | --- | --- | | `compatible` | Includes a fresh `plan` and the original `reviewDigest`. Application policy may reuse approval; authorize the action and apply promptly. Compatibility itself grants no permission. | | `changed` | `differences` identify changed policy/options, baseline entities/identity, or plan fields. Obtain a new review and approval. A `plan` is included only when fresh planning completed. | | `incompatible` | The graph ID, schema identity, or revision origin differs. Approval cannot be reused for this target; resolve the mismatch and create a new review. | Malformed/unsupported artifacts, mismatched digests, and missing required evidence return `MergeReviewError` (`GRAPH_MERGE_REVIEW`). Existing typed planning and constraint errors remain errors rather than compatibility statuses. A target change during evidence capture/planning returns `MergePlanningStaleError`. The V1 baseline is deliberately conservative: - Every original node and edge row, including tombstones and validity metadata, must remain unchanged. Editing an old audit record requires a new review even when the candidate's resolved writes would be identical. - Expected absences for candidate/write/guard references must remain absent. Same-ID nodes of other kinds are also guarded, because they can change implicit identity membership. Complete archival identity evidence must remain unchanged. - Newly added rows can coexist with approval only when fresh planning produces identical resolved writes, guards, conflicts, evidence, provenance, and other plan content. Candidate-derived anchors and the execution digest/fence are regenerated. There is no exemption for an “audit” kind. For an eligible revision-tracked graph, pass `reviewScope: "candidate"` to `planCandidateWriteSetReview()` to emit V2 candidate-scoped evidence. V2 fingerprints the candidate's node and edge ids, edge endpoints, resolved writes, and plan guards, including expected absences across kinds. On Operational Identity graphs it also records the reachable identity assertion and same-id peer closure, plus assertion-ID collision evidence. Revalidation expands that retained identity scope, rereads the referenced rows, and replans the candidate under a new target fence. An unrelated original row may change without invalidating V2 when it cannot affect the fresh resolved plan; V1 would report that row change. Applications whose approval policy needs the V1 whole-graph rule should omit `reviewScope`. The review artifact records its version and scope, so revalidation applies the rule originally reviewed. Candidate-scoped review refuses graphs outside those eligibility rules. On a `oneActive` graph, a custom backend must expose `findActiveEdgesBySourceV1` for candidate-scoped review; the complete-clone candidate planner and V1 review remain available when it does not. On an Operational Identity graph, a custom Store runtime must also expose endpoint-scoped and assertion-ID-scoped identity reads. Without both reads, ordinary candidate planning uses the complete working-copy clone and V1 review remains available; an explicit V2 candidate-scoped review request is refused. Applicable store constraints still run during atomic application. Compatibility does not promise that apply will succeed: new rows may introduce constraint conflicts, and any write between revalidation and apply causes `StaleMergePlanError`. A failed application commits no partial candidate node, edge, or identity writes. Revalidate again after a stale refusal; require reapproval if the result changes. `policy.id` identifies your policy implementation; `policy.context` explicitly records every opaque dependency that can change its decision. Include callback and resolver versions, model/prompt versions, external configuration or data versions, and any application state used to authorize approval reuse. Use an empty context only when no such dependencies exist. TypeGraph captures callback presence and serializable options, but cannot discover callback code, closure state, external reads, or hidden application policy dependencies. The producer must supply complete evidence, and the application must authenticate the entire stored review and its approval. Content addressing and SHA-256 detect content changes; anyone able to replace evidence can recompute a digest. A valid digest is neither proof that the baseline was complete nor authorization to reuse approval. Enforce artifact immutability and access control in your storage or application. The review contains candidate data and an entire reviewed plan, so protect it with the same care as graph data. V1 review capture and revalidation read and fingerprint the complete target graph and archival identity ledger. The artifact stores one fingerprint per original row plus expected absences. Budget graph-sized reads and artifact storage for V1. V2 candidate-scoped review uses bounded point and identity closure reads for its baseline on eligible graphs. The execution receipt above is a separate commit. If its write fails or the process stops after apply, the merge may already be committed without a receipt. Retain the original review and approval, and reconcile committed history and application operation identity before repairing the receipt. Do not treat a missing receipt as permission to replay the candidate; applying its old execution plan is stale, and replanning is not a duplicate-execution check. To commit the receipt atomically with the merge, create it in an `afterApply` callback as described in [Composing application checks and writes](#composing-application-checks-and-writes). Review revalidation alone does not add that guarantee. See the runnable [durable merge review example](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/27-durable-merge-review.ts) for the complete schema and lifecycle. ## Composing application checks and writes Pass execution callbacks to `applyMergePlan()` when a reviewed candidate and related application records must commit together: ```typescript const result = await applyMergePlan(target, reviewedPlan, { beforeApply: async (reads) => { const resource = await reads.nodes.Resource.getById(resourceId); if (resource?.owner !== "unclaimed") { throw new Error("Resource is already claimed"); } }, afterApply: async (tx, applied) => { await tx.nodes.Resource.update(resourceId, { owner: "accepted" }); const decision = await tx.nodes.Decision.create({ status: "accepted", changedNodes: applied.merged.nodes, }); await tx.edges.decides.create(decision.id, resourceId, {}); }, }); if (!isOk(result)) throw result.error; // The plan and application writes have now committed together. ``` `Resource`, `Decision`, and `decides` stand for types registered in your graph. Import `MergePlanApplyOptions`, `MergePlanReadContext`, and `MergePlanApplied` from `@nicia-ai/typegraph/graph-merge` to type reusable helpers. Callbacks are execution options: they are not stored in the artifact or covered by its digest. The transaction acquires the schema fence before the graph write lock, validates schema, revision origin, and revision, and then calls `beforeApply`. This context exposes node and edge collection reads and, for identity-enabled graphs, identity reads. Write methods, native SQL, and a root Store are absent. Application writes before plan application are deliberately unsupported: an uncommitted write can change the reviewed state without changing its durable revision yet. After the precheck, TypeGraph performs the existing plan preflight, identity checks, and writes. `afterApply` receives a `TransactionContext` whose reads see those writes. Its `applied.merged` contains only the plan's provisional counts; callback writes do not contribute to the final report's merge counts. Use the supplied contexts for every graph operation. Calling the original Store inside a callback does not enlist it in this transaction. Do not retain a context for later work, and await every operation before returning. Callbacks must resolve without a value. Throw or reject to abort; returning a value, including an `Err`, is refused with `InvalidMergeOptionsError`. A callback rejection, stale plan, merge failure, capture-flush failure, or commit failure rolls back the combined graph operation. Errors are converted to the outer `Result` after rollback; ordinary application errors are retained in the cause chain. Existing typed merge errors and constraint translation remain intact. Only the successful outer result confirms commit. Transaction conflicts (PostgreSQL serialization failures or deadlocks) retry the whole transaction up to **three attempts**, including both callbacks. Every attempt checks the fence again; an intervening committed write makes the plan stale rather than silently rebasing it. Keep callbacks safe to repeat. Do not send messages, call external services with side effects, or publish an outcome inside a callback. Perform those effects after successful completion, or write an application outbox record through `tx` for later delivery. Returned contexts and provisional outcomes are not durable notifications. Protection covers the target graph's transactional state and participating TypeGraph writers using its graph fence. It does not make an application policy a declarative constraint: every writer changing that policy's state must enforce it, for example through its own conditional operation. It does not cover other graphs, arbitrary SQL, or external systems. SQLite uses its writer transaction; Composed PostgreSQL applications use read-committed isolation and the graph write lock with or without history. The lock statement records the effective session isolation; incompatible or unknown isolation is refused before callbacks. Standalone revision-tracking-only applications retain serializable isolation. Unsupported transaction capabilities are refused before callbacks. Existing session-bound fence and recorded-capture isolation checks still apply. With history enabled, plan and application writes share the transaction's recorded capture and flush, producing one per-graph recorded revision. Without history, revision tracking likewise advances for the combined transaction. Failure leaves no live changes or recorded revision from the failed attempt. Existing valid-time bounds, including open bounds, retain their semantics. Optional persisted merge provenance remains separate from recorded history: provenance records are persisted only after successful graph commit, and a persistence failure remains a report warning. Callbacks do not receive a post-commit provenance result. Previously committed review records in the target still invalidate a plan's revision fence; this API does not relax plan staleness. For a durable candidate review, revalidate the stored review first, then pass the compatible result's fresh `plan` and these callbacks to `applyMergePlan()`. ## Scaling branches and interchange `revisionTracking: true` is the recommended mode for long-lived, repeatedly branched graphs. It advances one durable revision anchor inside each successful Store write transaction. The anchor combines a per-graph random origin with the monotonic commit clock, so a branch can only match the store that created it — not an independent database whose clock happens to share the same timestamp. A branch and its merge precondition then read that constant-size anchor instead of hashing every live node and edge. Stores created with `history: true` already have the same guarantee through their recorded-time commit clock. On PostgreSQL, the guarantee serializes writes to the same graph with a transaction-scoped advisory lock. That is the correct trade-off for a live graph whose branch merges must fail closed, but it can reduce throughput and increase write latency for a high-concurrency, single-graph workload. Partition that workload across graphs or leave revision tracking off when the content-fingerprint fallback is acceptable. Turning revision tracking off does **not** turn off all serialization. *Constrained* writes now take the same per-graph mutual exclusion regardless of `revisionTracking` or `history`, because their check-then-write is only sound if nothing else writes the graph in between: edge cardinality (`one`, `unique`, `oneActive`, and the `getOrCreateByEndpoints` create and resurrect legs), node-kind disjointness on create, and a `kindWithSubClasses` uniqueness constraint that actually expands to more than one kind — a scope covering a single kind probes exactly the row the uniques table's primary key then reserves, so that key is already its fence. Everything else — an unconstrained create, a delete, a cardinality-`many` edge — pays nothing, so the cost is proportional to the constraints you actually declared. On PostgreSQL that exclusion is the same transaction-scoped advisory lock; on SQLite it is the `BEGIN IMMEDIATE` writer slot the backend already takes. A backend running without transactions (D1, `neon-http`, or `transactionMode: "none"`) has neither and cannot be fenced. This unlocks: - Many concurrent agent, importer, or review branches without base-version validation growing with the graph. - Large graph copies, backup/export, and transfer pipelines that keep only one interchange batch resident at a time via `exportGraphStream()` and `importGraphStream()`. - A safe fast path for a live base: a branch is rejected if any tracked base write lands before its merge commits, rather than silently merging a stale plan. Streaming removes the graph-sized heap spike, but a physical working copy still copies `O(graph)` rows and snapshot merge staging still compares branch state to the base. Bundled backends page those comparisons across declared kinds, so unused kinds do not each cost a database statement; custom backends without the cross-kind read retain per-kind keyset pagination. Disposable candidate clones also skip statistics refresh. Copy-on-write logical branches and delta-only staging remain the next larger architectural step. Revision tracking covers writes through the Store API. Direct backend writes and raw graph-table writes through `tx.sql` bypass the anchor, so applications using either escape hatch must avoid them for a branchable graph or retain the default content-fingerprint validation. On transactional backends, streaming export holds one read-only repeatable-read transaction across nodes, edges, and identity assertions, so every chunk belongs to one committed snapshot. A snapshot stream cannot be piped directly into a target that writes through the same serialized connection: the same SQLite backend, distinct wrappers sharing one better-sqlite3 handle or one local (`file:`/`:memory:`) libSQL client, a bare `pg`/neon `Client` (a checked-out `PoolClient` included), a `pg` `Pool` capped at one connection (`{ max: 1 }`, and equally the uncoerced string forms `{ max: "1" }` and the legacy `{ poolSize: "1" }` that `max: process.env.PG_MAX` produces), a postgres-js client capped at one connection (`{ max: 1 }`, `?max=1` in the URL, or `PGMAX=1`), distinct PGlite backend wrappers sharing one in-process connection, or Cloudflare Durable Object storage, whose transaction frame is ambient on the storage object — materialize it first or import it into an independent backend. Pooled connections, HTTP drivers, remote libSQL, and separate handles on one database are deliberately not treated as serialized: each statement gets an independent connection there, so refusing would refuse work that succeeds. The exclusion is one **exclusive** lease per serialized connection, not a one-time check and not a cross-kind-only rule: at most one long-lived interchange stream of any kind holds a given connection, so all four pairings are refused — import behind export snapshot (even through a user-wrapped stream that no longer identifies its source backend), export snapshot behind streaming import, export behind export, and import behind import. Whichever long-lived stream starts second gets a typed `ConfigurationError` instead of both hanging; its `details.code` names the condition holding the connection and `details.requested` / `details.heldBy` name the pairing that was refused (see [Interchange serialized-connection guard codes](/errors#interchange-serialized-connection-guard-codes)). Every long-lived import claims that lease, not only the chunk-streaming one: `importGraph` holds it for the whole call and `trustedImportGraph` / `trustedImportGraphStream` for the whole trusted session, so those APIs can throw this `ConfigurationError` too — new in 0.46 for trusted import, which previously threw only `TrustedImportError`. TypeGraph's branch cloner detects the shared-client case and materializes its snapshot before importing it. Non-transactional backends can export identity-disabled graphs without this snapshot guarantee. Identity-enabled stores already require a transactional backend at construction, so every identity export has the snapshot guarantee. ## Entity resolution Resolution is configured **per node kind** in `resolve`. A kind that is omitted merges *by id only*: its new nodes and edges are copied through, but no fuzzy matching runs. Each configured kind composes up to three candidate sources, all feeding one shared scorer: | Source | What it matches | Configured by | | ------------ | ---------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | | Exact unique | Two staged nodes sharing all of a declared `unique` constraint's values — a *definitional* match that bypasses scoring | the graph's `unique` constraints | | Blocking key | Cheap pre-grouping so similarity only compares plausibly-related nodes | `block` (staged) / `blockIndex` (vs. committed base) | | Similarity | Fuzzy scoring of candidate pairs against a `threshold` | `similarity` + `threshold` | ```typescript resolve: { Patient: { block: (node) => node.mrn ?? node.birthDate, // cheap candidate grouping similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.78, // pairs scoring >= 0.78 merge }, } ``` ### Blocking: `block` vs `blockIndex` Blocking bounds the otherwise-`O(n²)` pairwise comparison by only comparing nodes that share a cheap key. - **`block(node) => string | undefined`** is an arbitrary function over staged nodes — a normalized email, a tenant id, a birth date, a `soundex(name)`. Returning `undefined` puts the node in the shared *unblocked* bucket. - **`blockIndex`** names a declared `defineNodeIndex` and is the **new-vs-base** block key: it lets the merge query *already-committed* nodes that share a staged node's index key and propose them as candidates. It powers incremental ingestion (see [Snapshot vs incremental](#snapshot-vs-incremental)) and is ignored on the snapshot `merge()` path. ```typescript import { defineNodeIndex } from "@nicia-ai/typegraph"; const patientCohort = defineNodeIndex(Patient, { name: "patient_cohort_idx", fields: ["cohort"] }); const graph = defineGraph({ /* ... */ indexes: [patientCohort] }); // In resolve, recall committed patients in the same cohort: resolve: { Patient: { blockIndex: "patient_cohort_idx", similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.85 } } ``` ### Keyless windows A node with no block key and no unique signature lands in the *unblocked* bucket, which is otherwise compared all-vs-all. For large unblocked sets, set `keyless` to switch to bounded single-pass **sorted-neighbourhood**: nodes are sorted by their similarity text and each is compared only to its next `window` neighbours — `O(n·window)` instead of `O(n²)`, still fully deterministic. ```typescript resolve: { Article: { similarity: { kind: "fulltext", fields: ["title"] }, threshold: 0.8, keyless: { window: 20 }, // compare each unblocked article to its 20 nearest neighbours }, } ``` ### Similarity strategies Four strategies cover the spectrum from zero-dependency to embedding-powered: | Strategy | Needs embedder? | Use case | | ---------- | --------------- | ---------------------------------------------------------------------------------------------------------------- | | `fulltext` | No | Portable in-memory Sørensen–Dice trigram score over one or more fields (e.g. `name`). The cross-backend default. | | `custom` | No | Your own deterministic `score(a, b) => number` — domain rules, weighted field blends, edit distance. | | `vector` | Yes | Cosine similarity over one field's embedding. Catches semantic near-duplicates. | | `hybrid` | Yes | Blend `vector` and `fulltext` by `weights` (default 0.5 / 0.5). | The `fulltext` scorer runs **in memory** over the staged candidate text — it deliberately does not consult database fulltext indexes, because branch candidates are staged working-copy rows, not indexed search results. That keeps scoring deterministic and identical across SQLite and Postgres. For `vector` / `hybrid`, supply an `embedder` (batched, async, deterministic — the same text must always map to the same vector): ```typescript const result = await merge(base, branches, { embedder: async (texts) => texts.map((text) => embedModel(text)), // text[] -> Float32Array[] resolve: { Article: { similarity: { kind: "hybrid", fields: ["title", "summary"], weights: { vector: 0.7, fulltext: 0.3 } }, threshold: 0.84, }, }, }); ``` A `vector`/`hybrid` strategy with no embedder configured fails with a typed `SimilarityUnavailableError`, never a silent no-op. ## Conflicts When merged contributors disagree on a property value, Graph Merge **resolves by an explicit, deterministic policy and records what it did** — it never lets arrival order decide. ### Property conflicts | Policy | Behavior | | -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `flag` (default) | Commit the deterministic survivor value (or the committed base value, for base-vs-branch) and record a `PropertyConflict` for review. The graph still gets a value; the disagreement is surfaced rather than resolved toward another branch. | | `lastWriteWins` | Pick the value from the highest-priority branch (earliest in `branchOrder`) — *logical* order, never wall-clock. | | `provenanceWeighted` | Pick the value from the highest-weight branch (see `provenanceWeights`). Ties fall back to branch order. | | function | Delegate: `(conflict) => JsonValue` lets application code decide per conflict. | There are **two** property-conflict knobs, deliberately separate so a fuzzy branch match can never silently overwrite committed data: - `onPropertyConflict` — staged branch vs. staged branch. - `onBasePropertyConflict` — committed base vs. a branch (new-vs-base merges). Defaults to `flag` independently, and does **not** inherit `onPropertyConflict`. `provenanceWeighted` reads per-branch trust weights you supply: ```typescript const result = await merge(base, branches, { onPropertyConflict: "provenanceWeighted", provenanceWeights: new Map([ [authoritativeFeed.id, 1.0], // the system of record wins ties of value [bestEffortAgent.id, 0.2], ]), }); ``` ### Delete / modify conflicts An inherited node or edge that one branch **deletes** while another **modifies** is neither a pure delete nor a pure modify. `onDeleteModifyConflict` governs it for both nodes and edges: | Policy | Behavior | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `flag` (default) | The modification survives **and** an unresolved `DeleteModifyConflict` is recorded — a merge must never silently destroy the only branch still carrying data. | | `deleteWins` | Honor the delete; discard the modification; record the conflict. | | `modifyWins` | Resurrect the row; keep the modification; record the conflict. | Independent edits to the *same* inherited row by different branches are **three-way merged against the base**: a field only one branch changed takes that change with no conflict; only fields multiple branches changed to differing values become conflicts. This holds for node *and* edge properties, so disjoint edits compose instead of clobbering each other. ## Edges follow their entities When nodes collapse, their edges must too. After clustering, Graph Merge: 1. **Repoints** every edge endpoint onto its cluster's canonical survivor. 2. **Drops** any edge whose endpoint was finally deleted (recorded in `dropped`). 3. **Dedupes** edges that repointing brought together, as a pure set keyed by `(from, type, to, props)` — so `x → a` and `x → b` both landing on `x → c*` collapse to one edge. 4. **Reconciles** edges that collapse that way but disagree on properties, via the same conflict policy as nodes — over the properties each side actually *changed*, so an inherited row's untouched value never competes with (or outvotes) a value some branch authored. Steps 3 and 4 are scoped to collisions **repointing caused**: edges are grouped by the endpoint pair they named *before* repointing, and one row per pair collapses. A TypeGraph store is a multigraph — nothing enforces uniqueness on `(from, kind, to)`, `create()` makes a parallel edge, and `getOrCreateByEndpoints()` is the opt-in set-semantics accessor — so a branch that created a parallel edge merges as a parallel edge, and a window claim lands on the row its author touched. A repointed edge landing on endpoints that already have several parallel rows merges into one of them; the rest keep their own properties and windows. What makes two staged edges "the same row" is their **edge id**, not equal properties: one inherited edge staged by several branches folds into a single write, while a branch-created edge is a new row even when its properties happen to match an existing one's. When such a collapse mixes an **inherited** edge with a branch-created one, the inherited row is the one kept: a collapse rewrites the row it keeps and does not end the rows folded into it, so writing onto the row the target already holds is what keeps a committed edge from being left beside the row that replaced it. This mirrors the node rule below, and it is also what the surviving edge id in `PropertyConflict`, window resolutions, and provenance names. A collapse of branch-created edges alone keeps the lexicographically-minimal edge id. Inherited edges that a branch **deleted** are removed from the target, and inherited edges **modified** by multiple branches go through the same base-aware three-way merge as nodes — so an edge's `since` edited by one branch and `note` edited by another keep *both* edits. The collapse in step 4 is base-aware for the same reason: a staged copy of an inherited row carries that row's whole property bag, and only the values it *changed* count as claims. The clearest case is a row staged solely to carry an end-of-validity — it authored no property, so it contributes no claim and raises no conflict, whatever its branch's rank. ## Ontology type reconciliation With `reconcileTypes: "ontology"`, two staged nodes that share an id but carry subtype-compatible kinds (via the graph's `subClassOf` closure) are collapsed to the **most-specific** common type, recorded as a `TypeReconciliation`. A base `Doctor` and a branch `SpecialistDoctor` reconcile to `SpecialistDoctor` instead of being dropped as incompatible. The default `"off"` keeps identity strictly `(kind, id)`. ```typescript const graph = defineGraph({ /* ... */ ontology: [subClassOf(SpecialistDoctor, Doctor)] }); const result = await merge(base, branches, { reconcileTypes: "ontology" }); ``` ## Choosing the survivor By default a cluster's canonical survivor is the member with the lexicographically-minimal id. A committed member always wins instead, so its committed identity and the edges already attached to it stay stable: on new-vs-base merges that is a committed base member, and on incremental merges it is also a node the live target committed after the fork point, such as one an earlier branch's merge added. Override the staged-vs-staged choice with `canonical`: ```typescript const result = await merge(base, branches, { canonical: (cluster) => preferGoldenSource(cluster.members), // pick which id survives }); ``` ## Scaling & safety Two guards keep a merge bounded and predictable on large or pathological inputs: - **`maxComparisonsPerKind`** caps fuzzy comparisons per kind. On overflow, `onComparisonCeiling` decides: `"error"` (default) fails with a typed error, or `"mergeByIdOnly"` skips similarity for that kind (still honoring exact unique matches) and records a warning. Tighten your `block` to shrink buckets rather than raising the ceiling blindly. - **`clusterMaxDiameter`** optionally splits over-broad clusters: if a cluster's single-link diameter exceeds the bound, the weakest edges are dropped deterministically until every sub-cluster fits. This stops a chain of near-matches (`a~b~c~…`) from fusing genuinely-distinct entities. ```typescript const result = await merge(base, branches, { maxComparisonsPerKind: 50_000, onComparisonCeiling: "mergeByIdOnly", clusterMaxDiameter: 2, }); ``` ## The merge report `merge()` returns `Result`. The report is the **application boundary** — show conflicts to an operator, write a review record, persist provenance, or feed a downstream step. ```typescript type MergeReport = { merged: { nodes: number; edges: number; identity: { asserted: number; retracted: number }; // ledger effects }; resolutions: EntityResolution[]; // collapse membership + decisive match evidence conflicts: PropertyConflict[]; // per-property disagreements + how they resolved deleteModifyConflicts: DeleteModifyConflict[]; // node/edge delete-vs-modify cases typeReconciliations: TypeReconciliation[]; // ontology kind collapses // Node drops (deleted endpoints, incompatible members), edge drops, identity // drops (identity:duplicate-assertion, identity:endpoints-collapsed, // identity:retraction-target-mismatch, identity:deletion-overruled), and // lower-bound deltas the commit cannot apply (window-not-applicable) dropped: DroppedItem[]; // Inherited rows whose end-of-validity the merge resolved. Each entry carries // validTo for a set/move or clearValidTo: true for a reopening. validityEnds: ValidityEndResolution[]; baseAmbiguities: BaseAmbiguity[]; // new-vs-base matches that spanned >= 2 committed entities provenance: ProvenanceIndex; // byBranch(id) -> { nodeIds, edgeIds } warnings: string[]; // non-fatal advisories (ceiling skips, provenance-persist failures) candidateDiagnostics?: CandidateDiagnostics; // bounded, opt-in scored comparisons provenancePersisted?: { graphId: string; count: number }; // when persistProvenance ran }; ``` A typical operator loop: auto-apply when `conflicts` and `deleteModifyConflicts` are empty; otherwise enqueue them for review alongside `resolutions` so the reviewer sees what merged and why. ### Why two entities matched Every multi-member `EntityResolution` has `decisiveEdges`: a deterministic minimal connectivity witness. A resolution over N distinct `(kind, id)` identities normally has N−1 edges. Endpoints retain both kind and id, so same-id nodes of different kinds remain distinguishable during ontology reconciliation. A same-id ontology retype remains a `TypeReconciliation`, rather than creating an id-merge resolution. Its optional `decisiveEdges` carries the accepted retype witness without changing the meaning of the existing resolution collection. Each edge records every candidate source that proposed the pair in stable order. Definitional evidence names the trusted rule, such as a unique constraint, and does not pretend the internal forced match was a perfect similarity score. Scored evidence records the strategy descriptor, actual score, and threshold used by the shared scorer: ```typescript type MatchEvidence = | { a: { kind: string; id: string }; b: { kind: string; id: string }; sources: MatchSource[]; decision: "definitional"; } | { a: { kind: string; id: string }; b: { kind: string; id: string }; sources: MatchSource[]; decision: "scored"; strategy: MatchStrategy; score: number; threshold: number; }; ``` Built-in source metadata distinguishes block, unique, base-unique, base-index, keyless, and ontology-retype proposals. Several sources proposing the same pair are all retained after deduplication. Strategy metadata describes `fulltext`, `vector`, `hybrid`, or `custom` configuration, never custom function source. Default evidence excludes the raw compared values and rejected pairs because those may contain PII and can make reports enormous. Candidate diagnostics are explicit and bounded: ```typescript const planned = await planMerge(base, branches, { ...options, candidateDiagnostics: { limit: 1_000 }, }); ``` When enabled, the report and reviewable plan include accepted and rejected scored comparisons in canonical order. A definitional edge removed by the base ambiguity or diameter guard is also retained with its exclusion reason, so the final partition remains explainable. The collection also carries `total`, `limit`, and `truncated`. The limit is deterministic: the same candidate set produces the same retained prefix regardless of branch, source, or backend enumeration order. Diagnostics still omit raw compared values; join their `(kind, id)` references to application data only in an appropriately protected evaluation environment. ## Provenance Provenance answers *which branch contributed each merged node and edge*. A contribution is anything a branch authored into the committed row — the properties it staged, the modification that survived, or the end-of-validity the merge applied. - **Report-only (default, `provenance: true`)** — `report.provenance.byBranch(id)` returns the `{ nodeIds, edgeIds }` that branch contributed. In-memory; it evaporates after the call. - **Durable (`persistProvenance: true`)** — one `{branch, sourceId} → canonical` row per contribution is upserted into a *sidecar* graph on the target's backend (its own namespaced tables; your domain schema is untouched). The sidecar is opened and claimed **before** the merge commits, so a sidecar graph id TypeGraph cannot claim refuses the whole merge and leaves the target unmodified; only the row write itself is post-commit and best-effort, where a transient failure surfaces as a `warnings` entry rather than a failed merge. Re-running the same merge upserts (deterministic ids), never duplicates. `openProvenanceStore` only ever opens a sidecar graph id it can prove it owns, and ownership is **marker-first**: a durable `ProvenanceOwner` marker row is the sidecar's first write of any kind, committed inside the schema fence *before* the sidecar schema is registered. A never-seen id is free to claim only when it holds no row in **any** per-graph table — nodes and edges, but equally recorded-time history, the revision clock and origins, identity assertions and their derived closure and separation, fulltext, and unique keys — because a plain `createStore` writes rows without registering a schema, so an unregistered id is not by itself evidence of a free namespace. Ownership is then the marker alone, checked independently of the schema hash, because an application is free to define the same `Provenance` shape at an unrelated id. Because the marker comes first, the resumable interrupted state is **marker without schema** (or a marker beside a pre-marker legacy schema): that resumes by registering or migrating the schema. The opposite state — the exact current sidecar schema with no marker — is one TypeGraph cannot produce, and is refused unconditionally whatever the graph contains, empty and provenance-shaped included, since contents an application could have written are not evidence of authorship. **What a claim costs, on PostgreSQL.** One writer class takes neither the per-graph fence nor the graph's active schema row: a schema-less raw `createStore` writer, or a direct `backend.insertNode` / `insertEdge` call. At READ COMMITTED its insert could commit between the claim's re-inspection and the claim's own commit, leaving the marker on an id an application had just made its own. To close that, the claim issues `LOCK TABLE , IN SHARE ROW EXCLUSIVE MODE` inside the fence and before the re-inspection. That mode excludes every `INSERT` / `UPDATE` / `DELETE` on those two tables **for every graph on the database** — they are shared tables — while still admitting readers. So while a claim runs, every node and edge write database-wide waits. The bound is what makes it acceptable: the lock is taken **only inside a claim**, which happens when a sidecar is created, upgraded from the pre-marker schema, or resumed after a crash — never on the common path, where an already-owned sidecar opens with no fence at all. Its duration is the re-inspection's probes plus one `INSERT`, with no caller code and no caller I/O inside it. The mode is `SHARE ROW EXCLUSIVE` rather than plain `SHARE` because it must be self-exclusive: two concurrent claims on different sidecar ids hold different advisory locks, so under `SHARE` both would acquire it and then both request `ROW EXCLUSIVE` for their own marker insert — a lock-upgrade deadlock PostgreSQL resolves by aborting one of them. SQLite takes no such lock; `BEGIN IMMEDIATE` already owns the engine's single writer slot. Refusals carry the code `GRAPH_MERGE_PROVENANCE_ID_COLLISION` and one of five `details.reason` values — `application-graph`, `empty-legacy-sidecar`, `unupgradeable-legacy-sidecar`, `unowned-exact-schema-graph`, or `corrupt-ownership-marker` — so the remediation matches what is actually there instead of generic advice; a backend with no transactional schema fence refuses an unclaimed sidecar with `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED` (an already-owned sidecar still opens there). Under `persistProvenance: true` both of those arrive as a typed `InvalidMergeOptionsError` naming `details.option: "persistProvenance"`, with the originating `ConfigurationError` as its `cause` — see [Merge provenance sidecar codes](/errors#merge-provenance-sidecar-codes). Query persisted provenance back later: ```typescript import { openProvenanceStore, readProvenance } from "@nicia-ai/typegraph/graph-merge"; const store = await openProvenanceStore(target); const fromAgentA = await readProvenance(store, { branchId: "agent-a" }); // what did agent A contribute? const whoMadeX = await readProvenance(store, { canonicalId: "patient-123" }); // who contributed node X? ``` Inspection tools that have a backend and graph id but not the target's `GraphDef` can use the standalone overload: ```typescript const store = await openProvenanceStore(backend, targetGraphId); ``` ## Snapshot vs incremental A branch is forked from a `base@V` — a token combining the base's schema hash with the store's durable revision anchor when `revisionTracking: true` or `history: true` is on, or a complete live-content fingerprint otherwise. The revision anchor is namespaced by a durable per-graph origin, which `Store.clear()` rotates. A lineage-capable untracked store whose backend supports that origin relation also carries it beside its content fingerprint. The two merge entry points differ in how they treat that token. The token is printable text, so it can be stored anywhere an application keeps descriptors, plans, and fork points, including PostgreSQL `text` and `jsonb` columns. Treat it as opaque: compare it whole and never parse it. Tokens minted by releases before this format, which separated components with a NUL character, are refused with a `BaseVersionMismatchError` whose `details.reason` is `"legacy-token-format"`. Re-branch or re-plan from the current target. Earlier `engine:` anchors and untracked content tokens without the active schema version also require re-branching; they cannot match the current target's token. The token is printable text, so it can be stored anywhere an application keeps descriptors, plans, and fork points, including PostgreSQL `text` and `jsonb` columns. Treat it as opaque: compare it whole and never parse it. Tokens minted by releases before this format, which separated components with a NUL character, are refused with a `BaseVersionMismatchError` whose `details.reason` is `"legacy-token-format"`. Re-branch or re-plan from the current target. **`merge()` is a snapshot merge.** Every branch must have forked from the target's *current* `base@V`. If the target advanced since the branch was taken, `merge()` returns a `BaseVersionMismatchError` rather than risk clobbering newer data. This is the right model for "fork, do work, merge back" within one round. **`mergeIncremental()` is a fork-point merge into a live target.** It merges branches that forked from a frozen `forkPoint` into a `target` that may have *moved on*. Additions are re-discovered against already-committed entities (via `blockIndex` / unique constraints) so a re-seen entity updates the committed row instead of duplicating it. Inherited node and edge modifications/deletions are also propagated through the same three-way planner, with the live target kept authoritative when it changed concurrently. ```typescript import { mergeIncremental } from "@nicia-ai/typegraph/graph-merge"; const result = await mergeIncremental({ forkPoint, // the frozen ancestor the branches forked from target, // the live committed graph (may have advanced) branches, options: { resolve: { Patient: { blockIndex: "patient_cohort_idx", similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.85 } }, onBasePropertyConflict: "flag", // required: never overwrite a newer committed value }, }); ``` `mergeIncremental()` requires `onBasePropertyConflict: "flag"` — any other value is refused with `InvalidMergeOptionsError` — so a stale branch value can never overwrite a newer committed value during new-vs-base recall. The `forkPoint` must stay **frozen for the duration of the call**: every branch diff is computed against it, and the commit transaction re-reads its `base@V` before applying anything, so a write landing on the fork point mid-merge is refused with `BaseVersionMismatchError` instead of committing diffs against an ancestor that no longer exists. Only the `target` may advance while the merge runs. If both the branch and the live target changed the same inherited row, the target value/deletion wins and the conflict is reported. Both `merge()` and `mergeIncremental()` commit **transactionally** and require a transaction-capable target backend. Managed targets also acquire the schema-version write fence; raw targets remain outside schema fencing. On PostgreSQL, serialization failures from either the target-content guard or the schema fence are retried automatically around the complete commit. ### Lineage and pruned diffs A backend may declare a `lineage` capability: an opaque, whole-database `revision(session)` it can report and compare, plus `changesSince(session, revision, graphId)`, which names every node and edge of one graph that changed (inserted, updated, deleted, or resurrected) after that revision — or admits `{ kind: "unbounded" }` when it cannot bound the answer. Bundled SQLite and PostgreSQL stores with `revisionTracking: true` also provide bounded lineage through a DML journal when history capture is disabled. `lineageRevisionNow()` mints a public anchor and `changesSince(anchor)` returns changed node and edge keys. Use that anchor API rather than `revisionNow()`, which returns a clock value without the graph's origin identity. The journal is installed when the store is provisioned through `createStoreWithSchema()`, or explicitly with `installRevisionChangesJournal(backend)` from `@nicia-ai/typegraph/schema` under a schema owner role. Existing installations must first adopt base schema version 4 through a privileged schema open or generated base-schema migration. Runtime lineage checks the journal and its triggers without issuing DDL; a revision-tracked store without history fails with `REVISION_JOURNAL_NOT_READY` when the journal is not ready. Short-lived clones that do not need this bounded lineage can set `revisionJournal: false`. Writes before the first anchor are outside that anchor's range. Node and edge inserts, updates, and deletes are recorded by database triggers. Identity-only revisions and revisions whose write provenance is incomplete produce `{ kind: "unbounded" }` rather than an incomplete key list. Custom backends must provide their own lineage capability to get bounded results. Each trigger is attached to a whole physical node, edge, or identity table; it records every write to that table and uses `graph_id` to identify the affected graph. On shared tables this captures writes from every graph, not only graphs whose stores enabled the journal. Journal rows are retained per revision and never cleaned up automatically; applications should avoid installing triggers on shared tables unless cross-graph capture is intended, and should plan an external retention policy that preserves every revision still used as a branch anchor. `resolveLineage(store)` selects backend lineage first, then captured history, then the first-party revision journal. A lineage source is consulted only to avoid rework; it never changes what a merge decides. `revision()` reports `:`, never the bare clock value alone: the durable, random per-graph revision-origin nonce (`typegraph_revision_origins`) plus the recorded-time clock. Two independently created stores that share a `graphId`, or the SAME store across a `Store.clear()` boundary, can mint numerically comparable clock values, and the origin is what keeps `changesSince` from mistaking one for the other — a revision whose origin no longer matches the graph's LIVE origin row is `unbounded`, regardless of what its numeric clock value is. The recorded-relations derivation's delta is trustworthy only when EVERY writer to the graph goes through a store that captures history — a precondition it can partially, but not fully, enforce itself. `changesSince` proves completeness directly rather than inferring it from a high-water mark: every integer revision between the requested one and the graph's current clock must carry direct evidence — a `recorded_from` or a non-sentinel `recorded_to` — in one of the three recorded relations (nodes, edges, identity assertions). This catches an incomplete record wherever the hole falls, including a `revisionTracking`-only `Store` (no `history`) that advanced the shared clock without inserting a row and was later FOLLOWED by a capturing commit — a later capturing commit cannot retroactively supply the missing evidence, so the gap is caught regardless of what comes after it. What it CANNOT detect: a non-capturing writer bypassing every `Store` entirely (a raw `GraphBackend` write, or an engine-side mutation outside TypeGraph), which leaves no evidence to be short of. Route every writer through a capturing `Store` if a `"keys"` delta from this source must be exhaustive. `session` is the connection the caller's decision is bound to — a session-less bag could never be pinned to anything, so this one always carries one. A caller planning outside any transaction (`branch()`'s fork-revision capture, the pruning below) passes the root backend it holds; a caller re-validating a content fingerprint inside an open commit transaction reads through that transaction's own handle, so the fingerprint observes the transaction's snapshot and establishes dependencies on the rows it covers. **Untracked stores use a complete fingerprint.** An engine-wide revision and node/edge-only `changesSince` result cannot fence an identity-only write. It also cannot establish read dependencies on the graph state used in planning. For this reason, a store without TypeGraph revision tracking fingerprints live nodes, edges, and current identity assertions even if its backend exposes `lineage`. Where supported, the token also carries the durable graph origin. The commit transaction checks the origin and recomputes the fingerprint before applying its writes. Previously minted `engine:` base tokens are retired; re-branch from the current store rather than applying an old merge. **`Store.clear()` rotates the revision origin.** For revision-tracked stores, `clear()` deletes and re-mints the per-graph origin in the same transaction. A branch forked before that clear cannot merge into the post-clear store even when its revision clock has the same numeric value. The origin row is also read fresh on every mint (`computeBaseVersion`, `Store.revisionOriginNow()`), never cached on a `Store` instance. Two live `Store` objects can legitimately observe the same graph — nothing requires that only one `Store` ever exists per database — and only one of them runs `clear()` at a time; a stale per-instance cache on the other would keep minting anchors from the origin that existed before the clear, so a branch it forks would fail every merge at commit until that `Store` happened to be recreated. Reading fresh means a second `Store` over a graph another `Store` just cleared sees the rotation immediately, with nothing to recreate. **Pruning the diff.** `branch()` also records a `forkRevision` on the returned `GraphBranch` — the fork's own `lineage.revision(session)`, read right after the working copy is created and before any write reaches it, with the working copy's own root backend as the session (this runs strictly outside any transaction). For the recorded-relations source this is origin-bearing like any other reading, so clearing and repopulating the FORK itself to the same revision count `forkRevision` held is caught the same way a cleared BASE store already is — there is no separate guard for the fork side to add, because the token itself now carries the check. When staging a branch for merge, its diff against the base is restricted to the union of two deltas: what changed on the *fork* since `forkRevision`, and what changed on the *base* since the anchor in its own `base@V` — instead of enumerating every live row on both sides. A key absent from both deltas cannot have changed since the fork point, so narrowing the read to their union cannot miss anything the full diff would have found; it only fetches fewer rows to compare. Pruning is a pure optimization with one rule: whenever either side cannot supply a bounded delta, the merge falls back to comparing every live row, exactly as it always has. That covers no `forkRevision` (a hand-built branch, or one whose store resolved no `lineage`); either side's `changesSince` answering `unbounded` or REJECTING (a transient engine error never fails a merge the full diff would have completed); and the base's own anchor failing to resolve against the base store's lineage at all — an origin mismatch between a revision-anchored `base` and the base store's live revision row, a revision anchor minted before the base store's first tracked write, or an old engine anchor that must be re-branched. Nothing about *what* a merge decides depends on whether its diff was pruned. ## Working copies `branch()` is backend-agnostic. The default `cloneWorkingCopyStrategy` exports the base through TypeGraph's interchange and imports it into a fresh store on a backend your factory provides — so it works identically across SQLite, Postgres, and in-process PGlite, and needs no schema changes. The import is fidelity-preserving: undeclared properties that `validateStore()` treats as healthy semi-structured data are carried through. Stripping them would make a later merge invent deletions against the original base. ```typescript // Each branch gets its own in-memory SQLite backend: import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; const makeBackend = async () => createLocalSqliteBackend().backend; const fork = unwrap(await branch(base, makeBackend, { id: asBranchId("worker-1") })); ``` For a custom isolation mechanism (e.g. a future copy-on-write namespace), pass a `WorkingCopyStrategy` as the fourth argument to `branch()` — its single `create` method receives the base store and the `BaseVersion` `branch()` already stamped off it, and returns an independently-mutable store over the same graph definition. **A branch is a data fork.** `branch()` records the clone's committed schema `(version, hash)` at fork time, and the merge refuses (typed, as `BaseVersionMismatchError`) any branch whose store ran a schema operation afterwards — `evolve()`, `migrateSchema()`, or `removeKinds()` — even a round-trip migration that restores the original document hash. Those operations mutate rows through their own preflights, and projecting the side effects into a merge would detach them from the schema change that caused them. Apply schema changes to the target first (or re-fork), then merge. ### PostgreSQL table-backed working copies `createPostgresWorkingCopyManager` allocates a private set of TypeGraph tables in the source PostgreSQL database. It derives the table inventory and base schema marker from TypeGraph's PostgreSQL schema contributions, copies the source graph with fenced `INSERT ... SELECT` statements, and records ownership in `typegraph_working_copy_allocations`. The control backend, source backend, and backends returned by `connect` must all reach the same database, and `control` and `connect` must run as the same role ([One database role](#one-database-role)). TypeGraph checks the allocation's private ownership token through each connection. The control backend must execute DDL inside its PostgreSQL transactions; its root `executeDdl` port is not required. ```typescript import { drizzle } from "drizzle-orm/node-postgres"; import { createPostgresBackend, createPostgresTables, } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { createPostgresWorkingCopyManager } from "@nicia-ai/typegraph/adapters/drizzle/postgres/working-copy"; import { asBranchId, branchDurable, destroyDurableBranch, reopenDurableBranch, unwrap, } from "@nicia-ai/typegraph/graph-merge"; const control = createPostgresBackend(drizzle(pool)); const copies = createPostgresWorkingCopyManager({ control, connect: (names, allocation) => Promise.resolve( createPostgresBackend(drizzle(pool), { tables: createPostgresTables(names), ...(allocation === undefined ? {} : { vector: allocation.vectorStrategy }), }), ), }); const { branch: copy, descriptor } = unwrap( await branchDurable(sourceStore, copies.durable, { id: asBranchId("candidate-42"), allocationId: "candidate-allocation-42", }), ); await copy.close(); // Releases the connection; the tables remain. const reopened = unwrap( await reopenDurableBranch(graph, descriptor, copies.durable), ); await reopened.close(); unwrap(await destroyDurableBranch(descriptor, copies.durable)); ``` The same manager exposes `ephemeral` for `branch()`; closing that branch drops its tables. `listUnsealedAllocations({ after, limit })` pages through durable allocations awaiting seal and ephemeral allocations. These rows may still have active owners; the ledger alone cannot identify a crashed process. After confirming that no live branch or allocation uses a row, call `abortAllocation(id)` to remove it. A durable branch's descriptor contains only the allocation ID, not connection credentials. Pass `sourceTableNames` when the source backend uses custom status table names; pass `reopenOptions` to restore process-local hooks or query options on a later process. An external `recordedRead` binding is refused because its relation is outside the owned table inventory. Reopen options cannot replace the allocation's schema, recorded-read binding, history mode, or revision-tracking mode. Pass `operations` to let `durable.operations` commit host mutations atomically with immutable evidence; see [Atomic operations and immutable evidence](#atomic-operations-and-immutable-evidence). Each durable allocation also owns an evidence table (`op_evidence`) in the allocation's schema, addressed through that schema rather than the connection's `search_path`; destroy refuses to drop it while undelivered evidence remains. The source backend and every backend returned by `connect` must expose the complete PostgreSQL `tableNames` inventory, including history, identity, and status relations. The manager refuses missing or mismatched bindings with a `BranchError` before cloning or opening a Store. For `ephemeral` and `durable`, `connect` runs after the allocation tables are created, so custom callbacks may inspect those tables; on binding failure, the manager removes the new tables and ledger row. `makeBackend` connects earlier, before it provisions anything. The table-backed strategy supports bundled tsvector fulltext, declared PostgreSQL B-tree, GIN, and trigram graph indexes, and pgvector sidecars. It builds each declared graph index on private tables under stable allocation-scoped physical names while keeping logical index names and schema hashes unchanged. `materializeIndexes()` can retry or repair indexes after reopen; destroy removes their owned tables and indexes. When `connect` receives an allocation vector strategy, pass it to `createPostgresBackend`; the strategy assigns stable table and index names from the ledger-reserved physical prefix. Allocation claims and all initial table and vector DDL commit together, so a colliding or failed provision leaves no partly owned sidecars. Source vector sidecars are copied under the same transaction locks as TypeGraph relations. The ledger stores every relation name declared by each slot's `ownedTables()` contribution, so destroy can remove them in reverse declaration order without a graph object. Reopening requires the graph's vector slots and owned-relation inventory to match the persisted allocation manifest. Older ledger rows that stored only `tableName()` remain readable as single-relation slots. A declared vector slot whose source sidecar is absent is refused because its contents cannot be snapshotted exactly. The `ephemeral` and `durable` copies have a fixed schema: `evolve`, kind removal, and deprecation refuse before mutation. Use `makeBackend`, below, when the working copy's schema must change. Custom fulltext strategies still need a host-level database fork. The source and every copy connection, including durable reopen, must use the bundled `tsvectorStrategy`: a custom strategy may own additional physical tables whose rows cannot be copied safely from the generic contribution inventory. A connection with fulltext disabled is refused for the same reason. System index maintenance remains available. Source table locks cover the entire TypeGraph relation set and vector sidecars while the SQL clone runs, so a large clone briefly blocks writes to other graphs in the same database. #### One database role The manager supports one deployment shape: the `control` backend and every session `connect` returns run as the **same PostgreSQL role**. TypeGraph reads `current_user` on both sessions and refuses a difference with a `ConfigurationError` whose `details.code` is `WORKING_COPY_ROLE_MISMATCH`, and the refused allocation is not left behind. The reason is ownership. A `control` session provisions and removes every allocation, but the Store that opens on a connected backend issues its own DDL: runtime-contribution markers, the revision journal and its triggers, system and declared indexes, and vector tables an evolved graph introduces. Only a table's owner (or a member of the owning role, or a superuser) can drop it, and the comparison is by role name, so a `connect` role that is merely a member of `control`'s role is refused rather than trusted. A different role would leave the tables it creates behind on close and `abortAllocation`. The shared role therefore needs `CREATE` on the schema. `makeBackend` calls `connect` before it writes the ledger row or any DDL and refuses a mismatch there, so nothing is allocated. `ephemeral` and `durable` call `connect` after their allocation tables exist, so they refuse right after it, before cloning or opening a Store, and remove the new allocation; a durable reopen refuses the same way and leaves the sealed allocation untouched. #### One schema per allocation Every allocation lives in one schema: the `control` session's current schema when the allocation is made, recorded in the ledger's `schema_name` column. No `search_path` decides where an allocation's relations are created or dropped, so a `connect` pool whose connections lead with different schemas cannot strand tables that removal never finds. - **Provisioning** fixes its transaction's search path to that schema before it claims the ledger row, so the tables it creates land there whichever pooled connection runs it, and the claim records the schema the statement itself observed. - **The connected backend** receives table names that carry the schema. A backend built with `createPostgresTables(names)` over that object runs the DDL it issues lazily (bundled tables a Store ensures on first use, fulltext and contribution storage, schema-write transactions) with the schema leading its search path, and the allocation's pgvector strategy names its tables and indexes through the schema. `CREATE INDEX CONCURRENTLY` cannot run in a transaction; it creates the index in the schema of the table it names, which is already the allocation's. The backend's catalog probes (table, index, and column lookups, including the recorded-time compatibility check a `history: true` Store runs) read the allocation's schema, not the session's current one. Extensions are database-global and create no allocation relation, but their DDL still runs through the same DDL runner wherever the write fence takes no lock: there, a backend built over a caller's own transaction is subject to the same session check as any other lazy DDL (below). Under a lock fence, a pooled backend installs the extension in its own transaction, as before; a backend built over a caller's own transaction runs it as a savepoint inside that transaction and makes no session check, because the extension creates no allocation relation. - **Refusals.** A connection whose backend was built over a *copy* of `names` (which carries no schema) is refused with a `BranchError`. A `connect` driver that cannot hold an interactive transaction (`drizzle-orm/neon-http`) is refused with a `ConfigurationError` (`ALLOCATION_SCHEMA_REQUIRES_INTERACTIVE_TRANSACTIONS`), because it cannot run its DDL under a fixed schema. A backend built over a caller's own transaction runs its lazy DDL and schema writes, and adopts that transaction for a schema write, only when that session's current schema is the allocation's; otherwise it is refused with a `ConfigurationError` (`ALLOCATION_SCHEMA_SESSION_MISMATCH`). The caller owns that session's search path, so it is checked rather than rewritten. - **Removal** (`close`, `abort`, `destroy`, `abortAllocation`) searches the catalog across every schema for relations named with the allocation's reserved prefixes. It drops those in the recorded schema, schema-qualified in one statement, and deletes the ledger row in the same transaction. If a drop fails (a view that depends on an allocation table, for example) the transaction rolls back, the row stays, and the allocation remains in `listUnsealedAllocations()` for `abortAllocation()` once the dependency is gone. If any such relation sits in a different schema, removal refuses with a `BranchError` that names the schemas found and keeps the row, because deleting the row would discard the only pointer to them. Three cases are worded differently. When the recorded schema holds none of them, the schema was renamed or the tables moved (the message says the relations are "not in its schema"; move the tables back or correct the row's `schema_name` and remove again). When every relation found elsewhere has a same-named relation in the recorded schema, it is a stale copy left in another schema, such as a backup or restore schema (the message says the allocation "also has relations" there; drop the copy and remove again, since the copy blocks removal until it is gone). When some relations moved and others stayed, for example one table moved to a backup schema while the rest remain, the allocation is split and the relations elsewhere may be the only copy (the message says the allocation "is split across schemas"; the suggestion drops nothing, so move the relations back or correct the row's `schema_name`). `details` carries `allocationId`, `schema`, `foundIn`, and `schemas`, and `suggestion` names the recovery step. If the allocation's relations exist nowhere (its tables were dropped entirely) there is nothing to recover, and removal deletes the ledger row, so a crashed owner's allocation cannot stay listed forever. - **Ledger rows from before the schema was recorded** (written by 0.72.0) carry no schema. They resolve through the session that removes them and reopen without binding, and follow the same removal rule: relations found in a schema other than the removing session's refuse removal and name that schema. `control` adds the column to an existing ledger the first time it runs. The connection must still be able to *resolve* the allocation's tables, so its `search_path` must include the schema, typically `public`. A per-role `"$user"` schema ahead of it is fine. The ledger itself lives where `control`'s session creates it, so run `control` with one consistent `search_path`. #### `makeBackend` for branches, candidate planning, and evolution previews `copies.makeBackend` is a `MakeBackend`, so PostgreSQL callers no longer hand-roll table prefixes, DDL, and cleanup. It fits every API that takes one: `branch`, `ingestionBranch`, `planCandidateWriteSet`, `planCandidateWriteSetReview` (including sparse staging), `branchForEvolution`, and `planCandidateWriteSetForEvolution`. ```typescript import { branch, branchForEvolution } from "@nicia-ai/typegraph/graph-merge"; const fork = unwrap(await branch(sourceStore, copies.makeBackend)); const preview = unwrap( await branchForEvolution(sourceStore, evolutionPlan, copies.makeBackend), ); ``` Each call allocates a fresh allocation in the same ledger, in the `ephemeral` state, and returns an **empty, schema-mutable** backend: the caller (or the branch API) seeds it and may commit new kinds and fields, which the fixed-schema `ephemeral` and `durable` copies refuse. Closing the backend drops the allocation. While it is live it appears in `listUnsealedAllocations()`, and if its owner crashes without closing it, `abortAllocation(id)` removes everything it owns. Because the graph is unknown when the backend is allocated: - **Vector tables.** A graph that declares embeddings creates its per-field pgvector tables after allocation, so the ledger manifest cannot list them. Dropping an allocation therefore also removes every table in its schema whose name starts with the allocation's reserved vector prefix. That prefix is fixed-length and never truncated, so it cannot match another allocation's tables. `connect` always receives the allocation vector strategy for `makeBackend`; bind it with `createPostgresBackend({ vector: allocation.vectorStrategy })`. A connection that binds any other vector strategy is refused with a `BranchError`, because it could create tables the allocation does not own. Pass `vector: false` to opt out of vector support. - **Graph indexes.** PostgreSQL index names are database-global, so a declared index cannot reuse its logical name on a private table. `makeBackend` scopes each declaration to the allocation (`gix_`) the first time the Store's `materializeIndexes()` sees it, leaving logical names and schema hashes unchanged and never touching the source's or another allocation's indexes. A backend you derive from the returned one with `deriveBackend` inherits the scoping; one you build by copying its members does not. - **Fulltext.** The same bundled `tsvectorStrategy` requirement applies as for the cloned copies. `control` and `connect` must run as the same role ([One database role](#one-database-role)). Both must also use the allocation's schema ([One schema per allocation](#one-schema-per-allocation)); a pooled connection's own `search_path` does not decide where anything is created. ### Forked working copies A second bundled strategy, `forkedWorkingCopyStrategy({ fork, connect })`, targets a fork-capable host instead of a streamed-interchange clone: `fork` asks the host itself to produce a complete, independent copy of the database `baseStore` is on, and `connect` opens a backend on that copy. ```typescript import { asBranchId, branch, forkedWorkingCopyStrategy, unwrap, type ForkHandle, } from "@nicia-ai/typegraph/graph-merge"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { decorateBackend } from "@nicia-ai/typegraph/backend"; import { drizzle } from "drizzle-orm/node-postgres"; import { Pool } from "pg"; // A host whose fork call returns a new connection string for the branch — // this is the shape of the copy-on-write branching APIs some Postgres hosts // offer (Neon and Supabase branches, for example), without either SDK. type HostBranch = ForkHandle & Readonly<{ connectionString: string }>; const strategy = forkedWorkingCopyStrategy({ fork: async () => { const created = await hostBranchApi.createBranch(baseDatabaseId); return { connectionString: created.connectionString, dispose: async () => hostBranchApi.deleteBranch(created.id), }; }, connect: async (fork) => { // `createPostgresBackend` takes a Drizzle database, not a pool — open // one here. Its `close()` deliberately does not end a caller-owned pool // (Drizzle leaves connection lifecycle to the caller), so compose the // pool's own shutdown into this fork's `close` through the public // `decorateBackend` (never a spread) — `branch()`'s composed close then // ends the pool along with releasing the fork. const pool = new Pool({ connectionString: fork.connectionString }); const backend = createPostgresBackend(drizzle(pool)); return decorateBackend(backend, { close: async () => { await backend.close(); await pool.end(); }, }); }, }); // `makeBackend` is ignored once an explicit strategy is supplied — pass a // factory whose only job is to reject if it is ever called by mistake. const rejectMakeBackend = () => Promise.reject(new Error("makeBackend must not be called")); const worker = unwrap( await branch( base, rejectMakeBackend, { id: asBranchId("worker-1") }, strategy, ), ); // ... write on worker.store, plan and apply the merge ... await worker.close(); ``` `TFork` must extend `ForkHandle` (`{ dispose?: () => Promise }`). `forkedWorkingCopyStrategy` supplies ephemeral copies only. Its base-version comparison checks the graph's schema and revision or live-content anchor; the host fork must preserve the full physical database, including TypeGraph sidecars and extensions. A durable host strategy must persist its branch ID and attest the sealed origin when reopening it. For a hosted PostgreSQL branch such as [Neon](https://neon.com/docs/get-started-with-neon/workflow-primer), `connect` must use that branch's connection string and compute endpoint for every pooled checkout and transaction. Reusing the source pool can appear to pass a base-version check while writing to the source. Doltgres can pin a connection through a [database revision specifier](https://www.doltgres.com/docs/reference/version-control/branches/); avoid session-level branch switching on a pool whose checkouts may retain different branch state. Doltgres exposes native branch and merge commands, but TypeGraph continues to use its own merge planner and apply path; native merge and Doltgres backend support require separate conformance testing. `create()` calls `fork(baseStore)`, then `connect(fork)`; the connected backend's `close` is composed with the fork's `dispose` through `deriveBackend` (never a spread), so `worker.close()` — the branch's public release call — releases both the connection and the fork. A `connect` failure disposes the fork before rethrowing, leaving nothing open and the base untouched. A fork inherits the base's WHOLE construction option set — hooks, upsert coalescing, the SQL schema (custom table names), the auto-refresh-statistics threshold, query defaults, and an externally-bound recorded-read relation — read once through `Store.workingCopyOptions`, plus `history`/ `revisionTracking`, matched to the base's own `historyEnabled`/ `revisionTrackingEnabled`. This is safe precisely because a fork is the SAME physical database as the base: a custom `schema` names relations the fork carries too, and an external `recordedRead` binding points at one. The clone strategy inherits only `revisionTracking` — its fresh backend is a distinct, empty database, so a schema naming the base's tables or a `recordedRead` binding populated nowhere on the clone would misdirect it. Because the fork's store reads and writes through the base's table names, `connect()`'s backend must bind those SAME names. `create()` compares the connected backend's own table bindings against the base's own resolved SQL schema (`Store.revisionSchema` — the base's explicit `schema` option, or its backend's own `tableNames` otherwise), and refuses with a `BranchError`, closing the backend first, when they disagree: a backend bound to different (often just the default) table names would read and write through tables the fork's rows were never written to. **A fork preserves what a clone drops, and that is why it is safe to merge.** The clone strategy above streams the base through public interchange with `includeDeleted: false`, so it omits every soft-deleted row entirely: the interchange `meta` schema has no `deletedAt` field, so a tombstoned row would otherwise round-trip as LIVE and read as a spurious resurrection on the clone's diff. It also regenerates `created_at`/`updated_at` on import — safe only because the merge's state diff always compares against the *original* base store, never the clone. A fork is never rebuilt through `exportGraphStream`/`importGraphStream`, so none of that applies: tombstones, `created_at`/`updated_at`, and the `version` column carry over unchanged, and — with `history: true` — the fork physically carries the base's recorded relations, so `store.asOfRecorded()` answers from that history. A clone-based branch never enables history, so the same call on it refuses outright. `create()` asserts `computeBaseVersion(forkStore) === base` right after attaching the store, where `base` is the token `branch()` already stamped off the ORIGINAL base store before invoking the strategy — cheap when the base has revision tracking (an O(1) anchor compare), an O(graph) content fingerprint otherwise, and computed exactly once either way. This proves base-token equality at the instant the fork was taken, not byte-for-byte physical identity: the untracked fingerprint deliberately omits tombstones, `created_at`/`updated_at`, the `version` column, and recorded history (the "A fork preserves what a clone drops" paragraph above) — providing those unchanged is the FORK MECHANISM's job, not something this assertion re-verifies on every branch. That is still the right fence: the merge's lost-update guard reads `version` and the diff reads tombstones/timestamps straight off the fork, so a `fork` that is not a true physical copy breaks them regardless of what the content fingerprint agrees on. A mismatch closes the backend first and refuses with a `BranchError` carrying `forkVersion`/`baseVersion` in `error.details`; `branch()` catches it and returns that `BranchError` as the `cause` of the outer `BranchError` it resolves with. Only a base-token mismatch is refused here — a fork taken while the base was mid-write, or a `fork` that returns a different graph; divergence confined to the physical state the token omits (tombstones, timestamps, row versions, recorded history) passes the fence, and keeping that state faithful remains the fork mechanism's contract. `create()` also refuses BEFORE ever attaching a store when `connect()`'s backend aliases the base's own backend: the same backend object, one derived from the other through `deriveBackend`, or two wrappers sharing one underlying connection. Without this check, a `connect()` that mistakenly hands back the base's own backend (a cached factory keyed by database name, say) would pass every fence below trivially — every write on the "fork" would actually mutate the base, and closing the working copy would close the base's own backend. The refusal disposes only the fork (never the aliased backend, which the base still owns) and throws a `BranchError` naming `connect()`. This cannot detect every aliasing shape: a fresh backend built over the base's own connection pool is indistinguishable from a real fork's connection when that pool audits as independent (the normal case for a default-size `pg.Pool`) — a pooled checkout genuinely is a different connection from the pool's perspective. `ingestionBranch()` stays clone-based. Its strategy derives a working-copy schema with node uniqueness deferred so an untrusted batch's repeated keys can reach entity resolution before validation; a host-level fork carries the base's schema exactly, uniqueness included, with no hook to relax it. :::caution[Suspend hazard] A fork-capable host that suspends idle compute to reclaim it between requests drops that compute's in-process state, including anything memoized against a particular connection or session. TypeGraph's own locking already assumes this rather than trusting a lock survives idle time: the recorded-write lock memo (`RecordedGraphLockMemo`, populated by `memoizeAcquiredRecordedGraphWriteLock`) and the schema-fence lease (`memoizeLeasedSchemaFence`) are both keyed weakly by the transaction-scoped backend object, so they hold for exactly one transaction's lifetime and re-acquire on the next one, and the write fence itself (see [Write fence declaration](/backend-setup#write-fence-declaration-writefence)) is resolved and its lock taken fresh per transaction, never cached across one. An ordinary sequence of separate `store` calls — each its own transaction — therefore tolerates a suspend between any two of them. What does NOT tolerate a suspend is a single `store.transaction` callback: every read and write the callback issues, and the lock it holds, runs on one native database transaction over one connection, so a suspend partway through drops that connection out from under the callback and aborts whatever was in flight. Keep a `store.transaction` callback's wall-clock duration short and free of anything that could let the host suspend underneath it — an external API call, a human approval step, a long queue wait — and commit a long-running workflow across multiple `store.transaction` calls instead of holding one open across such a wait. ::: ### Durable host-native branches `branchDurable()` is the persistent counterpart to `branch()`. A `DurableWorkingCopyStrategy` allocates a host branch, opens a Store on it, and returns a non-secret JSON locator. TypeGraph seals the immutable fork origin beside that allocation and returns a `DurableBranchDescriptor` that can cross a queue, process, deployment, or machine boundary. For a remote host, persist a chosen `{ id, allocationId }` before calling `branchDurable(base, strategy, { id, allocationId })`. `create()` receives both and must refuse an allocation ID that may already exist. If the host allocates a branch but its response is lost, use host tooling to inspect the ID and recover or remove the allocation before retrying. A failed create reports both IDs for that reconciliation. The host must never allocate a second physical copy for the same ID or return a sealed copy as though it were new. ```typescript import { applyDurableMergePlan, branchDurable, destroyDurableBranch, planMerge, reopenDurableBranch, unwrap, } from "@nicia-ai/typegraph/graph-merge"; const created = unwrap(await branchDurable(base, durableStrategy)); await created.branch.store.nodes.Person.create({ name: "Ada" }); // Releases this process's connection and writer lease. The host branch stays. await created.branch.close(); await queue.put(JSON.stringify(created.descriptor)); // A later process reconstructs the ordinary GraphBranch used by planning. const descriptor = JSON.parse(await queue.get()) as typeof created.descriptor; const reopened = unwrap( await reopenDurableBranch(graph, descriptor, durableStrategy), ); const plan = unwrap(await planMerge(base, [reopened])); // Applies the complete TypeGraph plan inside the target transaction. const report = unwrap( await applyDurableMergePlan({ target: base, branch: reopened, descriptor, strategy: durableStrategy, plan, }), ); await reopened.close(); unwrap(await destroyDurableBranch(descriptor, durableStrategy)); ``` Closing and destroying are deliberately separate. `GraphBranch.close()` closes the backend and releases its access lease, but leaves the persistent allocation reopenable. `destroyDurableBranch()` asks the strategy to attest the complete origin and delete or archive that allocation atomically. A descriptor is untrusted input: TypeGraph checks its allocation id, graph definition, branch id, base token, schema anchor, and engine revision against the origin the host sealed. The allocation id is independent of the caller's branch id, so swapping or relabeling a locator cannot authorize deletion of another copy even when two copies were given the same branch id. Strategies write new locators using `version` and may list older supported locator versions in `readableVersions`. Every method must understand each listed version, including destroy and evidence access. The strategy locator must be JSON-safe and **must not contain secrets**. Use a branch id, database id, or other lookup key, then resolve credentials from strategy-owned configuration. TypeGraph returns the locator to application code so a connection URL, password, or bearer token placed there can escape through ordinary descriptor storage. Framework cleanup errors deliberately omit the locator and raw host cleanup error from diagnostic details. #### Exact forks and access leases After `strategy.create()` returns, TypeGraph recomputes `base@V` from the source. A source write racing allocation therefore refuses and aborts the working copy instead of sealing a branch from the wrong ancestor. TypeGraph then accepts an exact matching working-copy token as the fast path. When a strategy creates an equivalent persistent copy with an independent revision namespace, TypeGraph instead verifies that its complete merge-visible graph state has no delta from the source, fencing the source again after enumeration. The host remains responsible for physical fidelity outside TypeGraph's graph semantics. To enable lineage-pruned merge diffs, `create()` may return `forkRevision` captured atomically with the physical fork. When it cannot prove that cut, omit the revision and TypeGraph compares the complete graph state; reading a later revision after the copy was opened could miss an intervening branch write. Every `create()` and `reopen()` also returns a `DurableWorkingCopyAccess`: - `engine-fenced` says the database provides sound cross-client isolation and change fencing for the full Store planning/apply access pattern, across every connection and process that could mutate the working copy. - `exclusive` carries an allocation-wide writer lease. The strategy must acquire it before returning and exclude every other process and backend instance. TypeGraph closes the backend first, then releases the lease; a failed release is retried by the next `close()` call. Do not use `engine-fenced` merely because one backend object serializes its own calls. A `caller-serialized` backend owns one in-memory queue per backend instance, so two reopened pools or two processes still race. Such an engine must use a host-wide `exclusive` lease, and a concurrent reopen must wait or refuse. Merge planning also assumes the working copy is quiescent while it is diffed. #### Native database branches A strategy may allocate a working copy using a database-native branch, but `applyDurableMergePlan()` always applies the approved TypeGraph plan through the target Store transaction. The former native-merge callback was removed: it could commit outside the transaction that checked the target revision. A future native merge capability needs a host-native compare-and-swap on the actual target, plus proof that the full physical diff equals the approved TypeGraph writes, including schema, history, identity, and sidecars. For a Doltgres strategy, pin each Store connection to the intended database branch. [Doltgres revision specifiers](https://www.doltgres.com/docs/reference/version-control/branches/) provide that connection-level selection. Its [`DOLT_BRANCH()` and `DOLT_MERGE()` functions](https://www.doltgres.com/docs/reference/version-control/dolt-sql-functions/) implicitly commit the current transaction, so a fence checked before those functions cannot by itself protect their target. #### Atomic operations and immutable evidence A `DurableWorkingCopyStrategy` may also expose an optional `operations` capability (`DurableOperationCapability`). It lets a durable host combine one opaque graph mutation with its immutable operation evidence in a **single host transaction**. TypeGraph owns descriptor validation, sealed-origin attestation, request canonicalization, and evidence validation; the host owns the database mechanics. ```typescript import { durableBranchHasUndeliveredEvidence, getDurableOperation, markDurableOperationDelivered, operateDurableBranch, scanDurableOperations, unwrap, } from "@nicia-ai/typegraph/graph-merge"; const request = { idempotencyKey: "statement-42", // Host-defined, JSON-safe description of the graph change to apply. mutation: { kind: "statement", op: "upsert", payload: { subject: "s-1" } }, // Host evidence, retained verbatim. TypeGraph never interprets either field. metadata: { source: "etl", schemaVersion: 3 }, }; const outcome = unwrap( await operateDurableBranch(descriptor, durableStrategy, request), ); if (outcome.outcome === "unsupported") { // The strategy applied no mutation and wrote no evidence; TypeGraph refuses // rather than emulating atomicity with best effort or callbacks that run // outside the evidence transaction. throw new Error(`Missing capabilities: ${outcome.dimensions.join(", ")}`); } console.log(outcome.outcome); // "applied" | "replayed" // Newly applied evidence is always false. A replay returns the current // committed delivery state, which may already be true. console.log(outcome.evidence.delivered); ``` Both `mutation` and `metadata` are **JSON-safe host values**. TypeGraph never interprets their application fields; it canonicalizes `metadata` plus `mutation` into the `operationDigest` and otherwise carries them through untouched. The digest covers the complete request except the idempotency key, so reusing a key with a different mutation *or* different metadata conflicts. Non-JSON content is refused before any host call. `metadata` is retained as evidence; `mutation` is the host's own description of the graph change it must apply atomically with the evidence row. The strategy attests the caller's `expectedOrigin` against the allocation the descriptor names, exactly as reopen and destroy do. Every committed operation returns `before`/`after` coordinates — the merge-visible `base` fingerprint and, when the working copy resolves lineage, the engine `revision`. TypeGraph validates that the returned evidence echoes the canonical request and digest; a host cannot forge a different digest, echo a different request, or return non-JSON metadata (`DurableOperationEvidenceError`). **Idempotency.** The strategy treats `idempotencyKey` as its unique key: - Identical key **and** digest: returns the previously committed evidence (`outcome: "replayed"`) and re-applies nothing. Because delivery marking is monotonic, a replay after delivery legitimately returns `delivered: true`. - Identical key with a **different** digest: refuses with `DurableOperationConflictError` and mutates nothing. A first application (`outcome: "applied"`) must return `delivered: false`. TypeGraph rejects `applied` evidence that is already delivered, so a host cannot bypass downstream delivery or the destroy fence. It also validates the complete host outcome envelope: malformed outcomes and empty, duplicate, or unknown `unsupported` dimensions return `DurableOperationEvidenceError`. **Evidence access and delivery.** - `getDurableOperation(descriptor, strategy, idempotencyKey)` reads one operation's evidence, or `undefined` when it was never committed. - `scanDurableOperations(descriptor, strategy, { after?, limit? })` returns `{ operations, cursor, hasMore }` in monotonic commit order, with ties broken deterministically. Pass the opaque `cursor` back as `after` to resume, even after `hasMore: false`; later commits must sort after that cursor. An empty page echoes `after`, and only an empty initial scan omits `cursor`. `limit` defaults to `DURABLE_OPERATION_SCAN_DEFAULT_LIMIT` (100) and may not exceed `DURABLE_OPERATION_SCAN_MAX_LIMIT` (1000); a larger page is refused. - `markDurableOperationDelivered(descriptor, strategy, idempotencyKey)` marks one operation delivered, idempotently: marking an already-delivered operation returns the same evidence and writes nothing, and an unknown key returns `undefined`. - `durableBranchHasUndeliveredEvidence(descriptor, strategy)` reports whether any committed evidence is still undelivered — the queryable half of the destroy fence below. `operateDurableBranch()` is the only orchestrator that tolerates a missing capability: a strategy with no `operations` returns the explicit `unsupported` outcome (`dimensions: ["atomicMutation"]`) having executed no host call. `get`, `scan`, `markDelivered`, and `hasUndelivered` instead refuse with a typed `DurableOperationUnsupportedError`. TypeGraph never emulates the atomic guarantee: a callback that runs inside the strategy's own evidence transaction (as `apply` does in the bundled PostgreSQL manager below) is the host's atomic mutation, while best effort or a callback outside that transaction is refused. **Destroy fence.** A strategy with `operations` MUST refuse destruction while undelivered evidence remains, throwing `DurableEvidenceUndeliveredError`; `destroyDurableBranch()` preserves that typed refusal instead of flattening it into a generic branch failure, so the caller can still recover the evidence. Deliver (or archive) the outstanding evidence before destroying the branch. Concurrent `operate` and `destroy` are serialized by the host's own transaction: either the operation commits first (destroy then observes undelivered evidence and refuses) or destroy commits first (the operation fails against the removed allocation). No partial state is ever observable. ##### Bundled PostgreSQL manager `createPostgresWorkingCopyManager` implements the capability when given an `operations` option. `apply` is how the host's opaque mutation reaches the graph; TypeGraph still never interprets `mutation`. ```typescript const copies = createPostgresWorkingCopyManager({ control, connect, operations: { graph, // Runs inside the transaction that commits the evidence row. A throw rolls // back both the mutation and the evidence. apply: async (transaction, mutation) => { await applyHostMutation(transaction, mutation); }, }, }); const outcome = unwrap( await operateDurableBranch(descriptor, copies.durable, request), ); ``` `operations.graph` is required because a capability member receives only the descriptor, so the manager must reopen the allocation from the graph the host names. Before any connection or transaction opens, every member checks that graph against the sealed allocation's attested origin: its graph id and its version-blind definition hash must equal the ones the branch was forked with, so a graph that reuses the id with a different definition is refused. `apply` receives the transaction-scoped context of the allocation's fixed-schema Store, the same context `store.transaction` provides, so the allocation's fixed schema applies. Without the option, `copies.durable.operations` is undefined and `operateDurableBranch()` returns `unsupported` (`atomicMutation`). Each durable allocation owns one evidence relation under its ledger-reserved physical prefix, created in the provisioning transaction and dropped by destroy. The ledger records whether an allocation has one (`operation_evidence`). `operate` takes the allocation lock on the allocation's own transaction session, attests the sealed origin, resolves idempotency, takes the graph write lock, computes the `before` coordinates, calls `apply`, computes the `after` coordinates once the transaction's revision bookkeeping has run, and inserts undelivered evidence, all in one transaction. The allocation lock is a transaction-scoped advisory lock keyed on the allocation id, in a namespace of its own so it can never collide with a graph's write lock. It serializes operations per allocation, so the evidence sequence that backs the opaque scan cursor is commit order, and each operation's `before` equals the previous operation's `after` whenever every writer to the allocation goes through `operate` or takes the graph write lock. Ordinary writes take that lock on an allocation that tracks history or revisions, so a direct write cannot commit between `before` and `apply`; on an allocation that tracks neither, a direct write is not fenced and the evidence's `before`/`after` pair may include it. The graph write lock is graph-wide. While `apply` runs, tracked writes to the source graph and to every sibling working copy of it wait on that lock, so keep `apply` short and do not wait on other graph writers inside it. Coordinates always carry `base`. They also carry `revision`, the engine revision, when the allocation resolves lineage, which is when it tracks history or revisions; both are read on the transaction's own session so they describe one state. An allocation that tracks neither reports no `revision`, and its `base` values are content fingerprints, which read the whole graph twice per operation. **Isolation is observed, not assumed.** `operate`, `markDelivered`, and destroy each request READ COMMITTED, and the statement that takes the allocation lock also reports the isolation level its session actually runs at. Any other level is refused before anything is read or written, with a `ConfigurationError` whose `details.code` is `WORKING_COPY_ISOLATION_UNSUPPORTED`, because the request is honored only where a backend supports it and a role or server default of REPEATABLE READ would otherwise give the fence and the idempotency lookup a snapshot older than the lock wait. A `control` or `connect` wrapper must therefore forward the transaction `isolationLevel` option. The refusal only fires when a wrapper drops the requested option and the session's default is not READ COMMITTED. The same check runs everywhere the manager drops an allocation, not only in destroy and `abortAllocation`: closing an ephemeral working-copy store, closing a `makeBackend` backend, and the cleanup after a failed allocation. The first two surface the refusal from `close()`. The cleanup swallows it so the allocation's original failure reaches the caller, which leaves the allocation behind. Every such orphan is discoverable with `listUnsealedAllocations` and is removed by `abortAllocation` once `control` forwards the option. **Destroy fence.** Destroy (and `abortAllocation`) takes the same allocation lock. `destroyDurableBranch()` refuses with `DurableEvidenceUndeliveredError` while undelivered evidence exists, even from a manager built without `operations`; delivering the evidence requires a manager built with `operations`. An in-flight `operate` and a destroy on one allocation serialize on the lock: whichever commits first decides the other's outcome. The destroy waits at most `cleanupLockTimeoutMs` (5000 ms by default); one that outwaits a long `apply` fails with the database's lock timeout having committed nothing, and can be retried after the operation settles. `get`, `scan`, and `hasUndelivered` take no allocation lock, so they never wait behind an `apply`. A destroy that commits after any member has attested the sealed row but before that member holds the allocation (before its connection is attested, before `operate` mints the revision origin, or before `get`, `scan`, or `hasUndelivered` reads the evidence relation) fails the member with one `BranchError` (`changed owner or was destroyed during the operation`), the same error `operate` and `markDelivered` raise against a removed allocation. It is never a raw missing-relation error or the "connection is not bound to the allocation database" refusal, which is reserved for a connection that reaches a different database than the one the ledger names. The fence follows the manager's [removal rule](#one-schema-per-allocation). It reads the evidence relation only in the allocation's own schema. Evidence relations that sit in another schema refuse removal before the fence runs and are kept. The fence runs only when the evidence relation is among the relations removal drops. An allocation whose evidence relation is gone has no evidence left to deliver and nothing to recover, so destroy removes its remaining relations and its ledger row, exactly as it does for an allocation whose relations exist nowhere. **Mixed-version deployments.** Only managers on this version take the allocation lock and honor the destroy fence. A manager from an earlier release that shares the ledger destroys an allocation without consulting its evidence, so undelivered evidence is lost with the allocation, and it does not drop the evidence relation, so a later `allocate` with the same id refuses because `op_evidence` exists without a ledger row. Upgrade every process that shares a working-copy ledger before any of them creates or destroys a durable allocation. To recover an orphaned evidence relation, read its undelivered rows (`WHERE NOT delivered`) and deliver them, then drop the relation the refusal names and retry. TypeGraph never drops it for you, because it may hold the only copy of undelivered evidence. An allocation provisioned by an earlier release has no evidence relation, and its ledger row says so without any statement that changes the database. `operate` returns `unsupported` with `dimensions: ["evidenceStore"]`. The only statement it runs is one read-only ledger `SELECT` through `control`; it runs no DDL, takes no lock, calls no `connect`, and applies and writes nothing. The read members report no evidence: `get` and `markDelivered` return `undefined`, `scan` returns an empty page (echoing `after`), and `hasUndelivered` returns `false`. Re-fork the branch to gain evidence. ### Constraint-aware ingestion branches For a bounded candidate batch, `planCandidateWriteSet()` hides the transient branch lifecycle completely. It accepts a validated, versioned JSON document, stages it through the same constraint-aware ingestion implementation, delegates to incremental merge planning, and closes the working copy on every outcome. The result is the ordinary `MergePlanArtifact`, so review and application use the same APIs as every other merge plan. On eligible revision-tracked graphs, planning seeds existing candidate rows, edge endpoints, cardinality peers, live same-id ontology peers, and any reachable current identity component into the disposable working copy. The resolver still queries the live target for declared unique and index peers, and the plan retains its ordinary provenance, conflicts, digest, and commit-time fences. Existing undeclared target properties survive staging; extra candidate properties are refused. A custom backend without the active-only source read uses the complete clone path for `oneActive` graphs. A custom backend without the keyed match-identity owner read, or a candidate whose owner is excluded from the clone projection, also uses that path. Other ineligible graphs use the complete clone path so staging still checks constraints that can depend on rows beyond the candidate's ids. ```typescript import { captureCandidateWriteSetTarget, planCandidateWriteSet, unwrap, } from "@nicia-ai/typegraph/graph-merge"; const writeSet = { formatVersion: 1, sourceId: "provider-a", target: await captureCandidateWriteSetTarget(store), nodes: [ { kind: "Patient", id: "provider-a:123", properties: { name: "Ana", mrn: "123" }, validFrom: "2026-01-01T00:00:00.000Z", }, ], edges: [], } as const; const plan = unwrap( await planCandidateWriteSet({ target: store, makeBackend, writeSet: JSON.parse(JSON.stringify(writeSet)), options: { resolve: { Patient: { blockIndex: "patient_mrn_candidates", similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.9, }, }, }, }), ); ``` `sourceId` is the stable attribution carried into conflicts, resolutions, and provenance; node and edge ids remain the contribution source ids. The target schema identity prevents a document authored against one graph contract from being staged against another. `validFrom` is required (and may be `null`) so replaying identical JSON cannot acquire a new import-time timestamp and change the plan digest. This adapter applies TypeGraph's existing entity/property merge semantics. Two distinct records that both validate do not conflict merely because an application interprets their subject, predicate, time, source, or value fields as disagreement. Domain-specific acceptance and Statement semantics remain in the consuming application. Use `ingestionBranch()` when an untrusted ingestion batch may contain aliases that deliberately repeat a canonical node's unique key. An ordinary `branch()` keeps the complete graph schema and rejects the duplicate during staging, before entity resolution can review and collapse it. An ingestion branch materializes an honest working-copy schema with only node uniqueness deferred; schema validation, edge endpoint checks, disjointness, and edge cardinality still apply immediately. ```typescript import { asNodeId } from "@nicia-ai/typegraph"; import { applyMergePlan, asBranchId, ingestionBranch, planMergeIncremental, unwrap, } from "@nicia-ai/typegraph/graph-merge"; import { importGraph } from "@nicia-ai/typegraph/interchange"; const incoming = unwrap( await ingestionBranch(base, makeBackend, { id: asBranchId("provider-a"), }), ); const imported = await importGraph(incoming, providerDocument, { onConflict: "error", onUnknownProperty: "error", }); if (!imported.success) throw new Error("Provider import was rejected"); const alias = await incoming.nodes.Patient.getById( asNodeId("incoming-patient"), ); if (alias === undefined) throw new Error("Imported patient was not found"); // `canonicalPatient` is an existing Patient read from the base before forking. // The repeated MRN and its identity evidence can be staged together. await incoming.identity.assertSame(canonicalPatient, alias); const plan = unwrap( await planMergeIncremental({ forkPoint: base, target: base, branches: [incoming], options: { onBasePropertyConflict: "flag", resolve: { Patient: { blockIndex: "patient_mrn_candidates", similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.9, }, }, }, }), ); const applied = unwrap(await applyMergePlan(base, plan)); await incoming.close(); ``` Both `importGraph()` and `importGraphStream()` accept the returned handle, so an interchange document can be staged without a hand-written collection copy loop. Import remains the single owner of node-first ordering, validity windows, edge endpoint order and reference validation. On an identity-enabled graph, the handle also exposes an assertion-only `IdentityAssertionWriteFacade` as `identity`: `assertSame`, `assertDifferent`, `bulkAssertSame`, and `bulkAssertDifferent`. This lets a batch stage aliases that repeat unique keys and the explicit identity evidence needed to reconcile them before merge-time constraint validation. Assertion contradictions and invalid endpoints are still refused while staging; only node uniqueness is deferred. The returned handle exposes those ingestion collections and identity assertion writes, not the branch's underlying `Store`. Identity reads and retractions, schema operations, transactions, and runtime internals remain unavailable, so callers cannot bypass the deferred-constraint contract. As with `Store`, the `identity` property is absent at the type level when the graph does not enable Operational Identity. The original graph definition remains the merge contract: `applyMergePlan()` validates node uniqueness against the entire resolved write set in the target transaction. Valid key handoffs and swaps are accepted as one set. If reviewed resolution leaves two live owners of the same unique key, the merge returns `MergeConstraintConflictError` and commits no graph or provenance writes. The derived schema is persisted on the working-copy backend, so the relaxed contract is auditable and an explicit reattachment with an equivalent graph definition verifies the same constraint behavior. `ingestionBranch()` does not expose a general reopen/resume API. Deferral is not an in-memory flag and does not disable database constraints ad hoc. Ingestion branches require a backend with the batch uniqueness operations needed for atomic final validation. Unsupported backends are refused rather than falling back to sequential checks. ## Valid-time windows **A new row's window travels with the merge.** An explicitly open-left row stays open-left through snapshot and incremental merges, including edge repointing. Reviewable plans serialize that lower bound as `validFrom: null`; an omitted plan field means the write states no lower-bound change. JSON export/import and plan application preserve the distinction. A branch-authored node or edge window — including a deliberately ended one on a resurrection — is written as-is by the commit rather than reset to merge time. When the incremental target itself also created the surviving row, the target's committed window wins. **An inherited row's end-of-validity is merged.** Both `update(id, {}, { validTo })` and `update(id, {}, { clearValidTo: true })` on a branch are ordinary writes. The merge carries the set, move, or reopening to the target even when the row's properties are untouched: ```typescript await fork.store.nodes.Patient.update(asNodeId("pat-1"), {}, { validTo: "2030-06-01T00:00:00.000Z" }); const report = unwrap(await merge(base, [fork])); // base now holds pat-1 with valid_to = 2030-06-01, and: report.validityEnds; // [{ entity: "node", kind: "Patient", id: "pat-1", // validTo: "2030-06-01T00:00:00.000Z", claimedBy: ["worker-1"] }] await fork.store.nodes.Patient.update(asNodeId("pat-1"), {}, { clearValidTo: true }); // A merge now reopens pat-1 and reports: // [{ entity: "node", kind: "Patient", id: "pat-1", // clearValidTo: true, claimedBy: ["worker-1"] }] ``` An ending is treated as a **sibling of deletion**, not as a property, because it makes the same kind of statement: *this stopped being true*. That single choice explains the whole contract: | Situation | Outcome | | --------- | ------- | | One branch ends the row | That end is written — including a *later* end, which extends the window. | | One branch reopens the row | The end is cleared with `clearValidTo: true`. | | Several branches end it differently | No conflict. The **earliest** end wins, and `report.validityEnds` names every claiming branch. | | Sibling branches end and reopen it | The end wins as the stronger monotone claim; every claimant remains visible in `report.validityEnds`. | | The incremental target already ended it | The target's end stands. A branch never re-windows a row the target itself windowed, and the row is left out of the merge's writes entirely — but the discarded claims are still reported, as an entry carrying `precedence: "target"` and the target's own instant. | | One branch ends it, another deletes it | Deleted, with **no** `DeleteModifyConflict` — the stronger statement absorbs the weaker one. | | A branch re-states the end the target holds | No write at all — nothing is staged, so there is no version bump or history row even with `coalesceUnchangedUpserts` off. | | No branch touched the window | Untouched. A properties-only edit never passes a window, so the committed one stands. | The earliest-end rule is fixed, not a policy knob: it is commutative and associative, so the merge stays order-independent, and `onPropertyConflict` never sees a property your schema does not have. **The branch that authored the committed end is credited.** An ending is authored state, so its author is a contributor to that row in `report.provenance` and in the durable sidecar — even when moving the window is the only thing that branch changed. Credit follows the *committed* end: when several branches end a row differently, only the branches whose claim equals the written instant are credited, while `validityEnds[].claimedBy` still names every claimant, winning or not. An ending a deletion absorbed commits nothing, so it credits nobody, and neither does an entry marked `precedence: "target"` — the merge committed none of that end. **Every claim the merge observed is visible in `validityEnds`, applied or not.** An entry with no `precedence` is one the merge *decided*: `validTo` is the instant it wrote, or `clearValidTo: true` says it reopened the row. An entry with `precedence: "target"` is one it did **not** — the incremental target had already changed that end, so the entry describes the target's set or clear, `claimedBy` names the branch claims that were thrown away, and nothing was written or credited for the row. A row no branch claimed at all produces no entry, since there was nothing to discard. `validityEnds` reports claims about rows inherited from the fork point. If the fork point is empty, every branch row is branch-created and the array is always empty. A demo or topology that needs to exercise this report must seed the row before branching, then end that inherited row on one or more branches. Because an ending is not a modification, `onDeleteModifyConflict` never sees one: a row whose *only* change is its window loses to a concurrent deletion even under `"prefer-modify"`, since there is no modification to prefer. A row with a properties edit *and* an ending keeps the usual delete/modify behavior on the properties, and its ending rides along only if that modification survives. **What is still NOT merged, and why.** On a row that is live in both the base and the branch, `validTo` is the only window field a branch can author *and* the commit can apply. A row's lower bound is immutable outside resurrection — `validFrom` is written only when a soft-deleted row is brought back — so that lower-bound delta remains observable in a fork but unapplicable: | Observed delta | Reachable how | Merged? | | -------------- | ------------- | ------- | | `validTo` set or moved | `update(id, {}, { validTo })` | **Yes** | | `validTo` cleared back to open | `update(id, {}, { clearValidTo: true })` | **Yes** | | `validFrom` changed | soft-delete + resurrect inside the fork | No | Rather than silently ignore it, the merge reports the lower-bound change in `report.dropped` with reason `"window-not-applicable"`. Reconciling a value the commit would then drop is worse than not merging it: the report would claim a change that never happened. Delete+resurrect can also make an ended base row appear open because resurrection creates a fresh window. When `validFrom` changed, that open end is part of the same non-applicable resurrection artifact; it is not treated as a branch-authored `clearValidTo`, and an incremental target artifact does not outrank another branch's explicit end claim. Full interval reconciliation (intersecting `[validFrom, validTo]` across branches) is deliberately out of scope — it needs a write path that moves a live row's lower bound, which contradicts the temporal model, and it would silently discard a branch's extension. ## Forking one graph namespace `forkGraphNamespace(sourceStore, privateBackend, operationKey)` copies one history-enabled graph into an independently allocated PostgreSQL database. It copies the graph's committed schema, current rows, tombstones, recorded-time relations, revision clock and journal, identity relations, and TypeGraph materialization records. It checks a repeatable-read source snapshot against a pre-cut `base@V` token, compares every copied row before target commit, and returns `{ store, proof, abort }`. One source transaction holds that snapshot for the entire copy, from its first source read through the target copy and digest checks. The source can accept writes after the snapshot cut, while the long-lived snapshot remains open until copying finishes; `proof.sourceBase` identifies the copied cut. ```typescript import { forkGraphNamespace, prepareNamespaceForkTarget, } from "@nicia-ai/typegraph/graph-merge"; // Run with the schema owner role before the runtime fork. await prepareNamespaceForkTarget(sourceStore, privateBackend); const fork = await forkGraphNamespace(sourceStore, privateBackend, "restore-42"); // Owner role again: builds IVFFlat indexes over the copied rows. await fork.store.materializeIndexes(); const historical = await fork.store .asOfRecorded(receipt.recorded) .nodes.Item.getById(receipt.itemId); // Publish the private database through your own placement registry only after // checking the fork and any application-specific restore invariants. // Before publication, await fork.abort() to discard an unchanged copy. ``` The caller provisions and owns `privateBackend`. It may contain other graph namespaces, but it must contain no rows for the source graph. TypeGraph refuses a connection to the source database, including an aliased backend object. `prepareNamespaceForkTarget()` is the owner-side step, and the fork itself issues no DDL. It installs the retry ledger, creates the graph's per-field pgvector tables, and builds every index the source has materialized for the graph with the DDL the source used. It writes no graph rows and no materialization records, so it can run before the target is empty-checked, and running it again is harmless. Indexes whose build never completed on the source are neither built nor required. IVFFlat indexes are the exception: IVFFlat clusters the rows present when it is built, so building one on an empty table gives poor recall. They are not built by preparation and their materialization records are not copied; run `fork.store.materializeIndexes()` after the fork to build them over the copied rows. Every other index the fork carried is already recorded, so that call only builds the IVFFlat ones. An IVFFlat index left on the target by an aborted fork has no record, so the next fork's `materializeIndexes()` drops and rebuilds it over the new rows. The target stays private until the caller changes its own placement pointer; TypeGraph does not publish it. `abort()` atomically removes the copied graph and operation marker while preserving unrelated namespaces, and refuses if the target has changed. A retry with the same operation key returns the same proof after checking the target digest and base token; a different key cannot reuse the populated target. This first-party copy supports the bundled PostgreSQL table layout, bundled `pgvector` embedding storage, and default `tsvector` fulltext storage. Embeddings are copied, digested, and verified like every other graph relation, and `abort()` removes them. A graph with embedding fields forks only between backends with the same vector storage: pgvector on both sides, or `vector: false` on both, where embeddings live only in node properties. A vector-disabled source never wrote the vector tables a pgvector target would search, so that pair is refused. The fork refuses custom table mappings, custom vector or fulltext strategies, and contribution-owned tables it cannot copy and validate. The current copy buffers one relation at a time and inserts rows in bounded batches, so operators should size the private copy process for its largest graph relation. It does not use interchange, whose payload lacks recorded history and tombstones. ## Determinism Graph Merge is built to be reproducible, which is what lets you retry, cache, diff, and test a merge with confidence: - Candidate sets are sorted before clustering; clusters resolve by stable keys. - Conflict resolution consults only the captured `branchOrder` (or lexicographic branch id) — never wall-clock. - The committed graph and the normalized report are a pure function of the *unordered* branch set. Use `branchOrder` to make preference explicit wherever a policy needs ordering: ```typescript const branchOrder = [systemOfRecord.id, agentA.id, agentB.id]; const result = await merge(base, [agentB, systemOfRecord, agentA], { branchOrder, onPropertyConflict: "lastWriteWins", // systemOfRecord wins, regardless of input order }); ``` ## Errors Most entry points return a `Result`; the error arm is a typed `TypeGraphError` subclass you can branch on. `applyMergePlanInTransaction()` instead throws a typed `MergeError` so a caller-owned transaction callback cannot resolve and commit after a partially applied failure: | Error | When | | ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `BranchError` | `branch()` or `ingestionBranch()` could not materialize a working copy. | | `BaseVersionMismatchError` | A branch forked from a different `base@V` than the target now has (snapshot `merge()`). Also the typed replan error `mergeIncremental()`'s in-transaction guards raise, and the by-ID freshness check both commit modes run, when the target moved in the plan→commit window. | | `IdentityMergeConflictError` | Code `GRAPH_MERGE_IDENTITY_CONFLICT`. Thrown by both `merge()` and `mergeIncremental()` for identity contradictions, assertion-ID collisions, and retract/reassert races. See the [identity guide](/identity/#interchange-and-branch-merge). | | `MergeConstraintConflictError` | Code `GRAPH_MERGE_CONSTRAINT_CONFLICT`. The resolved plan would violate a deterministic store constraint, such as edge cardinality or node uniqueness. Its category is `constraint`, its `cause` is the original typed store error, and its details expose the original constraint fields. No graph or provenance writes commit. | | `InvalidMergeOptionsError` | Code `GRAPH_MERGE_INVALID_OPTIONS`. The supplied option combination is invalid, `mergeIncremental()` was given the snapshot-only `target` option instead of silently ignoring it, or `mergeIncremental()`'s `onBasePropertyConflict` is not `"flag"`. | | `SimilarityUnavailableError` | A `vector`/`hybrid` strategy was requested with no `embedder`. | | `MergeConflictError` | A conflict could not be resolved under the configured policy. | | `MergePlanCapabilityError` | Public planning was requested for a target without the durable revision guarantee needed across processes and time. Enable `revisionTracking` or `history`. | | `MergePlanningStaleError` | The target moved while planning was reading it. This is an expected retry-and-replan outcome under concurrent writers: no plan was returned, so recapture the target and create a new plan before retrying. | | `StaleMergePlanError` | The target revision changed after planning, or this plan was already applied. Review a newly-created plan. | | `InvalidMergePlanError` | The input is not a valid plan artifact. More specific subclasses distinguish unsupported versions, digest changes, and target/schema/origin mismatches. | | `CandidateSourceError` | A built-in candidate source failed; details identify its source id, entity kind, and operation. | | `CandidateWriteSetError` | Code `GRAPH_MERGE_CANDIDATE_WRITE_SET`. Candidate JSON is malformed, targets another graph schema, cannot be staged, or violates the active graph contract. The accepted graph is unchanged. | | `MergeReviewError` | Code `GRAPH_MERGE_REVIEW`. Durable review evidence is malformed, unsupported, incomplete, or inconsistent, or review options cannot be represented safely. | | `DurableOperationError` | Code `GRAPH_MERGE_OPERATION`. System-category failure while calling a durable-operation host, including transport and strategy failures. | | `DurableOperationRequestError` | Code `GRAPH_MERGE_OPERATION_REQUEST`. User-category refusal for an invalid durable-operation request, descriptor, or scan option. | | `DurableOperationConflictError` | Code `GRAPH_MERGE_OPERATION_CONFLICT`. Constraint-category refusal when an idempotency key is reused with a different operation digest. The previously committed operation is returned untouched; nothing new is written. | | `DurableOperationUnsupportedError` | Code `GRAPH_MERGE_OPERATION_UNSUPPORTED`. The strategy's `operations` capability lacks a requested member; TypeGraph refuses rather than emulating the atomic guarantee. | | `DurableOperationEvidenceError` | Code `GRAPH_MERGE_OPERATION_EVIDENCE`. System-category failure because a host returned malformed or request-inconsistent operation evidence. | | `DurableEvidenceUndeliveredError` | Code `GRAPH_MERGE_OPERATION_UNDELIVERED`. `destroyDurableBranch()` was refused because committed operation evidence is still undelivered. Deliver or archive it first; the typed refusal is preserved so the evidence stays recoverable. | | `MatchEvidenceError` | Evidence could not be constructed safely, including a custom scorer returning `NaN` or infinity. | | `MergeError` | Any other merge failure (e.g. comparison-ceiling `"error"`, a non-transactional target). `MERGE_ERROR_CODES` enumerates the codes. | ## Example See [FHIR Graph Merge](/examples/fhir-graph-merge) for a complete runnable snapshot merge that reconciles two independently-extracted patient-care branches, and [Incremental Merge](/examples/incremental-merge) for live-target ingestion against an advancing base with persisted, queryable provenance. # Operational Identity > Assert, retract, query, and historize identity between graph nodes The TypeGraph Identity Profile records identity facts between **individual nodes**. It is deliberately smaller than OWL: `same` is symmetric and transitive, `different` is symmetric and class-lifted, and neither relation substitutes properties or automatically expands every graph query. ## Enable the profile Identity is graph-level and opt-in: ```typescript const graph = defineGraph({ id: "knowledge", nodes: { Person: { type: Person }, Author: { type: Author } }, edges: {}, identity: { sameIdAcrossKinds: "fold" }, }); ``` The option is serialized with the schema. Enabled graph types expose the full facade as `store.identity` and `tx.identity`, a read-only facade as `StoreView.identity`, and the identity traversal option. These surfaces use conditional **presence**: on an identity-disabled graph type, `identity` does not exist on `Store`, `TransactionContext`, or the read-only views at all — reaching for it is a compile error, not a `never`-typed property. A runtime `ConfigurationError` with details code `IDENTITY_NOT_ENABLED` backs those getters too, for widened or `any`-typed handles TypeScript can't check (a JavaScript caller, or a store handle that lost its precise graph type). Constraint-aware `IngestionBranch` handles follow the same conditional-presence contract, but expose only `assertSame`, `assertDifferent`, and their bulk forms. Reads and retractions stay unavailable so untrusted batches can stage identity evidence without gaining the full operational surface. See [Constraint-aware ingestion branches](/graph-merge/#constraint-aware-ingestion-branches) for the staging and merge workflow. At runtime, a disabled graph does no identity work: no identity locks, probes, closure computation, or identity SQL run. That guarantee is scoped to runtime behavior — a bundled backend still provisions the identity tables' schema (not work) when it bootstraps a fresh database, independent of whether the specific graph passed to `createStore`/`createStoreWithSchema` declares `identity`. `sameIdAcrossKinds: "fold"` preserves TypeGraph's structural ID rule: live nodes of different kinds with the same ID belong to one identity class. No assertion row is manufactured for that implicit membership. Use `sameIdAcrossKinds: "ignore"` to enable the assertion ledger without joining equal IDs across kinds; only explicit `same` assertions then join classes. ## Write and read identity ```typescript const alice = await store.nodes.Person.create( { name: "Alice" }, { id: "person-alice" }, ); const author = await store.nodes.Author.create( { penName: "A. Example" }, { id: "author-alice" }, ); const result = await store.identity.assertSame(alice, author); // result.action is "created" or "existing"; result.assertion is durable truth await store.identity.membersOf(alice); // [{ kind: "Author", id: "author-alice" }, // { kind: "Person", id: "person-alice" }] await store.identity.representativeOf(alice); await store.identity.nodesOf(alice); // hydrated, kind-discriminated nodes await store.identity.areSame(alice, author); await store.identity.assertionsOf(alice); await store.identity.explainSame(alice, author); const ended = await store.identity.retractAssertion(result.assertion.id); // ended?.validTo is the exact assertion end instant ``` The complete write surface is: - `assertSame(a, b)` and `assertDifferent(a, b)` - `bulkAssertSame(pairs)` and `bulkAssertDifferent(pairs)` - `retractAssertion(id)` - `retractSameAssertion(a, b)` and `retractDifferentAssertion(a, b)` - `bulkRetractAssertions(ids)` Bulk methods are eager and, on PostgreSQL, run under one graph identity lock (see [Operational notes](#operational-notes) — SQLite serializes through its single-writer lock instead). `bulkAssertSame` and `bulkAssertDifferent` preserve input order and return exactly one result per input pair. Reasserting a current semantic pair is idempotent; assertion results distinguish `action: "created"` from `action: "existing"`. Retraction methods return the ended assertion (or `undefined` for a missing current assertion). `bulkRetractAssertions` does **not** share that one-result-per-input shape: it dedupes the input ids and returns only the assertions that were actually open, in dense, first-occurrence input order — so the result does not align index-by-index with the input array. Self-assertions are rejected. Assertion IDs use the exported private-symbol-branded `IdentityAssertionId` type so unrelated strings cannot be passed accidentally. When you hold a plain assertion-ID string that came from persistence or an interchange document, re-enter the branded type with the `asIdentityAssertionId(value)` caster rather than a `as` assertion. Assertions may state an explicit half-open validity window. Scalar methods take the window as their third argument; bulk methods carry one window per pair: ```typescript await store.identity.assertSame(alice, legacyAlice, { validFrom: "2020-01-01T00:00:00.000Z", validTo: "2022-01-01T00:00:00.000Z", }); await store.identity.bulkAssertDifferent([ { a: alice, b: bob, validFrom: "2023-01-01T00:00:00.000Z" }, { a: alice, b: carol }, // ordinary current assertion semantics ]); ``` A past-ended window affects historical reads only. An open window beginning in the past affects both historical and current reads. Repeating the exact relation, pair, and window is idempotent. A second open window for an already current semantic pair is refused rather than silently collapsed onto a different `validFrom`. Empty objects retain the ordinary unwindowed semantics, including inside a mixed bulk call. Runtime-evolved nodes carry a nominal dynamic-node type, so they flow through the same identity surface without a cast: ```typescript const evolved = await store.evolve(extension); const person = await evolved.nodes.Person.create({ name: "Alice" }); const tag = await evolved .getNodeCollectionOrThrow("Tag") .create({ label: "author" }); await evolved.identity.assertSame(person, tag); await evolved.identity.membersOf(tag); ``` Reference reads return `IdentityNodeReference` values covering both compile-time graph kinds and registered runtime kinds. This widening is necessary even when a read starts from `person`, because its class can contain `tag`. Their IDs retain the appropriate nominal brand, and `nodesOf` hydrates the class into static kind-discriminated members or `DynamicNode` values for runtime members. A plain `{ kind: string, id: string }` does not prove that the kind came through the evolved Store; pass the dynamic node or a nominal dynamic reference returned by an identity read. Unknown and removed kinds still fail at runtime with `KindNotFoundError`. A missing, deleted, or coordinate-invisible input returns `undefined`, `[]`, or `false` according to the method. A visible singleton returns itself from `membersOf` and `representativeOf`, and `areSame(ref, ref)` is true. `areDifferent` lifts an explicit different assertion across both identity classes and also reflects ontology `disjointWith` constraints. Representatives are deterministic: the code-point-smallest `(kind, id)` visible member wins. `explainSame(a, b)` returns a shortest path of persisted `same` assertions and implicit same-ID folds connecting two visible references. Each step names its endpoints and either the assertion or `type: "same-id-fold"`. It returns `[]` for one visible reference and `undefined` when the references are distinct or not visible at the read coordinate. Use `store.asOf(instant).identity` for a historical explanation. Historical identity reads and identity-expanded traversals use the kinds registered on the current Store. Assertions involving a removed kind remain in recorded history but no longer connect active classes. `classes({ limit, kinds?, cursor? })` lists visible classes, including singletons, in representative order. A kind filter selects classes containing at least one visible member of the requested kinds; each result still includes all of that class's visible members. Pass `nextCursor` to the next call until it is absent. The cursor is exclusive and applies to the same graph, read coordinate, and kind filter. When `kinds` is omitted, the scan uses the registered runtime kinds present when each page is requested; adding a runtime kind during that scan changes the filter and invalidates its cursor. At current coordinates, the database finds visible representatives for the page and expands members only for those classes; discovering representatives still examines the visible node set. Historical coordinates reconstruct all visible classes before applying the page boundary. For paging across writes, use a recorded-time coordinate when recorded history is enabled: valid-time `asOf` reads still observe later changes to the live tables. ## Integrity and lifecycle Ordinary unwindowed assertions require live endpoints. Explicit windows require both endpoint rows to cover the assertion's whole half-open interval; an ended or late-starting endpoint raises `IdentityEndpointValidityError`. Future bounds and inverted windows raise `IdentityValidityWindowError`. Zero-width windows are accepted as empty history. Contradictions are checked throughout every overlapping segment, including transitive `same` paths; adjacent half-open windows do not overlap. `assertSame` fails when a current `different` assertion spans the two classes or when any member kinds are ontology-disjoint. `assertDifferent` fails when both endpoints are already in one class. These checks, folding, node deletion, import, schema-transition validation, and closure rebuild share one per-graph lock and one mutation coordinator. Soft-deleting a node ends its current assertions. Hard-deleting it removes every current and ended assertion touching the node from the live assertion ledger; when recorded history is enabled, earlier recorded coordinates remain queryable. On every graph, a `create()` or `upsertById()` for a soft-deleted same-`(kind, id)` row **resurrects** that row rather than erroring: its properties are replaced and its validity window is reset, so `validFrom` becomes the resurrection instant — unless the write carries an explicit window, which is honored as given (this is how merge preserves branch-authored windows). A resurrecting node write that supplies only a historical `validTo` takes the same **born-already-ended** exception a create takes: no lower bound is stored ("ended at T, start unknown") rather than a start after its own end, so the row reads back at every `asOf` before that end and `meta.validFrom` is `undefined`. One stated window reaches one stored shape whichever node path resets it — `create()` on a fresh id, `create()` on a tombstone, or a resurrecting `upsertById()`. (Edge resurrection instead keeps its stored lower bound, so `getOrCreateByEndpoints` can resurrect an edge directly into the ended state — but the end it names is held to that retained bound, so reviving an edge into a window that closed before the edge began is refused as a `ValidationError`, and means passing both bounds.) This graph-wide rule does not depend on the identity profile. Resurrection does not revive ended assertions, but folding runs again over the resurrected node when configured. Kind removal cascades assertion and closure rows for the removed kinds. Tightening ontology disjointness is rejected when it would make a persisted class contradictory. `rebuildIdentityClosure(store)` repairs the derived current closure from live nodes and current assertions. It validates integrity and never advances the content revision. Schema-managed rebuilds, including automatic startup repair of derived identity relations, pin the schema version used by the rebuild. If a concurrent migration advances that version first, repair refuses with `StaleVersionError` without overwriting the newer closure. Reopen using the current graph definition before retrying. ### The database-level backstop The checks above are code deciding whether a write is legal, and code can be wrong. Underneath them TypeGraph maintains a second derived relation — the **separation relation** — that holds one row per pair of identity classes a current `different` assertion keeps apart, keyed by the two class keys under a `CHECK (class_key_low < class_key_high)` constraint. Every transaction that fuses two identity classes relabels the affected separation rows in the same statement batch. Fusing two classes that were separated relabels both sides of their shared row to one key, the constraint rejects it, and the transaction aborts — in the engine, with no application code in the way. A write that reaches the ledger through a path that skipped identity validation therefore still cannot commit a contradictory graph; it fails with an `IdentitySeparationViolationError` naming the `different` assertion it contradicts. Nothing about the identity API changes. The relation is derived and maintained wherever the closure is, `rebuildIdentityClosure(store)` recomputes it from the ledger, and store-open validation checks it against that recomputation the same way it checks the closure. ## Temporal identity Integrity is **structural**; reads are **coordinate-visible**. Current reads use a materialized closure and then filter members through the same visibility predicate ordinary node reads use. `store.identity` and `store.asOf(now).identity` therefore agree. Non-current valid-time and recorded-time views reconstruct one fixed point over both explicit `same` assertions and same-ID folding edges. A structurally existing but coordinate-invisible bridge can conduct identity without being returned as a member. Recorded assertions are captured in the same commit as the truth-bearing write. The assertion's validity window and the commit that recorded it are independent coordinates. A retrospective assertion is therefore invisible before its recorded-time commit even when its valid-time window reaches farther into the past. Archival export includes the endpoint temporal bounds needed to validate those windows on import, and graph merge carries branch-authored bounded assertions without turning them into current truth. Identity profile and ontology rules are schema-level interpretation, not a third temporal dimension. Historical views apply the Store's pinned `sameIdAcrossKinds` mode and ontology to the assertions and nodes visible at the requested coordinate. Changing those schema rules can therefore reinterpret older coordinates; it does not rewrite the recorded assertion ledger. ```typescript const before = await store.recordedNow(); const historical = store.asOfRecorded(before!); await historical.identity.membersOf(alice); ``` ### Folds and time Implicit same-id folds (`sameIdAcrossKinds: "fold"`) conduct based on a node's **lifecycle** — whether it currently exists and is not soft-deleted — not its valid-time window. A node created today with a backdated `validFrom` is valid-time visible in the past (an ordinary node read at that past coordinate returns it), but it does not conduct a fold there: the fold only takes effect once the node actually exists. Symmetrically, a node with a future `validFrom` does not suppress its folds today — it already exists and is live, so it folds now even though it is not yet valid-time visible. Explicit `same` and `different` assertions are unaffected by this: they carry their own validity windows and conduct exactly when they are current. This keeps the fold computation tied to write events rather than to valid-time windows, so the materialized closure used by current reads and by `asOf(now)` reads is identical — a fixed-point reconstruction of "current" never needs to special-case valid-time skew on the folding edge itself. ## Identity-expanded traversal Traversal expansion is per hop and defaults off: ```typescript const results = await store .query() .from("Person", "person") .traverse("authored", "edge", { includeIdentityMembers: true }) .to("Document", "document") .select((ctx) => ({ edge: ctx.edge, document: ctx.document })) .execute(); ``` The hop considers coordinate-visible members of the source class, returns the physical edge and target rows, preserves their provenance, and deduplicates a physical edge within the step — with one legitimate exception: a self-inverse edge (`inverseOf(edgeKind, edgeKind)`) traversed with `expand` between two identity-folded peers can yield the same physical edge twice, once per direction/target it matches through the fold. That is not a dedup bug; the edge genuinely satisfies the traversal from both of its endpoints. Recursive traversal supports the same option. TypeGraph does not perform automatic graph-wide expansion and collection reads such as `getById` have no identity option. Both coordinates reach the candidate edge the same way — an ordinary indexed equality on the class member, never a membership test evaluated per candidate edge. How each one reaches the class differs, because what a class costs to compute differs. At the **current** coordinate the maintained closure already *is* the class relation, so each traversal step seeks into it from its own frontier rows: the frontier row's class through the closure's primary key, that class's members through the class index, each member's node for its visibility. Cost is proportional to the frontier and the size of its classes — never to how many identity classes the graph holds. Measured on SQLite with *n* Person nodes, each folded with a Company and an Alias peer sharing its id (a three-member class per source), all *n* acting as source rows and every edge leaving the Company peer: | source rows | fan-out | matching edges | before | after | | --- | --- | --- | --- | --- | | 250 | 1 | 250 | 67 ms | 6 ms | | 1000 | 1 | 1000 | 1077 ms | 9 ms | | 2000 | 1 | 2000 | 4616 ms | 19 ms | | 1000 | 8 | 8000 | 8611 ms | 13 ms | | 500 | 200 | 100,000 | 51,602 ms | 77 ms | Growth is linear in graph size where it used to quadruple per doubling: the hop no longer evaluates membership per candidate *(source row, edge)* pair. The number to plan around is the last row — a hundred thousand matching edges over a five-hundred-row frontier is where the old per-source rescan dominated. A **historical** hop — one under `asOf`, `asOfRecorded`, or a non-current `view()` — cannot use the materialized closure, because the closure represents only the present. Its rows come from a reconstruction of identity classes out of the assertion ledger, and under `sameIdAcrossKinds: "fold"` that reconstruction also has to consider the structural same-id relation, which is proportional to the number of live nodes in the graph. No frontier row narrows that fixed point, so it is built once per statement into a materialized relation every traversal step joins. Measured on the narrow-edge fixture that isolates the term (SQLite, *n* Person nodes each folded with a Company peer, all *n* acting as source rows, fan-out 1): | *n* | before | after | | --- | --- | --- | | 250 | 122 ms | 7 ms | | 500 | 486 ms | 7 ms | | 1000 | 1984 ms | 14 ms | | 2000 | 8261 ms | 28 ms | Growth is linear in graph size where it used to quadruple per doubling. The caveat that remains is the historical one, and it is worth planning around: a past-coordinate hop rebuilds the whole graph's classes even when you asked about one node, so its floor is a pass over the identity population regardless of how narrow the frontier is. A **current** hop has no such floor — a single-start-row hop over 50,000 folded triples measures 1 ms on SQLite against 387 ms when the class relation was still built graph-wide, and nine unrelated 501-member classes cost it nothing at all (0.5 ms on SQLite, 2.4 ms on PostgreSQL, against 564 ms and 568 ms). Pick the coordinate you actually need: reading the present is the cheaper question by a wide margin. ## Interchange and branch merge Interchange format `2.0` optionally carries an identity section. State export (the default) includes current assertions. Import into a populated target is target-oriented: an existing current semantic pair keeps its target assertion ID and `validFrom`. Working-copy branch cloning imports into an empty target and preserves source IDs and `validFrom` exactly. ```typescript const state = await exportGraph(store, { includeTemporal: true }); const archive = await exportGraph(store, { identityMode: "archival", includeDeleted: true, }); ``` Identity-enabled exports default `includeTemporal` to `true`, because importing identity truth must prove that both endpoints existed throughout each assertion window. Explicitly setting `includeTemporal: false` on an identity-enabled graph is refused. Archival mode also includes ended assertions. Those rows are restored after shape validation and do not affect current closure. An ending a node deletion caused carries that node as `endedBy`, so a round-trip preserves why each assertion ended and not merely that it did; import rejects an `endedBy` on an open assertion, or one naming a node that is not an endpoint of the assertion it ends. Ended assertions can reference soft-deleted nodes, and by default (`includeDeleted: false`) export joins every assertion against its endpoints' live rows — an assertion with a soft-deleted endpoint is silently **dropped from the export entirely**, not carried with a dangling reference. Pair `identityMode: "archival"` with `includeDeleted: true` to keep those assertions in the archive. Interchange documents carry no `deletedAt` field, so a node exported only because of `includeDeleted: true` re-imports as **live** — an `includeDeleted` archive resurrects its soft-deleted nodes on import rather than restoring them as deleted. Weigh that trade-off deliberately for a backup: without `includeDeleted`, soft-deleted endpoints and the assertions that reference them are silently absent; with it, those nodes come back alive. Recorded side tables are not part of interchange. Graph merge includes identity truth in staleness fingerprints and diffs. Duplicate current assertions use the earliest `validFrom`, then the code-point-smallest assertion ID — unless one candidate is already committed on the target with the exact staged truth, which always wins: the applier is idempotent per semantic pair, so a challenger could never actually be written. A node deletion cascades into ending the assertions touching it, at the node's own deletion instant, and records the deleted node on every row it ends — so the diff reads which endings that deletion caused and stages each one with its cause, however close in time the branch's own retractions fell. When a delete/modify conflict resolution keeps the node, an ending is dropped along with the overruled deletion that caused it (reported as `identity:deletion-overruled`), while a retraction a branch made itself survives the deletion being overruled — including one the deleting branch made before deleting the node, even in the deletion's own millisecond. A hard delete removes the assertion rows outright, taking the recorded cause with them and leaving nothing to separate cause from intent, so those endings count as cascades. `merge()` detects identity conflicts at plan time and returns them as a typed `IdentityMergeConflictError` — direct opposing relations on one endpoint pair, transitive contradictions reached through a chain of `same` assertions no single branch wrote, retract/reassert races, and an assertion over a node another branch deleted. A branch that retracts a pair and also reasserts it itself (convergent, not racing) merges cleanly. This is mechanical truth propagation, not semantic entity reconciliation. Plan time is the early surface, not the only one: any identity refusal that still escapes to the applier inside the commit transaction is translated into the same typed `IdentityMergeConflictError`, with the original error preserved as its cause (identity environment and storage-corruption codes pass through untranslated — they are not statements about merge truth). See [`IdentityMergeConflictError`](/errors/#identitymergeconflicterror) for the exact `merge()` signature and how to catch it. ### Independent targets and assertion IDs `mergeIncremental()` accepts a target that has moved on from the branches' fork point, so a branch's assertion IDs can meet a ledger that assigned those IDs independently. Snapshot `merge()` still requires its target to match the branches' base@V exactly, but the same by-ID contract governs the divergence a branch can create within its own lineage (hard-delete/recreate replacement) and the plan→commit window. The contract is by ID, on complete truth: - **One assertion ID, one complete truth.** A planned assertion whose ID the target's ledger — ended rows included — already binds to a different complete truth (relation, endpoints, validity) refuses at plan time as `IdentityMergeConflictError`. An exact match is applied idempotently. - **Retractions carry the truth they retract.** A branch retraction ends the target's current row for its ID only when that row *is* the truth the branch retracted. When the target reuses the ID for different truth, the retraction is skipped and reported in `MergeReport.dropped` as `identity:retraction-target-mismatch` — the branch's own assertion is already absent from the target, and ending the target's unrelated row would delete truth the branch never saw. - **Truth replacement is a conflict, not a silent keep.** Within one lineage a branch can legally rebind an assertion ID by hard-deleting an endpoint (which physically removes the row) and importing the ID for different truth. The diff stages that replacement as a retraction plus a new assertion; because the target's ledger still holds the ID's prior truth in an ended row, the plan-time one-ID-one-truth check refuses it typed rather than silently keeping either side's truth. - **The commit re-verifies IDs.** Both commit modes re-read every planned assertion and retraction ID inside the commit transaction and refuse plan→commit drift as `BaseVersionMismatchError` — retrying recomputes the plan from current state. One deliberate exception: a planned retraction whose row another writer already ended is accepted as a no-op, not drift. `MergeReport.merged.identity` reports the rows the applier actually created and ended; idempotent skips are excluded. - **The commit proves the result, not the plan.** After its identity writes, and still inside the same transaction, a merge re-derives the identity classes it touched from the written state and refuses a contradiction there as `IdentityMergeConflictError` — so a plan validated against state that has since moved cannot leave a contradictory ledger behind. The whole merge rolls back; there is no partial commit. If the derived classes disagree with the materialized closure, the closure is rebuilt inside the same transaction and the check re-runs, which repairs a lagging closure atomically with a merge that is otherwise sound. ## Operational notes On PostgreSQL, every identity-affecting node write on an identity-enabled graph serializes on a per-graph advisory transaction lock: at most one writer per graph proceeds at a time. This is a correctness guarantee for the assertion ledger and closure, and it is also a throughput ceiling — concurrent writers to the same graph queue behind the lock. Writes to other graphs, and all reads, are unaffected. First-time enablement is heavier than steady state. It takes a `SHARE` lock on the shared nodes table, which briefly blocks writes for **every** graph in that database, and it loads the whole graph to build the initial identity closure. Plan enablement for a quiet window on large databases. `evolve()` on an identity-enabled graph re-runs the same closure rebuild, so schema evolution carries a comparable one-time cost proportional to graph size. Changing `sameIdAcrossKinds` is a **breaking** schema change — a `fold`↔`ignore` flip rewrites the materialized identity closure and changes every `areSame`/`membersOf`/`includeIdentityMembers` answer against existing data — so it requires the same explicit `migrateSchema()` opt-in as any other breaking change; it never auto-migrates silently. Identity-relevant ontology changes (`disjointWith`, `equivalentTo`/deprecated `sameAs`, or `subClassOf`) are likewise persisted semantic migrations, not a local runtime toggle. `createStoreWithSchema` and explicit `migrateSchema()` both rebuild and validate the closure atomically with the schema commit that carries the change. While the flip is unapplied, store construction refuses with `ConfigurationError` details code `IDENTITY_PROFILE_MIGRATION_PENDING` whenever the identity change is the only breaking one in the diff; a migration that also breaks other schema surfaces raises the generic `MigrationError` enumerating everything. First-time identity *enablement* (`autoMigrate: false` on a graph newly declaring `identity: { ... }`) is a safe, additive change, and `createStoreWithSchema` refuses to return a Store while it is pending with `ConfigurationError` details code `IDENTITY_ENABLEMENT_PENDING`. The very first schema commit of an identity-enabled graph is an enablement too: a legacy database populated through an unmanaged `createStore` gets the same atomic fold scan, contradiction validation, and closure build during initialization — an empty database just makes them cheap no-ops. ## Migrating from type-level factories The ontology factories `sameAs(A, B)` and `differentFrom(A, B)` are deprecated: they relate **types**, not individual rows, and `differentFrom` never enforced instance identity. To migrate: 1. Add `identity: { sameIdAcrossKinds: "fold" }` to the graph. 2. Open it with `createStoreWithSchema` so the capability is persisted and existing cross-kind same-ID groups are validated and materialized. 3. Replace type-level facts with `store.identity` assertions between concrete node references. 4. Use `equivalentTo` or `disjointWith` when the intended relation is genuinely between kinds. On PostgreSQL, first-time enablement waits for in-flight node writes before it builds the initial identity closure. Quiesce or restart any store instances that were opened with the identity-disabled schema before allowing writes to resume; stale instances do not participate in identity locking. Identity requires interactive atomic transactions. Bundled SQLite and PostgreSQL drivers support it; Cloudflare D1 and `drizzle-orm/neon-http` reject an enabled graph with `ConfigurationError` details code `IDENTITY_REQUIRES_ATOMIC_BACKEND`. Identity-disabled graphs continue to work on those drivers. Durable entity handles, identity-group IDs, semantic reconciliation, automatic OWL property substitution, and graph-wide identity expansion are reserved future capabilities and are not implied by this profile. # Integration Patterns > Strategies for integrating TypeGraph into your application architecture This guide covers common integration patterns for adding TypeGraph to existing applications, from simple setups to production deployment strategies. ## Direct Drizzle Integration (Shared Database) If you're already using Drizzle ORM, TypeGraph can share your existing database connection. TypeGraph tables coexist alongside your application tables. ```typescript import { drizzle } from "drizzle-orm/node-postgres"; import { Pool } from "pg"; import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { createStore } from "@nicia-ai/typegraph"; // Your existing Drizzle setup const pool = new Pool({ connectionString: process.env.DATABASE_URL }); const db = drizzle(pool); // Add TypeGraph tables to your existing database await pool.query(generatePostgresMigrationSQL()); // Create TypeGraph backend using the same connection const backend = createPostgresBackend(db); const store = createStore(graph, backend); // For pure TypeGraph operations, use store.transaction() await store.transaction(async (tx) => { const person = await tx.nodes.Person.create({ name: "Alice" }); const company = await tx.nodes.Company.create({ name: "Acme" }); await tx.edges.worksAt.create(person, company, { role: "Engineer" }); }); ``` ### Mixed Drizzle + TypeGraph Transactions When combining TypeGraph operations with direct Drizzle queries in the same atomic transaction, create a temporary backend from the Drizzle transaction: ```typescript await db.transaction(async (tx) => { // Direct Drizzle operations await tx.insert(auditLog).values({ action: "user_created" }); // TypeGraph operations in the same transaction const txBackend = createPostgresBackend(tx); const txStore = createStore(graph, txBackend); await txStore.nodes.Person.create({ name: "Alice" }); }); ``` This pattern is only needed when you must combine both in one atomic transaction. **When to use:** - You want a single database to manage - Your graph data relates to existing tables - You need cross-cutting transactions **Considerations:** - TypeGraph tables use the `typegraph_` prefix to avoid collisions - Run TypeGraph migrations alongside your application migrations - Connection pool is shared, so size accordingly ## Drizzle-Kit Managed Migrations (Recommended) If you use `drizzle-kit` to manage migrations, you can import TypeGraph's table definitions directly into your schema file. This lets drizzle-kit generate migrations for all tables—both yours and TypeGraph's—in one place. ### Setup **1. Import TypeGraph tables into your schema:** ```typescript // schema.ts import { sqliteTable, text, integer } from "drizzle-orm/sqlite-core"; // Import TypeGraph tables (these are standard Drizzle table definitions) export * from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; // Or for PostgreSQL: // export * from "@nicia-ai/typegraph/adapters/drizzle/postgres"; // Your application tables export const users = sqliteTable("users", { id: text("id").primaryKey(), name: text("name").notNull(), email: text("email").notNull(), }); ``` **2. Generate migrations normally:** ```bash npx drizzle-kit generate ``` Drizzle-kit will now see all tables—TypeGraph's and yours—and generate migrations for them. **3. Apply migrations:** ```bash npx drizzle-kit migrate # Or for Cloudflare D1: wrangler d1 migrations apply your-database ``` **4. Create the backend:** ```typescript import { drizzle } from "drizzle-orm/better-sqlite3"; import Database from "better-sqlite3"; import { createSqliteBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; import { createStore } from "@nicia-ai/typegraph"; const sqlite = new Database("app.db"); const db = drizzle(sqlite); // Use the same tables that drizzle-kit manages const backend = createSqliteBackend(db, { tables }); const store = createStore(graph, backend); ``` ### Custom Table Names To avoid conflicts or match your naming conventions, use the factory function: ```typescript // schema.ts import { createSqliteTables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; // Create tables with custom names export const typegraphTables = createSqliteTables({ nodes: "myapp_graph_nodes", edges: "myapp_graph_edges", uniques: "myapp_graph_uniques", schemaVersions: "myapp_graph_schema_versions", embeddings: "myapp_graph_embeddings", fulltext: "myapp_graph_fulltext", indexMaterializations: "myapp_graph_index_materializations", kindRemovals: "myapp_graph_kind_removals", reconciliationMarkers: "myapp_graph_reconciliation_markers", }); // Export individual tables for drizzle-kit export const { nodes: myappGraphNodes, edges: myappGraphEdges, uniques: myappGraphUniques, schemaVersions: myappGraphSchemaVersions, embeddings: myappGraphEmbeddings, indexMaterializations: myappGraphIndexMaterializations, kindRemovals: myappGraphKindRemovals, reconciliationMarkers: myappGraphReconciliationMarkers } = typegraphTables; // SQLite fulltext is an FTS5 virtual table — drizzle-kit can't model // virtual tables, so this name is exposed as a string. The backend // creates the FTS5 table on first store boot via a focused // `ensureFulltextTable()` ensure (idempotent CREATE VIRTUAL TABLE // IF NOT EXISTS), so drizzle-kit-managed setups work without an // extra manual step. export const myappGraphFulltextTableName = typegraphTables.fulltextTableName; ``` For PostgreSQL with the default `tsvectorStrategy`, the factory **does** return a typed Drizzle table — `tables.fulltext` — alongside the others, so drizzle-kit-managed setups pick up the fulltext table automatically: ```typescript // schema.ts (PostgreSQL) import { createPostgresTables } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; export const typegraphTables = createPostgresTables({ // …same names as above… }); export const { nodes: myappGraphNodes, edges: myappGraphEdges, // … fulltext: myappGraphFulltext, indexMaterializations: myappGraphIndexMaterializations, // … } = typegraphTables; ``` If you swap in an alternate Postgres fulltext strategy (pg_trgm, ParadeDB / pg_search, pgroonga), the typed `tsvector`-shaped table won't match what your strategy needs. Override `tables.fulltext` in your schema barrel with your strategy's own Drizzle table, or skip the typed export and rely on the backend's runtime `ensureFulltextTable()` ensure to bootstrap your strategy's DDL. Then pass the same tables to the backend: ```typescript import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; import { typegraphTables } from "./schema"; const backend = createSqliteBackend(db, { tables: typegraphTables }); ``` ### Adding TypeGraph Indexes The table factory functions also accept `indexes`, which drizzle-kit will include in migrations: ```ts // schema.ts import { createSqliteTables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; import { defineNodeIndex } from "@nicia-ai/typegraph/indexes"; import { Person } from "./graph"; const personEmail = defineNodeIndex(Person, { fields: ["email"] }); export const typegraphTables = createSqliteTables({}, { indexes: [personEmail] }); ``` For PostgreSQL, use `createPostgresTables` from `@nicia-ai/typegraph/adapters/drizzle/postgres`. See [Indexes](/performance/indexes) for covering fields, partial indexes, and profiler integration. Beyond accelerating queries, a declared index powers `store.nodes..bulkFindByIndex(indexName, items)` — a batched lookup that returns, for each incoming record, the live nodes sharing its index key. This is the primitive for **import reconciliation** and **dedup-candidate discovery**: probe an entire import batch against the graph in one query to decide create-vs-merge per record (the key may be non-unique, so each record yields its own candidate list). See the [batched index lookup reference](/performance/indexes#batched-index-lookup-bulkfindbyindex) and the runnable [`examples/17-bulk-find-by-index.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/17-bulk-find-by-index.ts). If you only need PostgreSQL adapter exports, import from `@nicia-ai/typegraph/adapters/drizzle/postgres`: ```typescript import { createPostgresBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; ``` ### PostgreSQL with pgvector For PostgreSQL with vector search, ensure the pgvector extension is enabled before running migrations: ```sql CREATE EXTENSION IF NOT EXISTS vector; ``` When multiple allocations share one PostgreSQL database, give each backend a stable namespace so its pgvector tables and indexes remain physically isolated: ```typescript import { createPgvectorStrategy } from "@nicia-ai/typegraph"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; const backend = createPostgresBackend(pool, { vector: createPgvectorStrategy("tenant-a"), }); ``` Keep the namespace stable for the lifetime of the allocation. The default `pgvectorStrategy` continues to use the existing `tg_vec` / `tg_vecidx` names. Then in your schema: ```typescript // schema.ts export * from "@nicia-ai/typegraph/adapters/drizzle/postgres"; export const users = pgTable("users", { ... }); ``` **When to use:** - You already use drizzle-kit for migrations - You want a single migration workflow for all tables - You need Cloudflare D1 or other platforms that require drizzle-kit migrations **Advantages over raw SQL migrations:** - Single source of truth for schema - Type-safe schema in TypeScript - Drizzle-kit handles migration diffs automatically - Works with all drizzle-kit supported platforms ## Separate Database Use a dedicated database when you want isolation between your application data and graph data. ```typescript import { Pool } from "pg"; import { drizzle } from "drizzle-orm/node-postgres"; import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; // Application database (your existing setup) const appPool = new Pool({ connectionString: process.env.APP_DATABASE_URL }); const appDb = drizzle(appPool); // Dedicated TypeGraph database const graphPool = new Pool({ connectionString: process.env.GRAPH_DATABASE_URL }); const graphDb = drizzle(graphPool); await graphPool.query(generatePostgresMigrationSQL()); const backend = createPostgresBackend(graphDb); const store = createStore(graph, backend); ``` **When to use:** - Your primary database doesn't support required features (e.g., pgvector) - You want independent scaling for graph operations - Compliance requires data separation - You're adding graph capabilities to a legacy system **Considerations:** - No cross-database transactions (use eventual consistency patterns) - Sync data between databases via application logic or events - Separate backup/restore procedures ## In-Memory (Ephemeral Graphs) Use in-memory SQLite for temporary graphs, caching, or computation. ```typescript import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; function createEphemeralStore(graph: GraphDef) { const { backend } = createLocalSqliteBackend(); return createStore(graph, backend); } // Use case: Build a temporary graph for computation async function computeRecommendations(userId: string): Promise { const tempStore = createEphemeralStore(recommendationGraph); // Load relevant data into temporary graph const userData = await fetchUserData(userId); await populateGraph(tempStore, userData); // Run graph algorithms const results = await tempStore .query() .from("User", "u") .whereNode("u", (u) => u.id.eq(userId)) .traverse("similar", "s") .to("Product", "p") .select((ctx) => ctx.p) .execute(); return results; } ``` **When to use:** - Temporary computation graphs - Request-scoped graph state - Graph-based caching with expiration - Isolated test fixtures **Considerations:** - Data lost on process termination - Memory usage scales with graph size - No persistence—rebuild on restart ## Hybrid Overlay (Graph on Existing Data) Add graph relationships on top of existing relational data without migrating your data model. Your existing tables remain the source of truth; TypeGraph stores only the relationships and graph-specific metadata. Use the `externalRef()` helper to create type-safe references to external tables: ```typescript import { createExternalRef, defineEdge, defineGraph, defineNode, embedding, externalRef } from "@nicia-ai/typegraph"; import { z } from "zod"; // Define nodes that reference your existing tables const User = defineNode("User", { schema: z.object({ // Type-safe reference to your existing users table source: externalRef("users"), // Denormalized fields for graph queries (optional) displayName: z.string().optional(), }), }); const Document = defineNode("Document", { schema: z.object({ source: externalRef("documents"), embedding: embedding(1536).optional(), }), }); // Graph-only relationships not in your relational schema const relatedTo = defineEdge("relatedTo", { schema: z.object({ relationship: z.enum(["cites", "extends", "contradicts"]), confidence: z.number().min(0).max(1), }), }); const authored = defineEdge("authored"); const graph = defineGraph({ id: "document_graph", nodes: { User, Document }, edges: { relatedTo: { type: relatedTo, from: [Document], to: [Document] }, authored: { type: authored, from: [User], to: [Document] }, }, }); ``` The `externalRef()` helper validates that references include both the table name and ID, catching errors at insert time: ```typescript // Valid: includes table and id await store.nodes.Document.create({ source: { table: "documents", id: "doc_123" }, }); // Error: wrong table name (caught by TypeScript and runtime validation) await store.nodes.Document.create({ source: { table: "users", id: "doc_123" }, // Type error! }); // Use createExternalRef() for a cleaner API const docRef = createExternalRef("documents"); await store.nodes.Document.create({ source: docRef("doc_456"), }); ``` **Syncing with external data:** ```typescript // Sync helper: Create or update graph node from app data async function syncDocument(store: Store, appDocument: AppDocument) { const existing = await store .query() .from("Document", "d") .whereNode("d", (d) => d.source.get("id").eq(appDocument.id)) .select((ctx) => ctx.d) .first(); if (existing) { await store.nodes.Document.update(existing.id, { embedding: await generateEmbedding(appDocument.content), }); return existing; } return store.nodes.Document.create({ source: { table: "documents", id: appDocument.id }, embedding: await generateEmbedding(appDocument.content), }); } // Query combining graph traversal with app data hydration async function findRelatedDocuments(documentId: string) { // Get graph relationships const related = await store .query() .from("Document", "d") .whereNode("d", (d) => d.source.get("id").eq(documentId)) .traverse("relatedTo", "r") .to("Document", "related") .select((ctx) => ({ source: ctx.related.source, relationship: ctx.r.relationship, confidence: ctx.r.confidence, })) .execute(); // Hydrate with full data from app database const externalIds = related.map((r) => r.source.id); const fullDocuments = await appDb.select().from(documents).where(inArray(documents.id, externalIds)); return related.map((r) => ({ ...r, document: fullDocuments.find((d) => d.id === r.source.id), })); } ``` **When to use:** - Adding graph capabilities to an existing application - Semantic search over existing content - Relationship discovery without schema changes - Gradual migration from relational to graph thinking **Considerations:** - Maintain sync between app data and graph nodes - Decide what to denormalize (tradeoff: query speed vs. sync complexity) - The `table` field in `externalRef` enables referencing multiple external sources ## Background Embedding Workers Decouple embedding generation from request handling using background jobs. ```typescript // job-queue.ts - Define the embedding job interface EmbeddingJob { nodeType: string; nodeId: string; content: string; } // worker.ts - Process embedding jobs import { createStore } from "@nicia-ai/typegraph"; async function processEmbeddingJob(job: EmbeddingJob) { const { nodeType, nodeId, content } = job; // Generate embedding (expensive operation) const embedding = await openai.embeddings.create({ model: "text-embedding-ada-002", input: content, }); // Update the node const collection = store.nodes[nodeType as keyof typeof store.nodes]; await collection.update(nodeId, { embedding: embedding.data[0].embedding, }); } // api-handler.ts - Enqueue jobs on create/update async function createDocument(data: DocumentInput) { // Create node without embedding (fast) const doc = await store.nodes.Document.create({ title: data.title, content: data.content, // embedding: undefined - will be populated by worker }); // Enqueue embedding job (non-blocking) await jobQueue.add("generate-embedding", { nodeType: "Document", nodeId: doc.id, content: data.content, }); return doc; } ``` **Batch processing for bulk imports:** ```typescript async function backfillEmbeddings(batchSize = 100) { let processed = 0; while (true) { // Find nodes missing embeddings const nodes = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.isNull()) .select((ctx) => ({ id: ctx.d.id, content: ctx.d.content, })) .limit(batchSize) .execute(); if (nodes.length === 0) break; // Batch embed const embeddings = await openai.embeddings.create({ model: "text-embedding-ada-002", input: nodes.map((n) => n.content), }); // Batch update await store.transaction(async (tx) => { for (const [i, node] of nodes.entries()) { await tx.nodes.Document.update(node.id, { embedding: embeddings.data[i].embedding, }); } }); processed += nodes.length; console.log(`Processed ${processed} documents`); } } ``` **When to use:** - Embedding generation is slow (100-500ms per call) - You want fast API response times - Bulk importing existing content - Retry logic for API failures **Considerations:** - Handle job failures and retries - Consider rate limits on embedding APIs - Queries on `embedding` should handle null values during population ## Testing For test setup patterns, seed data strategies, and profiler-based index coverage checks, see the dedicated [Testing](/testing) guide. ## Deployment Patterns ### Edge and Serverless Deploy TypeGraph at the edge using SQLite-compatible runtimes. > **Note:** Edge environments cannot use `@nicia-ai/typegraph/adapters/drizzle/sqlite/local` > because it depends on `better-sqlite3`, a native Node.js addon. Instead, use > `@nicia-ai/typegraph/adapters/drizzle/sqlite` which is driver-agnostic. **Cloudflare Durable Objects (SQLite) — transactional:** ```typescript import { drizzle } from "drizzle-orm/durable-sqlite"; import { createAdapterStoreWithSchema } from "@nicia-ai/typegraph"; import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; export class GraphObject { constructor(private ctx: DurableObjectState) {} async fetch(): Promise { const db = drizzle(this.ctx.storage); const backend = createSqliteBackend(db); // auto-detects "do-sqlite" const [store] = await createAdapterStoreWithSchema(graph, backend); // Atomic across TypeGraph + the caller's own relational tables: await store.transaction(async (tx) => { await tx.nodes.Document.update(documentId, props); if (tx.sqlAvailability !== "available") { throw new Error(`Native transaction unavailable: ${tx.sqlAvailability}`); } const sqlTx = tx.sql; await sqlTx.insert(documentVersions).values(versionRow); }); return new Response("ok"); } } ``` Unlike D1, Durable Objects expose an interactive storage transaction runner, so `store.transaction()` / `store.withTransaction()` are fully atomic (`capabilities.execution.interactiveTransactions: true`). See [Backend Setup](/backend-setup#cloudflare-durable-objects-sqlite) and the [Cross-Store Transactions recipe](/recipes#cross-store-transactions-drizzle--typegraph). **Cloudflare Workers with D1:** ```typescript // worker.ts import { drizzle } from "drizzle-orm/d1"; import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; export default { async fetch(request: Request, env: Env): Promise { const db = drizzle(env.DB); const backend = createSqliteBackend(db); const store = createStore(graph, backend); // Handle request with graph queries const results = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 5)) .select((ctx) => ctx.d) .execute(); return Response.json(results); }, }; ``` **Turso (libSQL):** ```typescript import { createClient } from "@libsql/client"; import { drizzle } from "drizzle-orm/libsql"; import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; const client = createClient({ url: process.env.TURSO_DATABASE_URL!, authToken: process.env.TURSO_AUTH_TOKEN, }); const db = drizzle(client); const backend = createSqliteBackend(db); const store = createStore(graph, backend); ``` > For Turso and D1, use [drizzle-kit managed migrations](#drizzle-kit-managed-migrations-recommended) > to set up the schema. **Bun with built-in SQLite:** Bun runs locally, so you can use the Node.js-compatible path with better-sqlite3, or use bun:sqlite with drizzle-kit managed migrations: ```typescript import { Database } from "bun:sqlite"; import { drizzle } from "drizzle-orm/bun-sqlite"; import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; const sqlite = new Database("app.db"); const db = drizzle(sqlite); const backend = createSqliteBackend(db); const store = createStore(graph, backend); ``` > Use [drizzle-kit managed migrations](#drizzle-kit-managed-migrations-recommended) > to set up the schema with bun:sqlite. **When to use:** - Low-latency requirements (data close to users) - Serverless functions with graph queries - Read-heavy workloads **Considerations:** - SQLite limitations (single-writer, no pgvector) - Cold start times include DB initialization - Vector search (cosine/L2): sqlite-vec on the local better-sqlite3 backend; libSQL's built-in vectors on the libSQL / Turso backend ### Per-Request Connections (Cache the Verified Store) Some serverless Postgres setups — Cloudflare Workers behind Hyperdrive, or any platform that pools connections for you — want a **fresh connection per request**. `createVerifiedAdapterStore` reconciles the committed schema and checks index materialization at open time (a few `SELECT`s), so re-opening a verified store on every request adds that cost to every graph-backed route. Verify **once per isolate**, then build a zero-query store per request from the cached reconciled schema: ```typescript import { createAdapterStore, createVerifiedAdapterStore, getCommittedSchemaVersion, type GraphBackend, type ReconciledSchema } from "@nicia-ai/typegraph"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; type Cached = { reconciled: ReconciledSchema; version: number | undefined; }; // Per-isolate cache, plus the in-flight reconciliation. Memoizing the promise // collapses concurrent cold (or stale) requests onto ONE verify instead of each // running its own — otherwise a burst of first requests reproduces the fan-out // stampede this pattern exists to avoid. let cached: Cached | undefined; let inFlight: Promise | undefined; function reconcileOnce(verifyBackend: GraphBackend): Promise { inFlight ??= (async () => { const [store, result] = await createVerifiedAdapterStore(graph, verifyBackend); cached = { reconciled: store.reconciledSchema, version: result.status === "unchanged" ? result.version : store.reconciledSchema.version, }; return cached; })().finally(() => { inFlight = undefined; }); return inFlight; } export default { async fetch(request: Request, env: Env): Promise { const backend = createPostgresBackend(newPoolForThisRequest(env)); // Cold start: concurrent first requests all await the same reconciliation. let snapshot = cached ?? (await reconcileOnce(backend)); // A one-row probe detects a schema commit from another isolate; a moved // version refreshes through the same single-flight path. const committed = await getCommittedSchemaVersion(backend, graph.id); if (committed !== snapshot.version) snapshot = await reconcileOnce(backend); // Zero database round-trips. Reads and writes still validate against // runtime-committed kinds carried by the reconciled snapshot. const store = createAdapterStore(graph, backend, { reconciled: snapshot.reconciled }); const results = await store .query() .from("Document", "d") .select((ctx) => ctx.d) .execute(); return Response.json(results); }, }; ``` `store.reconciledSchema` is an opaque snapshot of the reconciled graph (compile-time kinds folded with any runtime-committed kinds) plus the committed version it reflects. `createAdapterStore(graph, backend, { reconciled })` issues **no** queries and validates writes against that snapshot, so kinds committed at runtime remain writable without re-verifying. If you already hold a verified store and only need to swap the connection, `store.withBackend(freshBackend)` returns an equivalent store bound to the new connection with no re-verify. The `getCommittedSchemaVersion` probe is your read-your-writes seam: one round-trip, far cheaper than the full verified open (which also reconciles the schema and checks index materialization), and re-verify only fires when the version actually moves. Skip the probe only if your schema changes exclusively during a deployment that also clears the cache. Otherwise reads may use the stale schema snapshot, and the write fence rejects managed writes until the cache is refreshed. ### Read Replica Separation Route heavy graph queries to read replicas while writes go to primary. ```typescript import { Pool } from "pg"; import { drizzle } from "drizzle-orm/node-postgres"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; // Primary for writes const primaryPool = new Pool({ connectionString: process.env.PRIMARY_DATABASE_URL, max: 10, }); const primaryDb = drizzle(primaryPool); const primaryBackend = createPostgresBackend(primaryDb); const primaryStore = createStore(graph, primaryBackend); // Replica for reads const replicaPool = new Pool({ connectionString: process.env.REPLICA_DATABASE_URL, max: 50, // Higher pool for read-heavy workloads }); const replicaDb = drizzle(replicaPool); const replicaBackend = createPostgresBackend(replicaDb); const replicaStore = createStore(graph, replicaBackend); // Route based on operation export const stores = { write: primaryStore, read: replicaStore, }; // Usage async function searchDocuments(query: string) { // Read from replica return stores.read .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 10)) .select((ctx) => ctx.d) .execute(); } async function createDocument(data: DocumentInput) { // Write to primary return stores.write.nodes.Document.create(data); } ``` **When to use:** - Heavy read workloads (semantic search, graph traversals) - Write/read ratio is heavily skewed toward reads - Need to scale read capacity independently **Considerations:** - Replication lag means reads may be slightly stale - Don't use replica for read-after-write scenarios - Monitor replication lag in production ### Multi-Tenant Architecture Four approaches for multi-tenant deployments, each with different tradeoffs. #### Option 1: Shared tables with tenant isolation (simplest) ```typescript import { defineNode, defineGraph } from "@nicia-ai/typegraph"; // Include tenantId in your node schemas const Document = defineNode("Document", { schema: z.object({ tenantId: z.string(), title: z.string(), content: z.string(), }), }); // Always filter by tenant in queries function createTenantQuery(store: Store, tenantId: string) { return { searchDocuments: (query: string) => store .query() .from("Document", "d") .whereNode("d", (d) => d.tenantId.eq(tenantId).and(d.embedding.similarTo(queryEmbedding, 10))) .select((ctx) => ctx.d) .execute(), createDocument: (data: Omit) => store.nodes.Document.create({ ...data, tenantId }), }; } // Middleware extracts tenant and creates scoped API function withTenant(req: Request) { const tenantId = req.headers.get("x-tenant-id")!; return createTenantQuery(store, tenantId); } ``` #### Option 2: Separate `graph_id` per tenant (divergent schemas, one database) Every TypeGraph row is keyed by `graph_id`, and so is the committed schema. Two graphs with different `id`s coexist in one database with **independent schemas** — tenant A can declare kinds tenant B has never heard of, and neither sees the other's nodes, edges, or kind namespace. ```typescript function tenantGraph(tenantId: string) { return defineGraph({ id: `tenant_${tenantId}`, // the isolation boundary nodes: { Document: { type: Document } }, edges: {}, }); } // Each tenant commits — and evolves — its own schema, in the same database. const [store] = await createStoreWithSchema(tenantGraph("acme"), backend); ``` What `graph_id` isolates: - **Kinds** — the kind namespace is per `graph_id`. Declaring `Invoice` in one graph does not create it in another. - **Data** — nodes and edges are filtered by `graph_id` on every read and write. - **Schema version and evolution** — each graph owns its committed schema document and version, so tenants migrate independently. This is the cheap alternative to N physical databases when you want divergent per-tenant schemas without N connections. See [Graph identity and the kind namespace](/schemas-stores#graph-identity-and-the-kind-namespace). **Caveat 1 — index names are database-global.** Materialized SQL index names are derived from `(kind, fields, shape)` and are **not** namespaced by `graph_id`, because a SQL index name is a database-global identifier. Two graphs declaring the same kind name *and* the same index therefore resolve to one physical index: - **Same shape** — the second graph reuses the first graph's index. For a graph-scoped index this is safe: the index is keyed by `graph_id` (or `graph_id, kind`), so each graph still gets its own region of the index. - **Different shape** (say one `unique`, one not) — materialization fails loudly with a signature-drift error instead of silently sharing a mismatched index. Rename the declaration, or drop the existing index and retry. **Caveat 2 — a unique index must stay graph-scoped.** `scope` decides which TypeGraph system columns prefix the index key: | `scope` | Key prefix | Unique constraint applies | |---------|-----------|---------------------------| | `"graphAndKind"` (default) | `(graph_id, kind)` | per kind, per graph | | `"graph"` | `(graph_id)` | per graph | | `"none"` | *(none)* | **across every graph in the table** | A `unique` index declared with `scope: "none"` omits `graph_id` from the key, so the database enforces that value as unique across **all** graphs sharing the table — one tenant's row will block another tenant's insert. That is a real cross-tenant effect, and it holds whether or not two graphs share the physical index. Keep unique indexes on the default `"graphAndKind"` (or `"graph"`) scope in a multi-graph database; reserve `scope: "none"` for non-unique indexes where you deliberately want one index spanning every graph. Subject to those two rules, per-`graph_id` isolation holds: reads and writes stay filtered by `graph_id`, and the coupling is confined to physical index reuse. #### Option 3: Schema per tenant (PostgreSQL) ```typescript import { sql } from "drizzle-orm"; async function createTenantStore(tenantId: string) { const schemaName = `tenant_${tenantId}`; // Create schema if not exists await pool.query(`CREATE SCHEMA IF NOT EXISTS ${schemaName}`); // Run migrations in tenant schema await pool.query(`SET search_path TO ${schemaName}`); await pool.query(generatePostgresMigrationSQL()); await pool.query(`SET search_path TO public`); // Create Drizzle instance with schema const db = drizzle(pool, { schema: { schemaName } }); const backend = createPostgresBackend(db); return createStore(graph, backend); } // Cache tenant stores const tenantStores = new Map(); async function getTenantStore(tenantId: string): Promise { if (!tenantStores.has(tenantId)) { tenantStores.set(tenantId, await createTenantStore(tenantId)); } return tenantStores.get(tenantId)!; } ``` #### Option 4: Database per tenant (strongest isolation) ```typescript interface TenantConfig { id: string; databaseUrl: string; } async function createTenantStore(config: TenantConfig) { const pool = new Pool({ connectionString: config.databaseUrl }); await pool.query(generatePostgresMigrationSQL()); const db = drizzle(pool); const backend = createPostgresBackend(db); return { store: createStore(graph, backend), close: () => pool.end(), }; } // Connection manager with LRU eviction class TenantConnectionManager { private stores = new Map Promise }>(); private maxConnections = 100; async getStore(tenantId: string): Promise { if (!this.stores.has(tenantId)) { if (this.stores.size >= this.maxConnections) { await this.evictOldest(); } const config = await fetchTenantConfig(tenantId); this.stores.set(tenantId, await createTenantStore(config)); } return this.stores.get(tenantId)!.store; } private async evictOldest() { const [oldestId, oldest] = this.stores.entries().next().value; await oldest.close(); this.stores.delete(oldestId); } } ``` **Comparison:** | Approach | Isolation | Complexity | Scaling | Cost | | ------------------- | --------------- | ---------- | --------------------------- | ------- | | Shared tables | Low (row-level) | Low | Single DB | Lowest | | Schema per tenant | Medium | Medium | Single DB, separate schemas | Low | | Database per tenant | High | High | Independent DBs | Highest | **When to use each:** - **Shared tables**: SaaS with many small tenants, cost-sensitive - **Schema per tenant**: Moderate isolation needs, PostgreSQL only - **Database per tenant**: Enterprise customers requiring data isolation, compliance requirements ## Next Steps - [Quick Start](/getting-started) - Basic setup and first graph - [Semantic Search](/semantic-search) - Vector embeddings and similarity - [Performance](/performance/overview) - Optimization strategies # Graph Interchange > Import and export graph data for backups, migrations, and external integrations TypeGraph provides a standardized interchange format for importing and exporting graph data. Use it for: - Backing up and restoring graph data - Migrating data between environments - Exchanging data with external systems ## Quick Start ```typescript import { importGraph, exportGraph, GraphDataSchema } from "@nicia-ai/typegraph/interchange"; // Export your graph const backup = await exportGraph(store); // Import into another store const result = await importGraph(targetStore, backup, { onConflict: "update", onUnknownProperty: "strip", }); console.log(`Imported ${result.nodes.created} nodes, ${result.edges.created} edges`); ``` ## Interchange Format The interchange format is a JSON structure validated by Zod schemas. You can use `GraphDataSchema` to validate data before import, or export the schema as JSON Schema for API documentation. ```typescript import { GraphDataSchema } from "@nicia-ai/typegraph/interchange"; // Validate incoming data const validated = GraphDataSchema.parse(jsonData); // Export as JSON Schema for API docs import { toJSONSchema } from "zod"; const jsonSchema = toJSONSchema(GraphDataSchema); ``` ### Format Structure ```typescript interface GraphData { formatVersion: "2.0"; exportedAt: string; // ISO datetime source: { type: "typegraph-export" | "external"; // Additional source-specific fields }; nodes: Array<{ kind: string; id: string; properties: Record; validFrom?: string | null; validTo?: string; meta?: { version?: number; createdAt?: string; updatedAt?: string; }; }>; edges: Array<{ kind: string; id: string; from: { kind: string; id: string }; to: { kind: string; id: string }; properties: Record; validFrom?: string | null; validTo?: string; meta?: { createdAt?: string; updatedAt?: string; }; }>; identity?: { profile: "typegraph-identity-v1"; mode: "state" | "archival"; assertions: Array<{ id: string; relation: "same" | "different"; a: { kind: string; id: string }; b: { kind: string; id: string }; validFrom: string; validTo?: string; }>; }; } ``` `validFrom` has three states: the key **absent** means it wasn't requested (`includeTemporal: false`, the default) — import defaults it to the import's own creation timestamp, unless the record also states a `validTo` at or before that instant, in which case it is imported with no lower bound ("ended at T, start unknown") rather than one past its own end. An **explicit `null`** means the source row is confirmed to have no lower bound (open-left validity) — import preserves that instead of re-stamping it. A **string** is an explicit value, carried through unchanged. ### Format Version Compatibility Exports always write `formatVersion: "2.0"`. The read side — both `importGraph`/`importGraphStream` and `GraphDataSchema.parse` — additionally accepts `"1.0"`. A 1.0 document is structurally a valid 2.0 document: the only 2.0 change is the additive optional `identity` section, so pre-existing 1.0 exports validate and import unchanged. You never need to rewrite the version field of an older backup; validation and import handle both. ## Exporting Data Use `exportGraph` to serialize your graph data: ```typescript import { exportGraph } from "@nicia-ai/typegraph/interchange"; // Export everything const fullExport = await exportGraph(store); // Export specific node kinds const peopleOnly = await exportGraph(store, { nodeKinds: ["Person", "Organization"], }); // Export specific edge kinds const relationshipsOnly = await exportGraph(store, { edgeKinds: ["worksAt", "knows"], }); // Include metadata (version, timestamps) const withMeta = await exportGraph(store, { includeMeta: true, }); // Include temporal fields (validFrom, validTo) const withTemporal = await exportGraph(store, { includeTemporal: true, }); // Include soft-deleted records const withDeleted = await exportGraph(store, { includeDeleted: true, }); // Identity-enabled graphs export current assertions by default. // Include ended assertion history explicitly: const archival = await exportGraph(store, { identityMode: "archival", }); // A self-contained archive pairs archival identity with includeDeleted: const selfContainedArchive = await exportGraph(store, { identityMode: "archival", includeDeleted: true, }); ``` **Archival identity and soft-deleted endpoints:** `identityMode: "archival"` also exports *ended* assertions, and an ended assertion can reference an endpoint that was later soft-deleted. A default export (`includeDeleted: false`) joins every assertion against its endpoints' live rows, so an assertion touching a soft-deleted endpoint is silently **dropped from the export** — not carried with a dangling reference. This is silent archive loss, not a dangling-endpoint problem. When the archive must stand alone (backup, cold storage), pair it with `includeDeleted: true` so those assertions and their endpoints travel with it. That pairing has its own honest trade-off: the interchange format has no `deletedAt` field, so a node included only because of `includeDeleted: true` carries no record that it was deleted. Re-importing that archive resurrects the node as **live**. Choose deliberately: without `includeDeleted`, a backup silently loses soft-deleted endpoints and the assertions referencing them; with it, those nodes come back alive on restore. On import, every ended assertion's endpoints must exist as node rows in the target (soft-deleted rows qualify) — historical reads conduct identity through ended assertions, so an endpoint that never existed would become a phantom bridge joining real nodes at past coordinates. The store's own exports satisfy this by construction; a hand-built document that fails it is recorded as an `entityType: "identity"` entry in `result.errors`. ### Export Options | Option | Type | Default | Description | |--------|------|---------|-------------| | `nodeKinds` | `string[]` | all | Filter to specific node types | | `edgeKinds` | `string[]` | all | Filter to specific edge types | | `includeMeta` | `boolean` | `false` | Include version and timestamps | | `includeTemporal` | `boolean` | `false` | Include validFrom/validTo fields | | `includeDeleted` | `boolean` | `false` | Include soft-deleted records | | `identityMode` | `"state" \| "archival"` | `"state"` | Export current identity assertions, or current plus ended assertions | | `signal` | `AbortSignal` | none | Cancel the export: roll its snapshot transaction back and release the connection. See [Cancelling an export](#cancelling-an-export) | `exportGraphStream` also accepts `idleTimeoutMs`, a positive integer with no default. It bounds how long a delivered chunk may remain unacknowledged before the stream settles itself. The clock stops as soon as the consumer requests the next chunk, so a slow database read does not count as consumer idleness. **Round-trip caveat:** with the default `includeTemporal: false`, exported records carry no `validFrom`/`validTo`. On import, an omitted `validFrom` defaults to the *import's own* creation timestamp (a born-already-ended record is the exception noted above, and keeps no lower bound) — so a plain `exportGraph` + `importGraph` round trip does **not** reproduce the source's original valid-time window; every imported record becomes valid from import time forward. Pass `includeTemporal: true` on export when the clone needs to match the source's `asOf` behavior exactly (this is what `branch()` does internally). Identity-enabled graphs are the exception: their exports default temporal fields on, because assertion windows cannot be validated against endpoints without the endpoint bounds. Explicit `includeTemporal: false` is refused for those graphs. **Repair a legacy graph before exporting it.** A row an older library version stored with a backwards window (`valid_from > valid_to`) exports as it is stored and is then refused **per row** on re-import, because the import validates the stated pair — so an unrepaired graph does not round-trip. Run [`repairInvertedValidityWindows`](/schema-management#repairing-inverted-validity-windows) first; it normalizes those rows to the open-left shape import accepts. **`includeTemporal: true` with `onConflict: "update"`:** an update leg sends the document's `validTo` and never its `validFrom`, because a live row's lower bound is history. A document whose `validFrom` names a different instant than the target row holds is therefore stating a bound the import will not apply, and that row is reported as a per-row error carrying [`IMMUTABLE_VALIDITY_LOWER_BOUND`](/errors/#immutable_validity_lower_bound) rather than updated under a bound it ignored. This is reachable whenever a temporal export is replayed over rows that were created separately — the same document imported into a fresh graph creates those rows with their stated bounds and is unaffected. To update props over existing rows from a temporal export, either omit `validFrom` from the update document, export with `includeTemporal: false`, or import into a fresh graph and swap it in. ### Cancelling an export On a backend reporting `capabilities.execution.interactiveTransactions`, an export holds one repeatable-read snapshot transaction for its whole life, and on a single-connection backend it holds that connection's exclusive interchange-stream lease with it. (A backend without transactions — SQLite `transactionMode: "none"`, the session-less HTTP Postgres drivers — opens neither: its export paginates statement by statement, so a write committed mid-stream can appear in the pages that follow. That is a declared capability gap, not something the stream papers over.) Every *cooperative* exit gives both back, because each one runs the stream's `finally`: `break` or `throw` out of a `for await`, and an explicit `iterator.return()`. A consumer that pulls `next()` and then simply **drops the iterator** has no cooperative exit. Async-generator `finally` blocks do not run on garbage collection, so that snapshot transaction stays open for the life of the process — and on a serialized connection every later export and every later import is then refused for a stream nobody is reading. If you might abandon an iterator, pass a `signal` or configure an idle timeout: ```typescript const controller = new AbortController(); const iterator = exportGraphStream(store, { batchSize: 1000, signal: controller.signal, idleTimeoutMs: 30_000, })[Symbol.asyncIterator](); try { for (;;) { const next = await Promise.race([ iterator.next(), deadline(30_000), // resolves to a sentinel, leaving the pull in flight ]); if (next === TIMED_OUT) { // Do NOT just walk away: this is the leak. Aborting rolls the snapshot // back and frees the connection. controller.abort(new Error("export deadline exceeded")); break; } if (next.done === true) break; await write(next.value); } } finally { controller.abort(); } ``` Aborting rejects the pull that is in flight — and any later pull from a consumer that walked away and came back — with [`ExportStreamCancelledError`](/errors#exportstreamcancellederror), carrying the signal's own reason as `cause`, so a cancelled export is never mistaken for a complete one. The message states what was actually settled: a snapshot rolled back and a connection released on a transactional backend, or merely abandoned reads on one that never held either. Aborting a signal *before* the first pull refuses the export outright: no transaction is opened and no lease claimed. Aborting one that has already finished does nothing, so a single controller can safely span a whole job. `exportGraph` accepts `signal` too — there it simply makes the call reject instead of running to completion. When `idleTimeoutMs` expires, a later pull rejects with [`ExportStreamIdleTimeoutError`](/errors#exportstreamidletimeouterror), with code `"INTERCHANGE_EXPORT_STREAM_IDLE_TIMEOUT"`. Its `details.graphId` and `details.idleTimeoutMs` identify the stream and configured bound. The timeout is stream-only: `exportGraph` owns and promptly advances its internal consumer, so it accepts `signal` but not `idleTimeoutMs`. The option has no default because a stream may intentionally spend an unbounded amount of time processing a chunk; callers that cannot choose a safe idle bound should retain an `AbortController` and abort on their own job deadline instead. There is deliberately no garbage-collection fallback. A `FinalizationRegistry` cannot close this gap: any cleanup state able to settle an abandoned stream has to reach the stream's internals, and a registry holds its state strongly, so doing so would keep the abandoned stream reachable and the finalizer would never run. Explicit cancellation and the idle timeout are the mechanisms. ## Importing Data Use `importGraph` to load data into a store: ```typescript import { importGraph } from "@nicia-ai/typegraph/interchange"; const result = await importGraph(store, data, { onConflict: "update", onUnknownProperty: "strip", validateReferences: true, batchSize: 1000, }); if (result.success) { console.log(`Created: ${result.nodes.created} nodes, ${result.edges.created} edges`); console.log(`Updated: ${result.nodes.updated} nodes, ${result.edges.updated} edges`); console.log(`Skipped: ${result.nodes.skipped} nodes, ${result.edges.skipped} edges`); console.log(`Identity: ${result.identity.created} created, ${result.identity.skipped} skipped`); } else { console.error("Import had errors:", result.errors); } ``` ### Durable edge match identities during import The target graph declaration, not the interchange document, determines an edge's durable `matchIdentity`. After validating and normalizing an incoming edge's properties, normal import builds the same canonical endpoint/property key used by collection creates and writes it with the edge row. An import therefore cannot bypass convergence by choosing a new edge id. An incoming create whose durable identity is already owned by a different row, or was claimed by an earlier edge in the same import slice, is recorded as a per-edge error using `EdgeMatchIdentityConflictError`. This decision is separate from `onConflict`, which handles an existing row with the incoming edge's own id; `skip` and `update` do not authorize taking another row's durable identity. Bundled backends discover existing owners with set-oriented exact endpoint-pair reads and insert claimless durable slices with bind-budgeted, conflict-arbitrated batch statements. A slice that also needs cardinality claims remains atomic. If an exceptional batch refusal cannot identify the losing row, TypeGraph rolls the slice back to a savepoint and retries its rows individually so `result.edges.created` and `result.errors` remain honest. A custom or non-transactional backend that cannot prove that rollback refuses the import with `IMPORT_EDGE_BATCH_RETRY_REQUIRES_SAVEPOINT`; retrying a possibly committed prefix could otherwise double-count or misattribute rows. When the document carries an `identity` section, `result.identity` reports `{ created, skipped }` counts for imported assertions (skipped covers an exact re-import of an assertion that already exists under the same id). A rejected assertion — an unknown endpoint, a contradiction against the target's existing identity truth, or a reused assertion id that names different truth — is recorded in `result.errors` with `entityType: "identity"`, `kind` set to the assertion's relation (`"same"` or `"different"`), and `id` set to the assertion id, mirroring how node/edge errors carry `kind`/`id`. **Partial-commit caveat:** identity assertions are applied one at a time, and a mid-batch failure (a contradiction or id conflict partway through the `identity.assertions` array) stops the identity import but does not roll back the assertions already applied before it — they remain committed. They are **not** reflected in `result.identity.created`, since that count is only reported on success; the count under-reports rather than invents a number for committed-but-unaccounted work. The failure that stopped the batch is the one error entry you see in `result.errors`. ### Import Options | Option | Type | Default | Description | |--------|------|---------|-------------| | `onConflict` | `"skip" \| "update" \| "error"` | required | How to handle existing entities | | `onUnknownProperty` | `"error" \| "strip" \| "allow"` | `"error"` | How to handle extra properties | | `validateReferences` | `boolean` | `true` | Verify edge endpoints exist | | `batchSize` | `number` | `1000` | Batch size for database operations. Each batch pays fixed per-round-trip costs, so undersized batches slow client/server imports; inserts are still split by the driver bind budget internally. | ### Trusted initial import `trustedImportGraph` and `trustedImportGraphStream` are a separate, intentionally trusted path for loading a fresh dedicated database. They do not turn off validation on `importGraph`; they bypass the normal store write pipeline entirely. ```typescript import { trustedImportGraphStream, type GraphInterchangeChunk, } from "@nicia-ai/typegraph/interchange"; async function* chunks(): AsyncIterable { yield { type: "header", header }; for await (const nodes of readNodeBatches()) { yield { type: "nodes", nodes }; } for await (const edges of readEdgeBatches()) { yield { type: "edges", edges }; } } const result = await trustedImportGraphStream(store, chunks()); console.log(result); // { nodes: 1000000, edges: 5000000 } ``` The contract is deliberately narrow: - The TypeGraph node and edge tables must be globally empty. A different graph in the same database also makes the database non-empty. - The caller guarantees property shapes, endpoint existence, edge endpoint types, cardinality, duplicate-free IDs, and duplicate-free durable edge match identities. Only stream ordering and known kind names are checked. Trusted import still derives each declared edge identity and stores it with the row, so a collision reaches the database arbiter and rolls back the complete trusted-import transaction. - Recorded-time history, revision tracking, node uniqueness constraints, `searchable()` fields, and `embedding()` fields are rejected in this first version because their sidecar writes would otherwise be skipped. - Operational Identity-enabled target stores are rejected with `details.reason === "identity_unsupported"`; identity-bearing input is rejected with `details.reason === "invalid_stream"`. The trusted session writes only the node and edge relations, so it cannot persist assertions or materialize the derived closure — refusing both cases keeps identity truth from being silently dropped. Use `importGraphStream` for an export that carries Operational Identity assertions. This restriction does not apply to an edge registration's `matchIdentity`, which trusted import materializes in the edge relation itself. - Nodes must precede edges. The `meta` timestamps and node version in an interchange row are not restored; the import creates new storage metadata. - The complete stream is one transaction. Data insertion, temporary secondary index removal, index rebuilding, and planner statistics either all commit or all roll back. - A schema-managed Store acquires and validates its schema-write fence inside that transaction before loading rows. A stale managed import fails before row DML; a raw Store remains explicitly outside the schema-fencing guarantee. Supported native paths are synchronous prepared-statement SQLite (`better-sqlite3` and Bun SQLite) and transaction-capable PostgreSQL adapters with raw execution support (including node-postgres, postgres.js, and PGlite). Remote libSQL/Turso, D1, and HTTP-only PostgreSQL adapters reject the call with `TrustedImportError` and `details.reason === "backend_unsupported"`. Use `importGraph`/`importGraphStream` for external or uncertain data, conflict handling, incremental loads, and any graph with the unsupported features above. Use collection `bulkInsert` when the data is trusted but the database is not a fresh dedicated target. ### Conflict Strategies **`skip`** - Keep existing data, ignore incoming: ```typescript // Useful for incremental imports where you don't want to overwrite await importGraph(store, data, { onConflict: "skip" }); ``` **`update`** - Merge incoming data into existing: ```typescript // Useful for syncing updates from an external source await importGraph(store, data, { onConflict: "update" }); ``` **`error`** - Fail if any entity already exists: ```typescript // Useful for initial imports where duplicates indicate a problem await importGraph(store, data, { onConflict: "error" }); ``` #### An edge id held by a different edge Edge ids are unique per graph, but the import's existence probe (`getEdge` / `getEdges`) is keyed on `(graph_id, id)` alone. So a document edge whose id is already held by a row with a different **immutable identity** — its `kind` or either of its endpoints — finds that row. That question — *is this the same edge?* — is prior to *what do we do about the same edge?*, so it is answered **before** the conflict strategy and all three strategies answer alike: the row is reported as a per-row entry in `result.errors`, whose `error` message is prefixed `INTERCHANGE_EDGE_KIND_CONFLICT` and names each component that differs alongside the value the document stated. The stored row is left untouched. ```typescript const result = await importGraph(store, data, { onConflict: "update" }); const identityConflicts = result.errors.filter((entry) => entry.error.startsWith("INTERCHANGE_EDGE_KIND_CONFLICT"), ); ``` One prefix covers the whole class rather than a second one for endpoint mismatches: the condition is a single fact and the recovery is a single action, and a caller that had to match two prefixes to catch one condition would eventually match only one. The token still reads `…_KIND_CONFLICT` because it is the published, branchable string; it now covers every identity component. `ImportError` carries no `code` field, so the message prefix is the branchable token — the same `CODE: message` idiom the validity-window import refusals use. Give the incoming edge a distinct id, or import it under the identity the stored row already carries. Previously both non-`error` strategies were silent about this: `update` wrote the incoming edge's properties onto the *other* row with nothing in `result.errors`, and `skip` counted the document's edge as already present when no matching edge existed anywhere — so it was never created and never reported. Comparing `kind` alone closed only half of it: because endpoints are immutable, a document naming the incumbent's kind and id but different endpoints still read as the same edge, so `update` overwrote the incumbent's properties and silently retained its old endpoints. The update is additionally issued with all five identity components in the statement's own `WHERE`, so the check cannot be raced by a concurrent hard-delete-and-recreate; an update that consequently matches no row is reported as the same per-row error rather than aborting the import. Nodes were never affected: their probe is `getNode(graphId, kind, id)`, which is kind-scoped, so a cross-kind id collision simply reads as absent. #### An update target that changed under the import `onConflict: "update"` is a read-then-write pair: the import probes the stored row, validates the document's validity window against that row's `valid_from`, and then writes. Every part of that verdict is restated in the UPDATE's own `WHERE` — for edges the five identity components above, and for **both** nodes and edges the effective validity lower bound, whenever the window check actually read it. A concurrent hard-delete-and-recreate between the probe and the write therefore matches no row instead of landing a decision computed for a row that is gone (which would have ignored a `validFrom` the document stated, or persisted a `validTo` below the new row's `validFrom`). The bound is read — and so restated — when the document states a `validFrom` to compare against it, or a lone `validTo` to check for an inverted window. A document that states **neither** makes no claim about the row's window, so its properties update is not fenced on the bound and a concurrent recreate that only moved the bound does not refuse it. This matches `store.nodes.*.update` exactly: a write asserts what its decision read, and nothing more. A write that matches no row is reported per row, so an import whose earlier rows are already written is not aborted for it: ```typescript const result = await importGraph(store, data, { onConflict: "update" }); const raced = result.errors.filter( (entry) => entry.error.startsWith("INTERCHANGE_NODE_UPDATE_TARGET_CHANGED") || entry.error.startsWith("INTERCHANGE_EDGE_KIND_CONFLICT"), ); ``` `INTERCHANGE_NODE_UPDATE_TARGET_CHANGED` is the node-side prefix; edges reuse `INTERCHANGE_EDGE_KIND_CONFLICT`, whose message now also names the validity lower bound. Re-export the source and retry. A node update refused this way leaves no partial trace, and neither does one refused for a uniqueness conflict. The row write and the uniqueness transition are one unit: the new keys are claimed first (the claim is what decides the conflict), the row write follows, and the old keys are released only once it lands — with the claims given back if it does not. Fulltext and embedding sidecars are written only after the row update reports a match. So a row that `result.errors` reports is a row the import did not change, even though the transaction around it commits. ### Unknown Property Handling When importing data that has properties not defined in your schema: **`error`** - Reject the import (default, safest): ```typescript await importGraph(store, data, { onUnknownProperty: "error" }); // Throws if data has { name: "Alice", unknownField: "value" } ``` **`strip`** - Remove unknown properties silently: ```typescript await importGraph(store, data, { onUnknownProperty: "strip" }); // { name: "Alice", unknownField: "value" } becomes { name: "Alice" } ``` **`allow`** - Pass through to storage: ```typescript await importGraph(store, data, { onUnknownProperty: "allow" }); // Behavior depends on your database and schema strictness ``` ## Backup and Restore ### Creating Backups ```typescript import { exportGraph } from "@nicia-ai/typegraph/interchange"; import fs from "fs/promises"; async function createBackup(store: Store, backupDir: string) { const timestamp = new Date().toISOString().replace(/[:.]/g, "-"); const filename = `backup-${timestamp}.json`; const data = await exportGraph(store, { includeMeta: true, includeTemporal: true, }); await fs.writeFile( `${backupDir}/${filename}`, JSON.stringify(data, null, 2) ); return filename; } ``` ### Restoring from Backup ```typescript import { importGraph, GraphDataSchema } from "@nicia-ai/typegraph/interchange"; import fs from "fs/promises"; async function restoreBackup(store: Store, backupPath: string) { const json = await fs.readFile(backupPath, "utf-8"); const data = GraphDataSchema.parse(JSON.parse(json)); const result = await importGraph(store, data, { onConflict: "update", // or "error" for clean restore onUnknownProperty: "error", }); if (!result.success) { throw new Error(`Restore failed: ${result.errors.map(e => e.error).join(", ")}`); } return result; } ``` ## Migration Between Environments Move data from development to staging, or staging to production: ```typescript import { createStore } from "@nicia-ai/typegraph"; import { exportGraph, importGraph } from "@nicia-ai/typegraph/interchange"; import { graph } from "./schema"; async function migrateData( sourceBackend: GraphBackend, targetBackend: GraphBackend, ) { const sourceStore = createStore(graph, sourceBackend); const targetStore = createStore(graph, targetBackend); // Export from source const data = await exportGraph(sourceStore); // Import to target const result = await importGraph(targetStore, data, { onConflict: "error", // Ensure clean migration onUnknownProperty: "error", validateReferences: true, }); return result; } ``` ## Building Custom Import Pipelines For complex import scenarios, you can build pipelines using the Zod schemas: ```typescript import { GraphDataSchema, InterchangeNodeSchema, InterchangeEdgeSchema, type GraphData, } from "@nicia-ai/typegraph/interchange"; // Transform external data to interchange format function transformExternalData(externalRecords: ExternalRecord[]): GraphData { const nodes = externalRecords.map((record) => ({ kind: "Document", id: record.externalId, properties: { title: record.name, content: record.body, source: { system: "external", id: record.externalId }, }, })); // Validate each node const validatedNodes = nodes.map((node) => InterchangeNodeSchema.parse(node)); return { formatVersion: "2.0", exportedAt: new Date().toISOString(), source: { type: "external", description: "Imported from external CMS", }, nodes: validatedNodes, edges: [], }; } ``` ## Error Handling Import returns detailed error information for partial failures: ```typescript const result = await importGraph(store, data, { onConflict: "error" }); if (!result.success) { for (const error of result.errors) { console.error( `Failed to import ${error.entityType} ${error.kind}:${error.id}: ${error.error}` ); } // Decide how to handle partial import if (result.nodes.created > 0 || result.edges.created > 0) { console.log("Partial import completed, some entities were created"); } } ``` ### Serialized-connection refusals Row-level failures are reported in `result.errors`, but one class of failure is thrown instead: two long-lived interchange streams cannot share a single serialized database connection, so whichever starts second is refused with a typed `ConfigurationError`. The lease is exclusive — one stream of any kind per connection — and every long-lived import claims it, so `importGraph`, `importGraphStream`, `trustedImportGraph`, and `trustedImportGraphStream` can all throw it, as can `exportGraphStream` when an import already holds the connection — there, on the stream's first pull, since the claim begins when the snapshot transaction opens rather than when the iterable is constructed. The codes (`INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT`, `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`, `INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS`) and the `details.requested` / `details.heldBy` pairing they carry are documented in [Interchange serialized-connection guard codes](/errors#interchange-serialized-connection-guard-codes). Which connections count as serialized is read off the driver, so a driver TypeGraph cannot identify is left unmarked and a stream pair on it can still wedge. `createSqliteBackend` and `createPostgresBackend` take a `serializedResource` declaration for both directions of that gap — `{ mode: "shared", resource: client }` marks a connection TypeGraph cannot see, `{ mode: "independent" }` escapes a detection that is wrong for your topology. The escape hatch lifts the shared-resource refusal between two distinct backends only: one SQLite backend exporting into **itself** stays refused with `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`, because that is one handle holding the one snapshot transaction its own import needs. That surviving refusal is SQLite-only, so on PostgreSQL a backend declared independent may export into itself. See [Serialized connections](/backend-setup#serialized-connections). ## Best Practices ### Validate Before Import Always validate external data before importing: ```typescript import { GraphDataSchema } from "@nicia-ai/typegraph/interchange"; const result = GraphDataSchema.safeParse(untrustedData); if (!result.success) { console.error("Invalid data:", result.error.format()); return; } await importGraph(store, result.data, options); ``` ### Use Transactions for Consistency Import operations use transactions when the backend supports them. When the Store carries a reconciled schema version, every import batch acquires and validates the same schema-write fence as collection writes; a stale managed Store fails before row DML. A raw Store remains outside that guarantee. On a raw Store without transaction support, consider smaller batch sizes to minimize partial-failure impact; a managed Store on that backend fails closed on its first write. ### Test with `onConflict: "error"` First When setting up a new import pipeline, use `onConflict: "error"` to catch unexpected duplicates early: ```typescript // Development/testing await importGraph(store, data, { onConflict: "error" }); // Production (after validation) await importGraph(store, data, { onConflict: "update" }); ``` ### Monitor Import Results Log import statistics for observability: ```typescript const result = await importGraph(store, data, options); logger.info("Import completed", { success: result.success, nodesCreated: result.nodes.created, nodesUpdated: result.nodes.updated, nodesSkipped: result.nodes.skipped, edgesCreated: result.edges.created, edgesUpdated: result.edges.updated, edgesSkipped: result.edges.skipped, identityCreated: result.identity.created, identitySkipped: result.identity.skipped, errorCount: result.errors.length, }); ``` ## Next Steps - [Data Sync](/data-sync) - Patterns for keeping external data in sync - [Schema Migrations](/schema-management) - Managing schema changes over time - [Integration Patterns](/integration) - Database setup and deployment # Limitations > Known constraints and backend-specific limitations This page documents TypeGraph's known limitations and constraints. ## Backends Without Atomic Transactions Some runtimes cannot hold a multi-statement database session and therefore cannot offer atomic transactions: - **Cloudflare D1** — the D1 binding has no interactive transaction primitive (`D1Database.batch(...)` is transactional but batch-only). - **`drizzle-orm/neon-http`** — Neon's HTTP driver issues each statement as an independent request; there is no session to bind a transaction to. Cloudflare **Durable Objects** SQLite is *not* in this list: a store backed by `drizzle(ctx.storage)` is auto-detected as `transactionMode: "do-sqlite"`, reports `capabilities.execution.interactiveTransactions: true`, and is fully atomic. An `AdapterStore` created from that backend also exposes the adapter-only `store.withTransaction` and `tx.sql` surfaces. See [Backend Setup](/backend-setup#cloudflare-durable-objects-sqlite). These backends report `capabilities.execution.interactiveTransactions: false`. Read-only `store.batch(...)` still runs, but each query may use an independent connection and observe a different database snapshot. (Whether the queries nonetheless reuse one connection is up to the adapter — the no-transaction path hands each query the same backend object.) Note this is a difference of degree, not of kind: on PostgreSQL, `batch()`'s implicit transaction runs at the default read-committed isolation, so queries there can also observe interleaved commits. Write behavior depends on how the Store was constructed. A schema-managed Store fuses its schema fence into a write's own statement when the write fuses, and fails closed for writes that need the transaction-scoped schema or constraint fence otherwise — see [The guard every fused write shares](#the-guard-every-fused-write-shares) below for which writes fuse and which refuse. A raw `createStore()` / `createAdapterStore()` without a reconciled snapshot still has no interactive transaction boundary. `store.transaction(fn)` refuses with a typed capability error rather than pretending to provide rollback; direct backend writes remain raw. Eligible operations that use a certified atomic SQL program can still be available on these roots, but that transport guarantee is separate from the interactive transaction capability. These backends cannot honor the `isolationLevel` option on `store.transaction(...)`; the method refuses before invoking its callback, so the collection-read snapshot recipe documented elsewhere does not apply here. ```typescript // On a raw D1 / neon-http Store, this refuses before the callback runs. await store.transaction(async (tx) => { await tx.nodes.Person.create({ name: "Alice" }); }); ``` **If you require atomicity or schema-version fencing, branch on the capability:** ```typescript if (store.capabilities.execution.interactiveTransactions) { await store.transaction(async (tx) => { /* atomic */ }); } else { // Use independent operations, or a supported certified atomic operation. const person = await store.nodes.Person.create({ name: "Alice" }); const company = await store.nodes.Company.create({ name: "Acme" }); await store.edges.worksAt.create(person, company, { role: "Engineer" }); } ``` If you need atomic writes from an edge runtime, use `drizzle-orm/neon-serverless` (WebSocket-backed Pool) instead of `drizzle-orm/neon-http`. ### Four kinds of write atomicity TypeGraph distinguishes an interactive transaction, a static adapter batch, a certified atomic SQL program, and an authoritative one-statement command. `store.transaction(...)` is the interactive Store API: it pins a session and groups the callback's operations. A static batch is adapter-internal (such as D1 `batch()` or a multi-row insert); it is not a public Store transaction and cannot make arbitrary Store calls atomic. A certified atomic SQL program is a closed ordered statement sequence whose transport preserves result slots and parameters and rolls back primary and sidecar writes when a later statement fails. An authoritative command is a single `commands.execute` write whose database statement returns the decision it made. It can provide a safe transactionless create/found path only when the backend has a durable arbiter. Operational Identity, single-edge claim/cardinality enforcement, and any undeclared dynamic `matchOn` convergence that may write still require an interactive transaction and fail closed on a backend that cannot provide one. Outside the native durable-convergence envelope, an all-live `ifExists: "return"` endpoint batch is read-only and can return from its set-oriented root read without a transaction. Inside the native envelope, the authoritative upsert program runs before the Store knows every identity is live. It preserves the logical `"found"` result in one exchange, but may take incumbent-row locks and produce write amplification. Eligible direct edge batches on bundled roots are a separate exception: their closed native program carries the claim sidecars inside one atomic exchange. A declared edge `matchIdentity` persists a canonical endpoint/property key and has a unique database arbiter; eligible root `getOrCreateByEndpoints` calls can therefore use the authoritative one-statement command. The durable identity does not make unrelated Store operations, claims, or history/revision side effects transactionless. ### The guard every fused write shares Every static batch and every certified atomic program asserts the active schema version inside the very statement that writes, never as a preceding check — the fused create's `WHERE … is_active` predicate, or the program's leading `schema_fence` CTE. A stale version makes that statement match zero rows, so the write commits nothing, and the store re-reads and reports `StaleVersionError` instead of writing against a version that already moved on. This is what lets a `"batch"`-tier backend (`capabilities.execution.unitOfWork === "batch"` — Cloudflare D1's `batch()`, Neon HTTP's `transaction(queries)`, which fix every statement before the first one runs and commit them together with no session in between) run schema-managed creates, updates and deletes, and bulk writes at all: the fence travels inside the one exchange it can hold, instead of needing a session to hold it separately. A singleton node update, `upsertById`, or delete fuses the same way as a create, through a one-entry certified atomic program, whenever its kind carries no declared unique constraint — except a node delete, which fuses even when the kind DOES carry one, because the atomic delete program releases that claim in the same statement. A singleton edge update or delete fuses the same way (`EdgeCollection` has no `upsertById`). A write that needs more than that one guarded statement — because it must read a value it wrote earlier in the same write, hold an interactive callback open across round trips, maintain Operational Identity's closure, hold history's per-graph lock across a whole write cascade, or hold one transaction across a schema commit's compare-and-swap — refuses on a `"batch"`-tier backend with `BATCH_WRITE_UNSUPPORTED`, naming which of those it needed: | `reason` | What it needs | | --- | --- | | `interactive-callback` | Hold an interactive callback transaction open across several round trips (`store.transaction(fn)`). | | `constraint-needs-probe` | Read a value it wrote earlier in the same write before deciding what to write next (a declared constraint's probe-then-write). | | `identity` | Read and write Operational Identity's closure across several round trips inside one held transaction. | | `history` | Hold the per-graph write lock and clock open across a whole write cascade (`history: true` / `revisionTracking: true`). | | `schema-commit` | Hold one transaction across its compare-and-swap read and its activating write (`commitSchemaVersion` / `setActiveVersion`). | A write that simply cannot fuse — an ineligible write kind, a singleton create/update/`upsertById` on a kind with a declared unique constraint, a tombstone-resurrection write a supplied id falls through to, or a derived backend — refuses with `SCHEMA_WRITE_FENCE_UNSUPPORTED` instead and carries no `batchRefusal` reason: that gate has no proven need to name, only its own plain limitation. See [`BATCH_WRITE_UNSUPPORTED`](/errors#batch_write_unsupported) for where each reason surfaces in an error's `details`. ## libsql Single-Connection Transactions For local `@libsql/client` connections (`file:` paths and `file::memory:`), `createLibsqlBackend` frames transactions with raw `BEGIN IMMEDIATE`/`COMMIT` statements on the client's single stable connection. It deliberately avoids `client.transaction()`, which hands the client's connection to the transaction and lazily opens a new one afterwards — for an in-memory database that new connection is a fresh, empty database ([tursodatabase/libsql-client-ts#229](https://github.com/tursodatabase/libsql-client-ts/issues/229)). In-memory databases therefore work for all operations, including transactions. Remote Turso connections (`libsql://`, `http(s)://`) run each transaction on its own stream via the driver. The trade-off of a single connection: a store-level operation awaited from **inside** a `store.transaction` callback (on the root store, rather than the `tx` context) can never run — the open transaction occupies the backend's serialized execution slot until it completes — so the backend rejects it with a `ConfigurationError` instead of deadlocking. ```typescript // ✅ In-memory works, including transactions const client = createClient({ url: "file::memory:" }); // ❌ Root-store access inside a transaction callback throws await store.transaction(async (tx) => { await store.nodes.Person.find(); // ConfigurationError — use tx.nodes await tx.nodes.Person.find(); // ✅ transaction-scoped access }); ``` ## Recursive Traversal Depth Variable-length traversals use two depth caps and an explicit cycle policy: 1. Unbounded traversals (no `maxHops` option) are capped at 10 hops. 2. Explicit `maxHops` values are validated up to 1000 hops (`maxHops: >1000` throws). 3. Cycle prevention is on by default. To skip cycle checks for speed, opt into `cyclePolicy: "allow"` (which may revisit nodes across hops). This prevents runaway queries while still supporting deep, intentionally bounded traversals. ```typescript // Implicitly limited to 10 hops store .query() .from("Person", "p") .traverse("reportsTo", "e") .recursive() .to("Person", "manager"); // Explicit limits up to 1000 are honored store .query() .from("Person", "p") .traverse("reportsTo", "e") .recursive({ maxHops: 200 }) // honored .to("Person", "manager"); // Explicit limits above 1000 throw store .query() .from("Person", "p") .traverse("reportsTo", "e") .recursive({ maxHops: 2000 }) // throws .to("Person", "manager"); ``` The unbounded-traversal limit is defined as `MAX_RECURSIVE_DEPTH`: ```typescript import { MAX_RECURSIVE_DEPTH } from "@nicia-ai/typegraph"; // MAX_RECURSIVE_DEPTH = 10 ``` ## Connection Management Managed Store factories own their local SQLite or PGlite connection, and their `store.close()` method releases it. The local backend factories `createLocalSqliteBackend` and `createLocalPgliteBackend` likewise expose an owned backend whose `close()` releases its resources. Bring-your-own adapter factories leave connection ownership with you. For `createSqliteBackend`, `createPostgresBackend`, and `createLibsqlBackend`, you are responsible for: 1. **Creating and configuring** the database connection 2. **Implementing connection pooling** for production use 3. **Closing connections** when done ```typescript import Database from "better-sqlite3"; import { drizzle } from "drizzle-orm/better-sqlite3"; import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; // You manage the connection const sqlite = new Database("app.db"); sqlite.exec(generateSqliteMigrationSQL()); const db = drizzle(sqlite); const backend = createSqliteBackend(db); const store = createStore(graph, backend); // You close the connection sqlite.close(); ``` For production deployments, use connection pooling: ```typescript import { Pool } from "pg"; import { drizzle } from "drizzle-orm/node-postgres"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; const pool = new Pool({ connectionString: process.env.DATABASE_URL, max: 20, // Maximum connections }); const db = drizzle(pool); const backend = createPostgresBackend(db); ``` In the bring-your-own example above, `store.close()` leaves the supplied driver open. Close that driver or pool through its own API. ## Predicate Serialization Where predicates in unique constraints cannot be serialized. If you use schema serialization for versioning or migration, predicates are stored as `"[predicate]"` and cannot be reconstructed. ```typescript // This predicate works at runtime... unique({ name: "email_unique_when_active", fields: ["email"], where: (props) => props.status.isNotNull(), }); // ...but serializes as: // { "where": "[predicate]" } ``` **Workaround:** For full schema serialization support, avoid predicates in unique constraints. Use application-level validation instead. ## Vector Search Backend Requirements Vector and hybrid search work across all primary backends via a pluggable `VectorStrategy`. Each backend advertises its capabilities through `backend.capabilities.vector` (`{ supported, metrics, indexTypes, maxDimensions }`): | Backend | Requirement | Metrics | |---------|-------------|---------| | PostgreSQL | pgvector extension (HNSW / IVFFlat) | cosine, l2, inner_product | | SQLite | sqlite-vec extension (`vec0` KNN) | cosine, l2 | | libSQL / Turso | built-in native engine (DiskANN); nothing to load | cosine, l2 | | D1 | Not supported | — | Note that `inner_product` is PostgreSQL-only — sqlite-vec and libSQL support cosine and l2 only. Using vector predicates on unsupported backends throws `UnsupportedPredicateError`: ```typescript try { await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(queryVector, 10)) .select((ctx) => ctx.d) .execute(); } catch (error) { if (error instanceof UnsupportedPredicateError) { // Vector search not available on this backend } } ``` ## Query Builder Type Inference Complex query chains may occasionally require explicit type annotations when TypeScript cannot infer the full type. This is rare but can occur with deeply nested selects or unions. ```typescript // If type inference fails, add explicit type const results = await store .query() .from("Person", "p") .select((ctx) => ({ name: ctx.p.name as string, // Explicit annotation })) .execute(); ``` ## Bulk Operation Limits Bulk operations (`bulkCreate`, `bulkInsert`, `bulkUpsertById`, `bulkDelete`) have practical limits based on your database: | Database | Recommended Batch Size | |----------|----------------------| | SQLite | 500-1000 items | | PostgreSQL | 1000-5000 items | For larger datasets, batch your operations: ```typescript const BATCH_SIZE = 1000; for (let i = 0; i < items.length; i += BATCH_SIZE) { const batch = items.slice(i, i + BATCH_SIZE); await store.nodes.Person.bulkCreate(batch); } ``` ### Native node bulk eligibility Bundled PostgreSQL roots using a recognized session-capable driver, Neon HTTP, Cloudflare D1, and libSQL roots can use one schema-fenced native atomic program for schema-managed `nodes.bulkInsert()` and `nodes.bulkCreate()` calls when the node has no Operational Identity, history, or revision work. The program can compose fulltext/vector projections with the complete supported uniqueness and disjointness claim set for every member. Session-capable PostgreSQL executes that program on one pinned transaction; Neon HTTP, D1, and libSQL submit one transport batch. Advertised same-kind or hierarchy-wide uniqueness claims, disjointness claims, and mixed families are acquired in canonical order; compatibility reads preserve rows written under legacy claim axes. Claim-free members may participate alongside claimed members. IDs may be generated, caller-supplied, or mixed, and `bulkCreate()` returns rows in input order. This is an internal optimization, not a general Store batch API. Identity-enabled nodes, history/revision tracking, a member beyond the executor's declared claim-input budget, missing schema-fence support, and other unsupported shapes fail closed to the existing transaction or fallback behavior. The transport inventory for the supported libSQL root records one client `batch` submission and zero client `execute` calls for both generated-ID claim-free batches, multiple-claim batches, cross-scope claims, and claim-plus-projection batches. This is a measured submission count, not a wall-clock RTT benchmark; fallback paths are intentionally not assigned a latency claim. On D1, claim work is chunked inside the same submission rather than imposing a batch-wide ceiling. Each member has 87 claim-input binds after its row and fence: a canonical claim costs six, each legacy hierarchy-wide uniqueness probe costs nine, and each legacy disjointness probe costs six. Custom executors should call the exported `atomicNodeClaimInputCost()` owner rather than reproduce this formula. A member beyond that bound retains the portable behavior. Direct `edges.bulkInsert()` and `edges.bulkCreate()` calls on those same roots use one schema-fenced native program when history and revision capture are disabled. Declared durable match identities and `one`, `unique`, or `oneActive` cardinality are maintained inside that exchange; any endpoint, identity, or cardinality refusal rolls the whole call back. Transaction-scoped stores, derived backends, custom backends without the corresponding exact-root semantic registration, and dynamic get-or-create convergence retain the interactive path. Direct edge `bulkDelete()` calls use the same exact-root exchange and refuse a foreign-kind ID atomically. Restricted node `bulkDelete()` also releases every unique or disjoint claim owned by rows it tombstones in the same program, while enforcing live connected edges in SQL. Identity, projection, history, revision, cascade, and disconnect shapes retain their transaction path. `bulkUpsertById()` remains a resolved mutation set because it must read and schema-validate a database preimage before its writes are known. Bundled serverless roots can submit an eligible distinct-ID, live-row resolved set as one native exchange after that read. Bundled session-capable PostgreSQL can bind the same program to the exact collection-opened, caller-supplied, or adopted transaction; this is a bounded statement sequence on the pinned session, not one network exchange. Update-only sets use a guarded update; sets containing both fresh creates and updates include a terminal database assertion that rolls the whole exchange back when any guarded postimage is absent. Repeated IDs, resurrections, temporal changes, claims, edge sidecars (including durable edge match identity), history/revision capture, ordinary derived backends, and unregistered sessions use the interactive path. On D1's 100-parameter budget, each native statement carries at most 17 node mutations or 6 edge mutations. Larger eligible sets are chunked inside the same atomic transport submission; each chunk has its own terminal postimage assertion, so one refusal rolls every sibling chunk back rather than weakening the set contract. A D1 submission is bounded to 512 node members or 187 edge members; larger sets fail closed to the portable path instead of building an unbounded request. Other backends derive their statement width from their declared bind budget and retain an absolute 512-member submission ceiling. The operation returns an explicit `unsupported` verdict before issuing program SQL; the Store never infers fallback safety from a missing result. Once a session program starts, a savepoint preserves the surrounding transaction for typed refusal diagnosis. Node `bulkReplaceById()` avoids that structural preimage read by accepting only complete replacement documents and distinct IDs. On an eligible bundled root, the complete call—including claim ownership changes and fulltext/vector sidecars—uses one atomic transport submission. Live rows retain their stored validity windows; tombstones receive a freshly stamped window. Operational Identity and history/revision capture use the portable path. Custom backends must register and semantically certify the independent `replaceNodes` family; transport registration or another node family is not evidence for replacement. Eligible singleton `update()` and `delete()` calls reuse those same registered families. Plain or projected node updates, unconstrained non-durable-identity edge updates, all direct edge deletes, and plain restricted node deletes remain two-exchange operations—one authoritative read/gate and one atomic mutation—because TypeGraph must validate merged update properties and must preserve the rule that a missing delete fires no operation hooks. This removes explicit transaction transport from the eligible shape; it does not turn claims, edge sidecars, temporal, captured, derived-backend, or caller-transaction writes into autocommit operations. That singleton update path uses optimistic convergence: the mutation asserts the row preimage it read and retries a moved preimage up to four times. Under sustained same-row contention it can throw `DatabaseOperationError` where an interactive transaction would have waited to serialize the writers. This applies to eligible `update()` calls and the live-row leg of `upsertById()` on registered exact-root atomic transports. Caller transactions and other ineligible shapes continue to use the serialized transaction path. Applications using an atomic root should retry the operation when sustained contention can move the row throughout all four attempts. ### One `bulkUpsertById` batch cannot hand a constrained value between rows `bulkUpsertById` applies items in order for the purpose of deciding each row's final props, but it groups the writes: every create in the batch runs before every update. A batch where one item **releases** a constrained value and a later item **claims** it therefore fails, where the same operations applied one at a time succeed. - Nodes: releasing and re-claiming a `unique` constraint value in one batch throws `UniquenessError` — the claiming create is checked while the releasing row still reserves the value. - Edges: ending the lone `oneActive` edge from a source while creating its replacement throws `CardinalityError`, for the same reason. Bulk semantics are set-like, not scripted — a batch states the rows you want, not an order to reach them in — so this is a stated limitation rather than a pending fix. It always surfaces as a typed error, never as a dropped write. Split the handoff across two batches (release, then claim), or apply the conflicting items one at a time — as sequential `upsertById` calls for nodes, and as `update` then `create` for edges, which have no single-item upsert. See [Data Sync](/data-sync#one-batch-cannot-hand-a-unique-value-from-one-row-to-another) for the worked example. ## Graph Analytics Limits TypeGraph ships focused algorithms on `store.algorithms.*` — shortest path (weighted and unweighted), reachability, k-hop neighborhoods, degree, exact weakly connected components, deterministic label propagation, and global/personalized PageRank. See [Graph Algorithms](/graph-algorithms) for the full API. The following heavier analytics are **not** provided: - Modularity-optimizing community detection such as Leiden or Louvain - Centrality measures beyond degree (betweenness, closeness, eigenvector) - Strongly connected components - Topological sort - Graph partitioning For these use cases, export your data via `.query().traverse()` or `store.subgraph()` and use a specialized library such as [graphology](https://graphology.github.io/) in memory, or move to a dedicated graph database. ## Single Database Deployment TypeGraph is designed for single-database deployments. It does not support: - Distributed storage across multiple databases - Sharding - Cross-database queries - Replication coordination For distributed graph workloads, consider a dedicated graph database. ## Temporal Query Limitations Temporal queries (`asOf`, `includeEnded`) work correctly but have some constraints: - Point-in-time queries cannot be combined with streaming (`.stream()`) - `validFrom` defaults to the record's own creation timestamp when omitted, so `asOf` queries work out of the box; an end boundary still requires an explicit `validTo` — an open `validTo` means "still valid". A record written with a `validTo` at or before its own creation instant is "born already ended" and stores no lower bound instead, so it reads back at every `asOf` before that end - Rows an **older library version** stored with a backwards window (`valid_from > valid_to`) are readable at no coordinate, and upgrading does not rewrite them. Making them observable is an explicit operator action: run `repairInvertedValidityWindows({ relations: "live-and-recorded", mode: "apply" })` while writers are stopped, then re-baseline any outstanding merge branches. See [Repairing inverted validity windows](/schema-management#repairing-inverted-validity-windows) - Clock skew between application servers can affect temporal accuracy ### Recorded / system time (`history: true`) Recorded-time capture (`createStore(graph, backend, { history: true })`) and `store.asOfRecorded(T)` add a second temporal axis with these constraints. Use `createAdapterStore(..., { history: true })` instead when the application must adopt a caller-owned transaction: - **Opt-in, no backfill.** Capture only sees changes committed after it is enabled; an entity that already exists is first recorded the next time it is written. Enable it on a fresh graph for complete history. - **TypeGraph-write capture.** Built-in capture records TypeGraph collection writes only. Out-of-band database writes and row-returning raw SQL paths are not captured into the recorded relations. - **Reconstructing reads only.** A recorded view exposes point reads (`getById` / `getByIds`), bounded deterministic `scan()` pages, `query()`, `subgraph()`, and the graph algorithms. Broad filtered collection reads (`find` / `count` / `findFrom`), `search`, and fulltext / vector predicates are refused — those indexes reflect current state and cannot answer a recorded-time query. - **Transactional backend required.** Capture needs a backend with atomic transactions and statement execution — the built-in SQLite / PostgreSQL backends qualify. A custom backend must implement `executeStatement` (optional on the `GraphBackend` interface, but required once `history: true` is set) or enabling capture throws a `ConfigurationError` at write time. On an `AdapterHistoryStore`, raw `tx.sql` is disabled under `history: true`; adopt external transactions with `store.withRecordedTransaction(...)` instead of `store.withTransaction(...)` (which is a compile error on a history store). - **Reconstruction cost.** Recorded reads rebuild from the history relations and are slower than live reads, most noticeably for full-graph subgraph / algorithm reconstructions on PostgreSQL. - **PostgreSQL capture requires `READ COMMITTED`.** Every captured commit advances a single recorded-clock row for the graph. TypeGraph refuses PostgreSQL `REPEATABLE READ` / `SERIALIZABLE` history-capture transactions because snapshot isolation cannot safely allocate that per-graph recorded clock inside the captured transaction. Omit the transaction isolation option, or set it to `read_committed`. - **Recorded anchors are per graph.** Each captured transaction advances a fixed-width logical revision and pairs it with a non-decreasing physical wall-time high-water mark. TypeGraph does not provide a cross-graph recorded anchor. See [Logical revision and physical time](/queries/temporal#logical-revision-and-physical-time). - **The preview schema needs an offline migration.** Timestamp-only anchors and PostgreSQL recorded relations using `timestamptz` predate numeric recorded revisions and the `r1::` API encoding. Run `migrateLegacyRecordedTime()` while writers are stopped, then use `migrateRecordedAnchor()` for checkpoints held outside TypeGraph. See [Migrating preview recorded time](/schema-management#migrating-preview-recorded-time). ## Schema Migration Constraints Automatic migrations (`createStoreWithSchema`) only handle additive changes: | Change Type | Auto-Migrated | |-------------|---------------| | Add new node type | Yes | | Add new edge type | Yes | | Add optional property | Yes | | Add required property | No | | Remove property | No | | Rename type | No | | Change property type | No | Breaking changes throw `MigrationError` and require manual migration. # LLM Support > Machine-readable documentation for AI assistants and coding tools TypeGraph documentation is available in formats optimized for Large Language Models (LLMs) following the [llms.txt specification](https://llmstxt.org/). All files are generated from the same source docs as the website. ## Recommended Retrieval Order For coding agents, use these files progressively: 1. Start with [`/llms-small.txt`](/llms-small.txt) for implementation and debugging tasks. 2. Use [`/llms-full.txt`](/llms-full.txt) only when you need deep reference content. 3. Load [`/_llms-txt/examples.txt`](/_llms-txt/examples.txt) only when you need full end-to-end patterns. ## Available Files | File | Purpose | Size | |------|---------|------| | [`/llms.txt`](/llms.txt) | Index with page titles, descriptions, and links | Small | | [`/llms-small.txt`](/llms-small.txt) | Core docs for implementation and debugging tasks | Medium | | [`/llms-full.txt`](/llms-full.txt) | Complete documentation in a single file | Large | | [`/_llms-txt/examples.txt`](/_llms-txt/examples.txt) | Complete application examples | Medium | ## Copy-Paste Agent Instructions Use this in repository-level agent instruction files (`AGENTS.md`, `CLAUDE.md`, etc.): ```md TypeGraph (`@nicia-ai/typegraph`) is a TypeScript-first embedded knowledge graph library with typed nodes, edges, queries, and schema management over SQLite and PostgreSQL backends. When working with TypeGraph code (graph definitions, node/edge schemas, store operations, query builder, backend setup, or migrations): 1. Load https://typegraph.dev/llms-small.txt first. 2. Use https://typegraph.dev/llms-full.txt only for deep API/reference lookup. 3. Load https://typegraph.dev/_llms-txt/examples.txt only for end-to-end implementation patterns. 4. Prefer current API docs over inferred behavior from old snippets. ``` # Materializing External Event Logs > How to project at-least-once event streams into TypeGraph without making TypeGraph an event-log product External logs are the transport. TypeGraph is the typed, entity-resolved materialization and merge layer. Use this pattern when agents or integration runtimes already run on an event log or stream: Electric Durable Streams, database changefeeds, message queues, or a custom append-only feed. The log owns delivery, ordering, replay, and offsets. TypeGraph owns the current graph, valid-time facts, recorded-time history, and mergeable working copies. The sibling [`agent-stream-graph`](https://github.com/nicia-ai/agent-stream-graph) package is the reference implementation of this posture. ## The Shape of the Problem External log consumers usually have three properties: - **At-least-once delivery.** A change can be delivered more than once, especially after a crash or reconnect. - **Resume from a cursor.** The consumer persists the last source offset it has safely processed. - **Replay.** Reprocessing old events is normal: for recovery, backfills, or rebuilding a derived graph. That means a projector must be idempotent. Re-delivering the same source change should converge on the same graph state, not create duplicates. ## Idempotent Projectors Use stable source ids as TypeGraph ids whenever the source has them. For nodes, that usually means `upsertById`. For edges, prefer `getOrCreateByEndpoints`. Avoid `create` in a log projector unless the source event itself carries a unique id you pass as the TypeGraph id. ```typescript async function projectChange( tx: TransactionContext, change: Change, ) { const issue = await tx.nodes.Issue.upsertById( change.issueId, { title: change.title, state: change.state, }, { validFrom: change.issueValidFrom, onImmutableLowerBound: "preserve", }, ); const actor = await tx.nodes.Actor.upsertById(change.actorId, { name: change.actorName, }); await tx.edges.changedBy.getOrCreateByEndpoints( issue, actor, { action: change.action }, { ifExists: "update", validFrom: change.relationshipValidFrom, validTo: change.relationshipValidTo, onImmutableLowerBound: "preserve", }, ); } ``` The important rule is that the second delivery of the same change takes the same code path and reaches the same row identities. The `"preserve"` policy makes `validFrom` create/resurrection-only input for both node and edge writes: a later revision updates props and `validTo` without trying to rewrite the live row's start. Without it, the default `"refuse"` policy raises `IMMUTABLE_VALIDITY_LOWER_BOUND` when a revision states a different start. The edge also explicitly selects `ifExists: "update"`; the default is `"return"`, which is right for create-once relationships but writes neither revised props nor a closing `validTo` when the edge already exists. ### `matchOn` widens the identity key — don't reach for it by default `getOrCreateByEndpoints` matches on the endpoints `(from, to)` alone unless you pass `matchOn`. Endpoints-only is the **more** idempotent choice and is right for most projectors: a re-delivered edge between the same two nodes converges on the one existing edge regardless of how its properties drifted between deliveries. `matchOn` adds the named property fields to the match key, so it *widens* identity — two edges between the same endpoints are now distinct if they differ on a matched field. Use it only when the relationship model genuinely allows several parallel edges between one pair (say, one `changedBy` edge per distinct `action`), and know the footgun: if a re-delivered change carries a **changed** value in a matched field, it no longer matches the earlier edge and you get a **second** edge instead of convergence. Reach for `matchOn` when the domain needs the extra edges, not as a reflex. Validity timestamps do not become part of this identity key. If the same endpoints can have multiple application-time periods, include a stable period or source-event identifier in the edge schema and in `matchOn`. This keeps a re-delivery of one period convergent without collapsing a later period into the same edge. ## Cursor Bookkeeping A cursor is application state: the last source offset you have safely processed. It should advance only at a source offset boundary, after every change in that batch has been projected. Where the cursor lives — a row in your own relational table, or a node in the graph — decides which guarantees you can get. ### Exactly-once with an adopted transaction To commit the projected batch **and** the cursor as one unit, let the caller own the transaction and adopt it with [`store.withRecordedTransaction(externalTx, fn)`](/schemas-stores/#transaction-receipts). The graph writes and your own cursor write land on the same connection inside the same commit: either both persist or neither does, so the cursor can never advance past a batch the graph did not durably record. Two constraints make this the *only* sanctioned transactional recipe on a store created with `createAdapterStore` or `createAdapterStoreWithSchema` and `{ history: true }` — which the Transaction Receipts and Bitemporal sections below both require: - **Write your own tables through the external handle you passed in**, never through `tx.sql`. Under history capture the typed transaction context omits `sql` (raw SQL would bypass recorded-time capture); suppressed access reaches a runtime guard and raises a [`ConfigurationError`](/errors/#recorded-capture-guard-codes). The external handle *is* the pinned connection, so writing your cursor row through it keeps both layers in the one transaction. - **`store.withTransaction()` — the non-recorded sibling — is a compile error on a history store** (its `externalTx` argument is rejected against a message type), and its runtime guard throws `RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION`. It has no flush point before the caller commits, so recorded-time capture could not seal. Use `withRecordedTransaction` instead. **Async drivers (Postgres / libsql)** open the boundary with `db.transaction`: ```typescript const receipt = await db.transaction(async (dbTx) => { const outcome = await store.withRecordedTransaction(dbTx, async (tx) => { for (const change of batch.changes) { await projectChange(tx, change); } }); // The cursor row goes through the external handle, in the same transaction. await dbTx .insert(streamCursors) .values({ sourceId: batch.sourceId, offset: batch.endOffset }) .onConflictDoUpdate({ target: streamCursors.sourceId, set: { offset: batch.endOffset }, }); return outcome.receipt; }); // one COMMIT / ROLLBACK across both layers ``` **Synchronous `better-sqlite3`** cannot adopt an `async` transaction callback (its driver rejects a promise-returning `db.transaction`), so the caller frames the boundary by hand with `BEGIN IMMEDIATE` / `COMMIT` / `ROLLBACK` on the single connection: ```typescript await db.run(sql`BEGIN IMMEDIATE`); try { const { receipt } = await store.withRecordedTransaction(db, async (tx) => { for (const change of batch.changes) { await projectChange(tx, change); } await db.run(sql` INSERT INTO stream_cursor (source_id, offset) VALUES (${batch.sourceId}, ${batch.endOffset}) ON CONFLICT (source_id) DO UPDATE SET offset = excluded.offset `); }); await db.run(sql`COMMIT`); // persist receipt.recorded as the offset's replay anchor — see below } catch (error) { await db.run(sql`ROLLBACK`); // graph writes and cursor roll back together throw error; } ``` The graph writes and your own statements share the caller's one pinned connection. TypeGraph serializes the statements its collections issue; sequence your own raw statements yourself (don't `Promise.all` them with graph writes) so two queries never race on that connection. For an adapter-backed materializer that runs against several backends, branch on capability rather than message-matching: use [`tx.sqlAvailability`](/recipes/#cross-store-transactions-drizzle--typegraph) to decide whether raw SQL is usable inside `store.transaction`, and [`isRecordedCaptureGuardError(error, code?)`](/errors/#recorded-capture-guard-codes) to recognize a history-store guard when you catch one. ### At-least-once with a separate cursor store When the runtime already owns checkpointing, or the backend cannot provide atomic transactions (`backend.capabilities.execution.interactiveTransactions === false` — Cloudflare D1, `drizzle-orm/neon-http`), keep the cursor outside the graph transaction. The pattern is at-least-once plus idempotence: a crash after the graph writes but before the cursor write replays the batch, which is safe precisely because the projector converges. This fallback requires a raw Store. A schema-managed Store refuses writes on a non-transactional backend because it cannot hold the schema-version fence. Use a transactional driver, or deliberately construct a raw Store and own schema/write coordination yourself. ```typescript for (const change of batch.changes) { // Each successful projection may commit before a later projection or cursor write fails. await projectChange(store, change); } await cursorStore.save({ sourceId: batch.sourceId, offset: batch.endOffset, }); ``` This at-least-once path plus an idempotent projector is the workload TypeGraph is built for. It is also the one that churns recorded history the hardest: every re-delivery of a byte-identical change rewrites its row, allocating a fresh recorded instant and a new history row per delivery. Enable [`coalesceUnchangedUpserts: true`](/schemas-stores/#createstoregraph-backend-options) on the store to suppress that. A node `upsertById` and an edge endpoint get-or-create update perform no write, history row, or revision advance when its validated props and requested window already equal the live row. Their bulk forms have the same behavior. See [Transaction Receipts](#transaction-receipts) for how a coalesced upsert reads on a receipt. Every captured transaction receives one versioned recorded instant: a strict per-graph logical revision paired with a non-decreasing physical wall-time high-water mark. High commit rates consume revisions without pushing the timestamp beyond observed wall time. A backward clock correction holds the physical component at its prior value until the clock catches up, preserving cumulative diagonal checkpoint replay. Group changes by their durable replay/checkpoint boundary so one addressable source position consumes one recorded instant where practical. Cap transaction size independently: a source may expose one coarse checkpoint for a very large initial sync, but that does not make an unbounded transaction safe. Recorded clocks are independent per graph, and there is no cross-graph `recordedNow()` snapshot. See [Logical revision and physical time](/queries/temporal#logical-revision-and-physical-time) for the anchor encoding and replay semantics. **Coalescing eliminates *re-delivery* churn, not replay cost.** The win is scoped to re-delivery of the current value — the realistic at-least-once case, where a change that was already applied arrives again (a crash-window replay, a duplicate) and is value-identical to the live row. A full **replay-from-zero** over the current state is different: if the stream contains in-place updates, replaying `insert a=1 … update a=2` re-applies `a=1` over the live `a=2` — a genuine backward change that writes — and then `a=2` restores it. Both writes are correct (the replay faithfully re-walks each historical state), but "coalescing makes replay free" holds only for streams whose rows never supersede each other. It also leaves a spurious `a=2 → a=1 → a=2` band in the live store's recorded history, stamped at replay time. To rebuild without either cost, replay into a **fresh store** and publish it, rather than re-applying the log over the current state. **Historical ends need a historical start on the creating event.** When a fresh store creates a row without a stated `validFrom`, TypeGraph uses the ingest instant. A later replayed event whose historical `validTo` precedes that ingest instant is therefore an `INVERTED_VALIDITY_WINDOW`, even if the source timeline itself was ordered. An event-time decoder must emit `validFrom` on the event that first creates each node or edge; `onImmutableLowerBound: "preserve"` then lets later revisions carry their source bound without trying to move the stored start. ### In-graph cursors and the receipt A cursor can also live inside the graph as an ordinary node — convenient, and it travels with the graph. But if you also use the transaction receipt (next section) to detect a projector that dropped a change, an in-graph cursor **corrupts that signal**: the receipt counts writes per transaction with no attribution, so the cursor's own upsert is indistinguishable from the projector's writes. A projector that drops a change in a transaction that also checkpoints an in-graph cursor produces `writes.total === 1` from the cursor alone — `writes.total > 0` no longer means "the projector wrote," the drop goes undetected, and the cursor advances past the lost event. Two ways out: - **Scope the projector with `tx.measure`.** On a receipt-enabled context (`transactionWithReceipt` or `withRecordedTransaction`), [`tx.measure((scopedTx) => …)`](/schemas-stores/#scoped-receipts-txmeasure) hands your callback a **scoped context** and returns a sub-receipt that counts exactly the writes made **through that scoped context** (`scopedTx.nodes` / `scopedTx.edges`). Run the projector through `scopedTx`; write the cursor through the outer `tx` (or your own table). Attribution is by which context you write through, not by timing, so the cursor's write counts only in the outer receipt and the scope reflects the projector alone. This is what makes an in-graph cursor and drop-detection composable; see the full loop below. - **Keep the cursor in your own relational table.** The [exactly-once recipe](#exactly-once-with-an-adopted-transaction) makes that atomic anyway, and it keeps the belief graph pristine: an in-graph cursor node still lands in every `asOfRecorded` reconstruction of the graph, so a consumer that wants recorded-time reads to show only projected facts should keep the cursor out of the graph entirely. ## Transaction Receipts When you need to know what a projector did, use `store.transactionWithReceipt` (TypeGraph owns the boundary) or `store.withRecordedTransaction` (you adopt an open transaction); both return a `TransactionOutcome` with a `receipt`. The receipt carries **two signals that deliberately disagree**, and a materializer needs both. `receipt.writes.total` counts completed write intents at the collection surface; `receipt.recorded` (on a `{ history: true }` store) is the recorded commit instant this transaction allocated, or `undefined` when nothing was captured or explicitly requested. The common, load-bearing case is where they diverge: | case | `writes.total` | `recorded` | | --------------------------------------------- | -------------- | --------------- | | projector wrote | `> 0` | defined | | no-op delete of an absent key (a real intent) | `1` | **`undefined`** | | coalesced upsert (value-identical, opt-in) | `1` | **`undefined`** | | projector dropped the change | `0` | `undefined` | | explicit recorded revision request | `0` | defined | A no-op delete completes a write *intent* but captures nothing; a [coalesced upsert](/schemas-stores/#createstoregraph-backend-options) is the same shape by design. In both, `writes.total` counts (the method resolved) but `recorded` is `undefined`. **An offset whose transaction reports `recorded === undefined` must carry the prior anchor forward** — otherwise replay-by-offset breaks at exactly the offsets where nothing changed. Call `requestRecordedRevision()` when that offset must instead receive its own anchor despite making no entity change. Two counting rules bite materializers specifically, both worth internalizing before you read `writes.total` as "the projector did work": - **Bulk methods count by input length**, so `bulkCreate([])` contributes `0`. A projector that filters a batch down to nothing and issues an empty bulk call must not read as a writer. - **A method that rejects counts `0`** — even on SQLite, where a failed statement does **not** abort the surrounding transaction. A projector that swallows a write error and commits can persist rows the receipt never counted, so do not read the receipt as rows-affected in that scenario. ### The full materializer loop Putting the pieces together: an adopted transaction for exactly-once cursors, a `tx.measure`-scoped projector so a single dropped change is caught within a multi-change batch — the outer receipt only tells you the whole batch wrote nothing, whereas a `measure` scope attributes writes per change by having the projector write through the scoped context it receives (any cursor written through the outer `tx` stays out of that count) — `receipt.recorded` as the per-offset replay anchor, and `writes.total === 0` on a non-delete change as the drop signal. `withRecordedTransaction` flushes recorded-time capture and resolves **before** the caller's commit, so `outcome.receipt.recorded` is already known inside the `db.transaction` callback. Write the cursor advance **and** its replay anchor through `dbTx` there, in the same commit as the graph writes. Persisting the anchor after the commit — as a separate step — would reopen the exactly-once gap the adopted transaction exists to close: a crash between the commit and the anchor write leaves the cursor advanced with no anchor, and that offset can never be replayed. If an accepted source position must be addressable even when the projector makes no entity changes, request an explicit revision in the callback. Await the `withRecordedTransaction` outcome, then insert the application-owned cursor and returned anchor through the still-open native transaction before its commit: ```typescript await db.transaction(async (dbTx) => { const { receipt } = await store.withRecordedTransaction(dbTx, async (tx) => { tx.requestRecordedRevision(); await projectBatch(tx, batch); }); await dbTx.insert(cursors).values({ source: batch.source, offset: batch.offset, recorded: receipt.recorded, }); }); ``` ```typescript let lastAnchor: RecordedInstant | undefined = await loadLastAnchor(); // on resume lastAnchor = await db.transaction(async (dbTx) => { const outcome = await store.withRecordedTransaction(dbTx, async (tx) => { for (const change of batch.changes) { // The projector writes through the scoped context, so `projected` // counts its writes alone — nothing else in the transaction. const projected = await tx.measure((scopedTx) => projectChange(scopedTx, change), ); // A non-delete change that wrote nothing was silently dropped. if ( projected.receipt.writes.total === 0 && change.operation !== "delete" ) { throw new DroppedChangeError(change); // rolls the whole batch back } } }); // `recorded` is undefined when the batch captured nothing (all drops, no-op // deletes, or coalesced upserts) — carry the prior anchor forward so replay // by offset still resolves. Anchor comes from the receipt, never from a // post-commit store.recordedNow() (see below). const anchor = outcome.receipt.recorded ?? lastAnchor; // Cursor and anchor commit atomically with the graph writes: no window where // the cursor has advanced past an offset whose anchor was never persisted. await dbTx .insert(offsetAnchors) .values({ sourceId: batch.sourceId, offset: batch.endOffset, recorded: anchor }); await dbTx .insert(streamCursors) .values({ sourceId: batch.sourceId, offset: batch.endOffset }) .onConflictDoUpdate({ target: streamCursors.sourceId, set: { offset: batch.endOffset }, }); return anchor; // updates lastAnchor only once the transaction commits }); ``` **Take the replay anchor from `receipt.recorded`, never from a post-commit `store.recordedNow()`.** `recordedNow()` is the graph-global recorded high-water mark, advanced by **any** writer to the graph. Between your commit and your read of it, a concurrent writer can advance it, and `asOfRecorded(that)` then reconstructs a belief your stream never produced. The receipt hands you the instant *this* transaction allocated; that is the only anchor that reconstructs exactly what this offset materialized. ## Bitemporal Mapping External streams usually carry domain time and delivery time. Keep those separate: - **Event time belongs in valid time.** If a source change says a fact became true on January 1, pass that timestamp as `validFrom`; if it ended on January 31, pass `validTo`. - **Ingest time is recorded time.** TypeGraph records when the graph committed the write. Recorded time is allocated by the backend and cannot be backdated. - **Backfills collapse recorded instants to now.** Replaying historical events today writes historical valid-time facts with today's recorded-time anchors. That is correct SQL:2011 bitemporal behavior, not a bug. To replay by source offset, load the anchor you saved for that offset and read a recorded-time view. The `receipt.recorded` you persisted is a branded `RecordedInstant`, but round-tripping through your cursor table stores it as a plain string — re-brand it with `asRecordedInstant` on the way back before passing it to `asOfRecorded`: ```typescript import { asRecordedInstant } from "@nicia-ai/typegraph"; const stored = await offsetAnchors.anchorFor(offset); // plain string from storage const anchor = asRecordedInstant(stored); // validates + re-brands const graphAtOffset = store.asOfRecorded(anchor); const issue = await graphAtOffset.nodes.Issue.getById(issueId); ``` If the cursor table contains timestamp-only anchors from the recorded-time preview, migrate the TypeGraph relations first and remap those cursor values with `migrateRecordedAnchor({ backend, graphId, anchor: stored })`. See [Migrating preview recorded time](/schema-management#migrating-preview-recorded-time). That answers "what did the materialized graph know after offset X?" even if later corrections changed or deleted rows. See [Recorded time](/queries/temporal/#recorded-time-bitemporal) for the full view surface. ### Refresh planner statistics after a large replay A **custom** replay or backfill loop — one built from the projector recipes above — runs its writes **inside a caller-provided transaction**, which never auto-refreshes the query planner's table statistics: `ANALYZE` from another connection cannot see rows that are still uncommitted, so the store deliberately skips the automatic refresh it does after large autocommit bulk writes. Left alone, the planner keeps pre-load row estimates and can pick an order-of-magnitude-slower plan. After a large custom replay, refresh once: ```typescript await replayEverything(); await store.refreshStatistics(); // once, after the bulk replay commits ``` The interchange path handles this for you: `importGraph` and `importGraphStream` call `refreshStatistics()` once after the import commits (see [Bulk Copy Between Stores](#bulk-copy-between-stores)), so a bulk copy needs no manual refresh. ## Bulk Copy Between Stores To copy a materialized graph into another store — most often a graph-merge working copy — stream interchange directly from source to target with `exportGraphStream` / `importGraphStream`. This is the same path [graph-merge](/interchange/) uses internally, so a copy produces byte-identical merge results, conflicts, and provenance to a native branch: ```typescript import { exportGraphStream, importGraphStream, } from "@nicia-ai/typegraph/interchange"; const result = await importGraphStream( branch.store, exportGraphStream(beliefStore, { nodeKinds: ["Belief", "Claim"], edgeKinds: ["supports"], includeTemporal: true, }), { onConflict: "update" }, ); if (!result.success) { throw new Error(`copy failed with ${result.errors.length} import errors`); } ``` Two option defaults are exactly right here and worth stating because they are not obvious: - **`includeDeleted` defaults to `false`, and the copy clones live state — it does not synchronize deletions.** The exporter simply omits soft-deleted rows; it cannot round-trip `deletedAt` at all (the wire format carries no deletion flag). So a fact deleted on the source is merely *absent* from the stream: on a fresh target it never appears, but on a populated target an existing live row **stays live** — the copy never deletes it. If the target must reflect deletions, apply them through your projector, not the bulk copy. - **`includeTemporal` must be set to `true`** (it defaults to `false`). It is what carries each fact's original `validFrom` / `validTo` across the copy; without it the import re-stamps every fact with the *copy's* wall clock, destroying valid-time fidelity in the merged branch. `importGraphStream` preserves ids, routes existing rows through normal `onConflict` handling, validates edge endpoints (`validateReferences` defaults to `true`), and refreshes planner statistics once after the import commits. ## Cursor-Based Resumption and Electric The examples above assume a per-change offset. **Electric does not provide one** — every change in a `ShapeStream` catch-up batch shares the stream's `lastOffset`. A cursor keyed on Electric's offset can therefore only advance at a **batch boundary**, after the whole batch is projected. Advancing mid-batch is unsafe: Electric's `read(after)` is strictly-after, so resuming from a mid-batch offset permanently skips that batch's remaining changes. Project the whole batch, then checkpoint the cursor once at its boundary. # Multiple Graphs > Using separate graph definitions for different domains in the same application TypeGraph supports multiple graphs for applications that have distinct data domains that benefit from separate graph definitions. ## When to Use Multiple Graphs Use separate graphs when you have: - **Distinct domains**: A RAG system for documents and a business network for suppliers have different node types, edge semantics, and query patterns - **Independent lifecycles**: One graph might evolve rapidly while another is stable - **Team ownership**: Different teams own different graphs, with separate schema review processes - **Different retention policies**: Document chunks might be ephemeral while business relationships are long-lived **Don't use multiple graphs** when: - You need cross-graph queries or traversals (use a single graph with ontology relations instead) - The domains are closely related (e.g., Users and Documents that Users author) - You're trying to solve multi-tenancy (use tenant isolation patterns instead) ## Example: Documents and Business Network A company needs two graphs: 1. **Documents graph**: Powers semantic search over internal documents 2. **Organization graph**: Tracks suppliers, partners, and contracts ### Defining the Graphs ```typescript // graphs/documents.ts import { z } from "zod"; import { defineNode, defineEdge, defineGraph, embedding } from "@nicia-ai/typegraph"; const Document = defineNode("Document", { schema: z.object({ title: z.string(), source: z.string(), createdAt: z.string().datetime(), }), }); const Chunk = defineNode("Chunk", { schema: z.object({ content: z.string(), embedding: embedding(1536), position: z.number().int(), }), }); const hasChunk = defineEdge("hasChunk"); export const documentsGraph = defineGraph({ id: "documents", nodes: { Document: { type: Document }, Chunk: { type: Chunk }, }, edges: { hasChunk: { type: hasChunk, from: [Document], to: [Chunk] }, }, }); ``` ```typescript // graphs/organization.ts import { z } from "zod"; import { defineNode, defineEdge, defineGraph, subClassOf } from "@nicia-ai/typegraph"; const Organization = defineNode("Organization", { schema: z.object({ name: z.string(), domain: z.string().optional(), }), }); const Supplier = defineNode("Supplier", { schema: z.object({ name: z.string(), domain: z.string().optional(), category: z.enum(["materials", "services", "logistics"]), }), }); const Partner = defineNode("Partner", { schema: z.object({ name: z.string(), domain: z.string().optional(), partnershipLevel: z.enum(["bronze", "silver", "gold"]), }), }); const Contract = defineNode("Contract", { schema: z.object({ title: z.string(), value: z.number(), startDate: z.string().datetime(), endDate: z.string().datetime().optional(), status: z.enum(["draft", "active", "expired"]).default("draft"), }), }); const supplies = defineEdge("supplies"); const hasContract = defineEdge("hasContract"); export const organizationGraph = defineGraph({ id: "organization", nodes: { Organization: { type: Organization }, Supplier: { type: Supplier }, Partner: { type: Partner }, Contract: { type: Contract }, }, edges: { supplies: { type: supplies, from: [Supplier], to: [Organization] }, hasContract: { type: hasContract, from: [Organization], to: [Contract] }, }, ontology: [ subClassOf(Supplier, Organization), subClassOf(Partner, Organization), ], }); ``` ### Creating Stores Both graphs can share the same database backend. Each graph's data is isolated by its `id`. ```typescript // stores.ts import { createStore } from "@nicia-ai/typegraph"; import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { drizzle } from "drizzle-orm/node-postgres"; import { Pool } from "pg"; import { documentsGraph } from "./graphs/documents"; import { organizationGraph } from "./graphs/organization"; const pool = new Pool({ connectionString: process.env.DATABASE_URL }); const db = drizzle(pool); const backend = createPostgresBackend(db); // Same backend, different stores export const documentsStore = createStore(documentsGraph, backend); export const organizationStore = createStore(organizationGraph, backend); ``` ### Using the Stores Each store is fully independent with its own typed API: ```typescript // Semantic search in documents async function searchDocuments(query: string, embedding: number[]) { return documentsStore .query() .from("Chunk", "c") .whereNode("c", (c) => c.embedding.similarTo(embedding, 10)) .select((ctx) => ({ content: ctx.c.content, position: ctx.c.position, })) .execute(); } // Business queries in organization async function getActiveSuppliers(category: string) { return organizationStore .query() .from("Supplier", "s") .whereNode("s", (s) => s.category.eq(category)) .traverse("hasContract", "e") .to("Contract", "c") .whereNode("c", (c) => c.status.eq("active")) .select((ctx) => ({ supplier: ctx.s.name, contract: ctx.c.title, value: ctx.c.value, })) .execute(); } ``` ## Coordinating Across Graphs Since cross-graph queries aren't supported, coordinate at the application level. ### Shared Identifiers Use consistent IDs when entities relate across graphs: ```typescript // When ingesting a supplier's documents, use the supplier ID as a reference async function ingestSupplierDocument( supplierId: string, title: string, content: string, embedding: number[] ) { // Store document with supplier reference in metadata const doc = await documentsStore.nodes.Document.create({ title, source: `supplier:${supplierId}`, createdAt: new Date().toISOString(), }); const chunk = await documentsStore.nodes.Chunk.create({ content, embedding, position: 0, }); await documentsStore.edges.hasChunk.create(doc, chunk, {}); return doc; } // Later, find documents for a supplier async function getSupplierDocuments(supplierId: string) { return documentsStore .query() .from("Document", "d") .whereNode("d", (d) => d.source.eq(`supplier:${supplierId}`)) .select((ctx) => ctx.d) .execute(); } ``` ### Application-Level Joins Combine results from multiple graphs in your application: ```typescript interface SupplierWithDocuments { supplier: { name: string; category: string }; documents: Array<{ title: string }>; } async function getSupplierOverview( supplierId: string ): Promise { // Parallel queries to both graphs const [supplier, documents] = await Promise.all([ organizationStore.nodes.Supplier.getById(supplierId), getSupplierDocuments(supplierId), ]); return { supplier: { name: supplier.name, category: supplier.category, }, documents: documents.map((d) => ({ title: d.title })), }; } ``` ### Event-Driven Sync For loose coupling, use events to keep graphs in sync: ```typescript // When a supplier is created, set up document ingestion eventBus.on("supplier.created", async (event) => { const { supplierId, name } = event.payload; // Create a placeholder document node for future ingestion await documentsStore.nodes.Document.create({ title: `${name} - Supplier Profile`, source: `supplier:${supplierId}`, createdAt: new Date().toISOString(), }); }); // When a supplier is deleted, clean up related documents eventBus.on("supplier.deleted", async (event) => { const { supplierId } = event.payload; const docs = await documentsStore .query() .from("Document", "d") .whereNode("d", (d) => d.source.eq(`supplier:${supplierId}`)) .select((ctx) => ctx.d.id) .execute(); for (const docId of docs) { await documentsStore.nodes.Document.delete(docId); } }); ``` ## Separate Backends For stronger isolation, use separate database connections: ```typescript // Documents in PostgreSQL with pgvector for embeddings const documentsPool = new Pool({ connectionString: process.env.DOCUMENTS_DATABASE_URL, }); const documentsBackend = createPostgresBackend(drizzle(documentsPool)); export const documentsStore = createStore(documentsGraph, documentsBackend); // Organization data in a separate database const orgPool = new Pool({ connectionString: process.env.ORG_DATABASE_URL, }); const orgBackend = createPostgresBackend(drizzle(orgPool)); export const organizationStore = createStore(organizationGraph, orgBackend); ``` **When to separate backends:** - Different performance profiles (vector search vs. relational queries) - Compliance requirements (PII in one database, analytics in another) - Independent scaling needs - Different backup/retention policies ## Schema Management Each graph has independent schema versioning: ```typescript import { createStoreWithSchema } from "@nicia-ai/typegraph"; // Each graph tracks its own schema version const [documentsStore, docsSchemaResult] = await createStoreWithSchema( documentsGraph, backend ); const [orgStore, orgSchemaResult] = await createStoreWithSchema( organizationGraph, backend ); // Check migration status independently if (docsSchemaResult.status === "migrated") { console.log("Documents schema was migrated"); } if (orgSchemaResult.status === "migrated") { console.log("Organization schema was migrated"); } ``` ## Inspecting What a Database Holds Graphs sharing a backend are separated by `graph_id` inside TypeGraph's tables. Two reads answer the questions an operator asks about that layout without depending on it: which graphs live in this database, and how many rows one graph holds. ### `listGraphIds(backend, options?)` Lists the graph ids that hold data, one bounded page at a time: ```typescript import { listGraphIds } from "@nicia-ai/typegraph"; let after: string | undefined; for (;;) { const page = await listGraphIds(backend, { prefix: "tenant-", after, limit: 100 }); if (page.length === 0) break; for (const graphId of page) console.log(graphId); after = page.at(-1); } ``` | Option | Meaning | | -------- | ----------------------------------------------------------------------------------- | | `prefix` | Only ids starting with this exact, case-sensitive text. `%` and `_` are not wildcards. | | `after` | Exclusive cursor: only ids ordered after this one. Pass the last id of the previous page. | | `limit` | Page size from 1 to 1000. Defaults to 100. Anything else throws `ConfigurationError`. | Ids come back in byte order (UTF-8 code point order) on every backend, so `Tenant-x` sorts before `tenant-a` on SQLite and PostgreSQL alike and a cursor resumes exactly where the last page ended, whatever the database collation. The reserved deployment marker id that TypeGraph uses for deployment-scoped contribution markers is never listed. A graph appears while it has nodes, edges or a committed schema version, which are exactly the relations a default `store.clear()` empties. A cleared graph therefore stops being listed even though `store.clear()` keeps its contribution markers unless you pass `preserveContributionMaterializations: false`, and even when a revision-tracked store reseeded its `recordedClock` row during the clear. Each page walks graph ids by index seek, one seek per graph per relation, instead of reading every row. The walk starts at the cursor or prefix and stops after the page, so a page costs about `limit` seeks wherever it sits, however many graphs the database holds and however many rows they contain. SQLite serves the seeks from the `graph_id`-leading primary keys, which are already in byte order. PostgreSQL orders ordinary text indexes by the database collation, so it serves them from a byte-ordered (`COLLATE "C"`) `graph_id` index that base-schema version 5 adds to `nodes`, `edges` and `schema_versions`; see [Base-schema version 5](/backend-setup#base-schema-version-5-byte-ordered-graph_id-indexes-postgresql) for what it costs and how to build it ahead of an upgrade. On a 20,000-graph, 50-rows-per-graph PostgreSQL 18 database a page takes about 3 ms, where the same read took about 400 ms before the index; at 200 graphs of 5,000 rows it is about 3 ms either way. A database without the index (its base schema not adopted yet, or DDL managed by hand) lists the same ids by reading and de-duplicating every row of those relations for each page, measured at 47 to 105 ms a page at these sizes. A backend that declares no recursive traversal does the same. Use the listing for operator tooling, not on a request path. Rows that exist only outside those relations, such as orphaned recorded history or contribution markers, do not make a graph appear; `inspectGraphStorage` counts every relation. The read runs in one read-only transaction where the backend supports it. It needs the backend's catalog probes to tell a table that was never provisioned from an empty one, and throws `ConfigurationError` on a custom backend that has none. ### `inspectGraphStorage(store)` Counts one graph's rows in every relation that can hold them: ```typescript import { inspectGraphStorage } from "@nicia-ai/typegraph"; await store.clear(); const { graphId, relations, totalRows } = await inspectGraphStorage(store); const leftovers = relations.filter((relation) => relation.rows > 0); // [{ relation: "contributionMaterializations", table: "typegraph_contribution_materializations", rows: 2 }] ``` `relations` lists every graph-scoped relation under its logical key (`nodes`, `edges`, `uniques`, `edgeClaims`, `identityAssertions`, `recordedNodes`, `fulltext`, `schemaVersions`, and so on) with the physical `table` it resolved to on this backend, so custom table names are reported as configured. The graph's per-field vector tables come from its vector slots and the active vector strategy and are reported as `vector:.`. A relation whose table the database never provisioned counts as `0` rather than failing. Use it to verify that `store.clear()` left nothing behind. Two relations can legitimately hold a row after a clear, by design: - `contributionMaterializations` is preserved unless you pass `preserveContributionMaterializations: false`. - `recordedClock` is reseeded inside the clear transaction on a store with live revision tracking (without history). Every other relation reads `0` after a clear, and other graphs in the same database are untouched. #### Consistency of the counts Each relation is counted by its own statement, so `relations` and `totalRows` describe one state of the graph only when every statement read the same snapshot. The result carries a `consistency` field that says whether they did: | `consistency` | Meaning | | --- | --- | | `"snapshot"` | Every count came from one snapshot: `relations` and `totalRows` describe a state the graph was in. | | `"per-statement"` | Each relation was counted independently. A write between two counts can leave the result describing a state that never existed together, for example rows in `nodes` beside an empty `schemaVersions`. | The read asks for a read-only `repeatable read` transaction, but it does not trust the request: a transaction wrapper can drop the isolation option, and a role or database can default the level. The effective level is read on the counting session itself, inside the first count statement, so the answer costs no extra round trip. What that gives on each backend: - **SQLite (better-sqlite3, libSQL, and other drivers with interactive transactions):** always `"snapshot"`. A SQLite transaction reads one snapshot whatever level was requested. - **PostgreSQL (`pg`, `postgres-js`, PGlite):** `"snapshot"` when the session was observed at `repeatable read` or `serializable`, which is what the request produces. `"per-statement"` when it ran at `read committed`, which happens when a wrapper around `backend.transaction` does not forward its options and the role or database defaults to `read committed`; the same wrapper under a `repeatable read` default still reports `"snapshot"`, because the level is observed, not requested. A backend that declares no session isolation read cannot be observed and reports `"per-statement"`. - **Backends without interactive transactions (Cloudflare D1, `neon-http`):** `"per-statement"`, because there is no transaction to share a snapshot. The exception is a graph with at most one provisioned relation, which is one statement and so trivially consistent. The evidence proves the isolation of the session that ran the first count. A backend wrapper that violates the transaction contract by handing the root pool through as its transaction backend can run later counts on other sessions, which no observation on the first one can detect. The read never refuses on a weaker level: it is a diagnostic. Treat `"per-statement"` counts as an approximation. To verify a clear with them, make sure nothing else writes the graph while you read, or read twice and compare. ## Shared Subgraph Helpers When multiple graphs share a common set of node and edge types, you can write reusable helpers that accept any store containing that shared subgraph. The `StoreProjection` utility type makes this type-safe without coupling to a specific graph definition. ### Defining shared types and graphs Start with the shared node and edge types, then define the graphs that use them: ```typescript import { createStore, defineNode, defineEdge, defineGraph, type Node, type StoreProjection, } from "@nicia-ai/typegraph"; const Document = defineNode("Document", { schema: z.object({ title: z.string() }), }); const Chunk = defineNode("Chunk", { schema: z.object({ text: z.string() }), }); const Comment = defineNode("Comment", { schema: z.object({ text: z.string() }), }); const hasChunk = defineEdge("hasChunk", { from: [Document], to: [Chunk] }); const aboutChunk = defineEdge("aboutChunk", { from: [Comment], to: [Chunk] }); const reviewGraph = defineGraph({ id: "review", nodes: { Document: { type: Document }, Chunk: { type: Chunk }, Comment: { type: Comment }, Label: { type: Label }, }, edges: { hasChunk, aboutChunk, hasLabel }, }); const catalogGraph = defineGraph({ id: "catalog", nodes: { Document: { type: Document, unique: [{ name: "title_unique", fields: ["title"], scope: "kind", collation: "binary" }], }, Chunk: { type: Chunk }, Comment: { type: Comment }, Category: { type: Category }, }, edges: { hasChunk, aboutChunk, inCategory }, }); ``` ### Projecting a shared subgraph Define a projection against either graph — it picks only the shared keys: ```typescript type CoreStore = StoreProjection< typeof reviewGraph, "Document" | "Chunk" | "Comment", "hasChunk" | "aboutChunk" >; ``` ### Writing a reusable helper ```typescript async function addComment( store: CoreStore, chunk: Node, text: string, ) { const comment = await store.nodes.Comment.create({ text }); await store.edges.aboutChunk.create(comment, chunk); return comment; } ``` ### Using across different graphs The same `addComment` function works with any store whose graph includes the projected nodes and edges — even if the graphs diverge on other types or unique constraints: ```typescript const reviewStore = createStore(reviewGraph, backend); const catalogStore = createStore(catalogGraph, backend); await addComment(reviewStore, chunk, "needs revision"); await addComment(catalogStore, chunk, "good categorization"); ``` The projection also works inside transactions — `TransactionContext` is structurally assignable to `StoreProjection` for the same keys: ```typescript await reviewStore.transaction(async (tx) => { await addComment(tx, chunk, "transactional comment"); }); ``` ### What the projection strips `StoreProjection` erases node constraint names, making constraint-based methods like `findByConstraint` uncallable through the projection. This is intentional: unique constraints are graph-registration-level details that typically differ between graphs sharing the same node types. If you need constraint access, type the helper against a specific `Store` instead. ## Caveats **No cross-graph queries**: You cannot traverse from a node in one graph to a node in another. If you need this, consider: - Merging the graphs into one with clear ontology separation - Using application-level joins as shown above **Separate ontology closures**: Each graph computes its own `subClassOf`, `implies`, etc. closures. Ontology relations don't span graphs. **Independent transactions**: A transaction in one store doesn't include the other. For cross-graph consistency, use sagas or eventual consistency patterns. **Shared tables**: When using the same backend, both graphs write to the same `typegraph_nodes` and `typegraph_edges` tables, differentiated by `graph_id`. This is fine for most cases but means a database-level issue affects both graphs. ## Next Steps - [Multi-Tenant SaaS](./examples/multi-tenant) - Isolating data by tenant within a single graph - [Schema Migrations](./schema-management) - Versioning and migrations - [Integration Patterns](./integration) - More deployment strategies # Ontology & Reasoning > Semantic relationships, type hierarchies, and inference ## When Do You Need an Ontology? An ontology captures **meaning** about your data—relationships that exist at the type level, not just instance level. You need ontology when: - **Type hierarchies**: "A Podcast is a type of Media" (query for Media, get Podcasts too) - **Concept relationships**: "Machine Learning is narrower than AI" (topic navigation) - **Constraints**: "A Person cannot also be an Organization" (prevent invalid data) - **Edge implications**: Query `knows` through more-specific `marriedTo` rows when explicitly requested - **Bidirectional queries**: "manages and managedBy are inverses" (traverse in either direction) Without ontology, you'd implement these manually—if statements scattered throughout your code, hand-rolled validation, duplicate queries. Ontology centralizes this logic in your schema. ## How It Works TypeGraph treats semantic relationships between types as **meta-edges**—edges at the type level rather than instance level: ```typescript // Instance edges: relationships between INSTANCES // "Alice knows Bob" const knows = defineEdge("knows"); // Meta-edges: relationships between TYPES // "Employee subClassOf Person" subClassOf(Employee, Person); ``` When you define an ontology, TypeGraph: 1. **Precomputes closures** at store initialization (not query time) 2. **Expands only the query operations that explicitly opt in** (except inverse traversal, whose store default is `"inverse"` and can be changed) 3. **Enforces the documented constraints** when building a registry or writing data It does not run a general reasoner, materialize implied edges, substitute properties between types, or automatically expand every query. ## Verified Support Matrix | Relation / feature | Runtime contract | | --- | --- | | `subClassOf` | Transitive registry closure, write-path endpoint assignability, and opt-in node-query expansion with `includeSubClasses` | | `disjointWith` | Same-ID collision enforcement, propagated through interleaved `subClassOf` and `equivalentTo` closure (`sameAs` remains a deprecated equivalence alias) | | `implies` | Transitive registry closure and opt-in traversal expansion with `expand: "implying"`; endpoints are validated | | `inverseOf` | Single inverse partner, endpoint reversal validation, and traversal expansion with `expand: "inverse"` (the default store setting) | | `equivalentTo` | Registry lookups and graph-merge type reconciliation; no automatic query or property behavior. `sameAs` is folded in as a full alias — the merge type reconciler and the registry treat a `sameAs` declaration identically to `equivalentTo` | | `broader` / `narrower` | Transitive registry introspection only | | `partOf` / `hasPart` | Transitive registry introspection only | | `relatedTo` | Symmetric direct registry introspection through `getRelatedKinds` only | | Type-level `sameAs` | Deprecated name for `equivalentTo` (see above); prefer calling `equivalentTo` directly | | Type-level `differentFrom` | Deprecated and decorative — never enforced instance identity; migrate to the graph-level TypeGraph Identity Profile | | Custom `metaEdge()` properties | Serialized introspection metadata only; custom transitivity, symmetry, inverse, and inference settings are not executed | ## Core Meta-Edges TypeGraph provides a standard set of meta-edges: ```typescript import { subClassOf, broader, narrower, equivalentTo, sameAs, differentFrom, disjointWith, partOf, hasPart, relatedTo, inverseOf, implies } from "@nicia-ai/typegraph"; ``` ### Subsumption (Type Inheritance) **`subClassOf`**: Defines type inheritance where instances of the child are also instances of the parent. ```typescript subClassOf(Podcast, Media); subClassOf(Article, Media); subClassOf(Company, Organization); ``` **Query Behavior:** Subclass expansion is **opt-in** via `includeSubClasses: true`: ```typescript // Without expansion: returns only nodes with kind="Media" const mediaOnly = await store .query() .from("Media", "m") .select((ctx) => ctx.m) .execute(); // With expansion: returns Media, Podcast, AND Article nodes const allMedia = await store .query() .from("Media", "m", { includeSubClasses: true }) .select((ctx) => ctx.m) .execute(); // Results include nodes of kind "Media", "Podcast", and "Article" ``` This is a fundamental difference from traditional ORM inheritance—TypeGraph stores the concrete type (`kind: "Podcast"`) in the database, and expands at query time when requested. ### Hierarchical (Concept Hierarchy) **`broader`** and **`narrower`**: Define conceptual hierarchy without identity. ```typescript broader(MachineLearning, ArtificialIntelligence); broader(DeepLearning, MachineLearning); broader(ArtificialIntelligence, Technology); ``` **Important**: This is different from `subClassOf`. A topic instance of "ML" is related to "AI", but is **not** an instance of "AI". ```typescript // Get all topics narrower than Technology const narrowerTopics = registry.expandNarrower("Technology"); // ["ArtificialIntelligence", "MachineLearning", "DeepLearning", ...] ``` ### Equivalence **`equivalentTo`**: Defines semantic equivalence between types or external IRIs. ```typescript equivalentTo(Person, "https://schema.org/Person"); equivalentTo(Organization, "https://schema.org/Organization"); ``` **`sameAs`** and **`differentFrom`** are deprecated type-level factories. `sameAs` is currently a type-equivalence alias; `differentFrom` is decorative. For durable individual identity, enable the graph-level TypeGraph Identity Profile and use `store.identity`. That ledger deliberately does not provide OWL property substitution or automatic graph-wide query expansion. ### Constraints **`disjointWith`**: Declares that two types cannot share the same ID. ```typescript disjointWith(Person, Organization); disjointWith(Podcast, Article); ``` Disjointness is inherited by subclasses. If `Company subClassOf Organization`, then `disjointWith(Person, Organization)` also makes `Person` and `Company` disjoint. **Effect**: Attempting to create a node that violates disjointness throws `DisjointError`: ```typescript // Create a Person with ID "entity-1" await store.nodes.Person.create({ name: "Alice" }, { id: "entity-1" }); // Throws DisjointError: Person and Organization are disjoint await store.nodes.Organization.create({ name: "Acme" }, { id: "entity-1" }); ``` **Coherence rules**: `disjointWith` cannot contradict the rest of the ontology. A kind disjoint with itself, a kind disjoint with one of its own subclass ancestors, a common subclass of two disjoint parents, and a kind declared both `equivalentTo` and `disjointWith` another are all rejected, including overlaps reached through mixed equivalence/subclass paths. These checks run both when you construct a graph and when a persisted schema is reloaded, so a document written by an older, more permissive version can fail validation on load with a `ConfigurationError` whose details code is `ONTOLOGY_DISJOINT_CONFLICT`. To recover, fix the graph definition and, for a persisted schema, correct the stored document before upgrading (or rewrite it through the previous minor version, which still accepts it). The same construction-and-reload rule applies to the other ontology coherence checks (duplicate relations, hierarchical self-loops and cycles, and inverse-partner uniqueness). ### Composition **`partOf`** and **`hasPart`**: Define compositional relationships. ```typescript partOf(Chapter, Book); hasPart(Book, Chapter); partOf(Episode, Podcast); hasPart(Podcast, Episode); ``` ### Edge Relationships **`inverseOf`**: Declares two edge kinds as inverses of each other. ```typescript inverseOf(manages, managedBy); inverseOf(cites, citedBy); inverseOf(follows, followedBy); ``` **Effect**: You can query in either direction using the registry: ```typescript const inverse = registry.getInverseEdge("manages"); // "managedBy" ``` You can also expand traversals to include inverse edge kinds at query time: ```typescript const relationships = await store .query() .from("Person", "p") .traverse("manages", "e", { expand: "inverse" }) .to("Person", "other") .select((ctx) => ({ other: ctx.other.name, via: ctx.e.kind, })) .execute(); ``` For symmetric relationships, declare an edge as its own inverse: ```typescript inverseOf(collaboratesWith, collaboratesWith); ``` An edge may have only one distinct inverse partner. Every allowed pair must be compatible with a reversed pair in its partner, in both traversal directions, using equal kinds or `subClassOf` assignability. Matching the independent source and target unions is insufficient for source-dependent edges. A self-inverse edge must satisfy the same reversed-pair check against itself. **`implies`**: Declares that one edge kind implies another exists. ```typescript implies(marriedTo, knows); implies(bestFriends, friends); implies(friends, knows); ``` **Effect**: Query for `knows` can include `marriedTo`, `bestFriends`, and `friends` edges: ```typescript const connections = await store .query() .from("Person", "p") .traverse("knows", "e", { expand: "implying" }) .to("Person", "other") .select((ctx) => ctx.other) .execute(); ``` **Endpoint compatibility is required.** `implies(edgeA, edgeB)` only makes sense if every node kind `edgeA` can connect could also, in principle, satisfy `edgeB`'s own domain/range — otherwise `expand: "implying"` would traverse rows whose kinds don't match what the traversal actually asked for. Every allowed pair in `edgeA` must match a single allowed pair in `edgeB`: both endpoints must be assignable — equal, or a `subClassOf` descendant — to their corresponding endpoint in that pair. For [source-dependent targets](/core-concepts#source-dependent-targets), finding the source in one entry and the target in another does not suffice. An incompatible pair (say, `Author -> Paper` implying `Paper -> Topic`) throws `ConfigurationError` wherever the graph is built into a store or committed as a schema version (`createStore`, `createStoreWithSchema`, `store.evolve({ ontology })`) — including relations authored through a graph extension, not just `implies()` calls in code. ## Using the Ontology ### In Graph Definition ```typescript const graph = defineGraph({ id: "knowledge_base", nodes: { ... }, edges: { ... }, ontology: [ // Type hierarchy subClassOf(Podcast, Media), subClassOf(Article, Media), subClassOf(Company, Organization), // Concept hierarchy broader(MachineLearning, ArtificialIntelligence), broader(DeepLearning, MachineLearning), // Constraints disjointWith(Person, Organization), disjointWith(Media, Person), // Composition partOf(Episode, Podcast), // Edge relationships inverseOf(cites, citedBy), implies(marriedTo, knows), ], }); ``` ### Registry Lookups The type registry (accessed via `store.registry`) provides methods to query the ontology: ```typescript const registry = store.registry; // Subsumption registry.isSubClassOf("Podcast", "Media"); // true registry.expandSubClasses("Media"); // ["Media", "Podcast", "Article"] // Hierarchy registry.expandNarrower("Technology"); // ["AI", "ML", "DL", ...] registry.expandBroader("DeepLearning"); // ["ML", "AI", "Technology"] // Constraints registry.areDisjoint("Person", "Organization"); // true registry.getDisjointKinds("Person"); // ["Organization", "Media", ...] // Edge relationships registry.getInverseEdge("cites"); // "citedBy" registry.getImpliedEdges("marriedTo"); // ["knows"] registry.getImplyingEdges("knows"); // ["marriedTo", "bestFriends", "friends"] registry.getRelatedKinds("MachineLearning"); // ["DataScience", ...] ``` ## Custom Meta-Edges Define domain-specific meta-edges for serialized introspection metadata: ```typescript import { metaEdge } from "@nicia-ai/typegraph"; // Custom meta-edge for prerequisite relationships const prerequisiteOf = metaEdge("prerequisiteOf", { transitive: true, inference: "hierarchy", description: "Learning prerequisite (Calculus prerequisiteOf LinearAlgebra)", }); // Custom meta-edge for superseding relationships const supersedes = metaEdge("supersedes", { transitive: true, inference: "substitution", description: "Replacement relationship (v2 supersedes v1)", }); ``` ### Meta-Edge Properties Each custom meta-edge can carry these properties as metadata. In the current release they do **not** make the registry compute a custom closure or make the query builder execute custom inference. Only the built-in relations in the support matrix have runtime behavior. | Property | Type | Description | | ------------ | --------------- | ------------------------- | | `transitive` | `boolean` | A→B, B→C implies A→C | | `symmetric` | `boolean` | A→B implies B→A | | `reflexive` | `boolean` | A→A is always true | | `inverse` | `string` | Name of inverse meta-edge | | `inference` | `InferenceType` | How this affects queries | ### Inference Types For custom meta-edges, `inference` is descriptive metadata for consumers: | Type | Description | | ---------------- | -------------------------------------------- | | `"subsumption"` | Query for X includes instances of subclasses | | `"hierarchy"` | Enables broader/narrower traversal | | `"substitution"` | Can substitute equivalent types | | `"constraint"` | Validation rules | | `"composition"` | Part-whole navigation | | `"association"` | Discovery/recommendation | | `"none"` | No automatic inference | ## Closure Computation TypeGraph precomputes transitive closures at store initialization: ```typescript // subClassOf closure // If: Podcast subClassOf Media, Episode subClassOf Media // Then: expandSubClasses("Media") = ["Media", "Podcast", "Episode"] // implies closure // If: marriedTo implies partneredWith, partneredWith implies knows // Then: getImpliedEdges("marriedTo") = ["partneredWith", "knows"] ``` This makes queries efficient—expansion happens at query compilation time, not execution time. ## Best Practices ### Separate `subClassOf` from `broader` These have different semantics: - `subClassOf`: Type membership (a Podcast instance is also a Media instance) - `broader`: Conceptual relation (ML **relates to** AI, but ML instance ≠ AI instance) ```typescript // CORRECT: Type hierarchy subClassOf(Podcast, Media); // CORRECT: Concept hierarchy broader(MachineLearning, ArtificialIntelligence); // WRONG: Don't mix them // subClassOf(MachineLearning, ArtificialIntelligence); ``` ### Use Disjoint Constraints Prevent impossible combinations: ```typescript // Good: Prevent ID conflicts disjointWith(Person, Organization); disjointWith(Person, Product); disjointWith(Organization, Product); ``` ### Model Edge Hierarchies with Implies ```typescript // Relationship hierarchy: specific → general implies(marriedTo, partneredWith); implies(partneredWith, knows); implies(parentOf, relatedTo); implies(siblingOf, relatedTo); implies(relatedTo, knows); ``` ### Use InverseOf for Bidirectional Queries ```typescript inverseOf(manages, managedBy); inverseOf(follows, followedBy); inverseOf(cites, citedBy); ``` This lets you query efficiently in either direction without duplicating edges. ## API Reference ### Ontology Functions #### `subClassOf(child, parent)` Declares type inheritance. ```typescript function subClassOf(child: NodeType, parent: NodeType): OntologyRelation; ``` #### `broader(narrower, broader)` Declares hierarchical relationship (narrower concept to broader concept). ```typescript function broader(narrower: NodeType, broader: NodeType): OntologyRelation; ``` #### `narrower(broader, narrower)` Declares hierarchical relationship (broader concept to narrower concept). ```typescript function narrower(broader: NodeType, narrower: NodeType): OntologyRelation; ``` #### `equivalentTo(a, b)` Declares semantic equivalence between types or with external IRIs. ```typescript function equivalentTo( a: NodeType | string, b: NodeType | string ): OntologyRelation; ``` #### `sameAs(kindA, kindBOrIri)` Deprecated type-level alias of `equivalentTo`, including the equivalence with external IRIs. Migrate to the graph-level TypeGraph Identity Profile for individual identity. ```typescript function sameAs(kindA: NodeType, kindBOrIri: NodeType | string): OntologyRelation; ``` #### `differentFrom(a, b)` Deprecated decorative type-level relation. Migrate to the graph-level TypeGraph Identity Profile for individual identity. ```typescript function differentFrom(a: NodeType, b: NodeType): OntologyRelation; ``` #### `disjointWith(a, b)` Declares mutual exclusion (types cannot share the same ID). ```typescript function disjointWith(a: NodeType, b: NodeType): OntologyRelation; ``` #### `partOf(part, whole)` Declares compositional relationship (part to whole). ```typescript function partOf(part: NodeType, whole: NodeType): OntologyRelation; ``` #### `hasPart(whole, part)` Declares compositional relationship (whole to part). ```typescript function hasPart(whole: NodeType, part: NodeType): OntologyRelation; ``` #### `relatedTo(a, b)` Declares a symmetric association available through `registry.getRelatedKinds(kind)`. It has no query behavior. ```typescript function relatedTo(a: NodeType, b: NodeType): OntologyRelation; ``` #### `inverseOf(edgeA, edgeB)` Declares edge types as inverses of each other. ```typescript function inverseOf(edgeA: AnyEdgeType, edgeB: AnyEdgeType): OntologyRelation; ``` #### `implies(edgeA, edgeB)` Declares that one edge type implies another exists. ```typescript function implies(edgeA: AnyEdgeType, edgeB: AnyEdgeType): OntologyRelation; ``` Each allowed pair in `edgeA` must be assignable to one allowed pair in `edgeB` (equal, or a `subClassOf` descendant, on both endpoints). Throws `ConfigurationError` when the graph is built into a store or committed as a schema version if they aren't — see [Edge Relationships](#edge-relationships) above. #### `metaEdge(name, options?)` Creates a custom meta-edge for domain-specific relationships. ```typescript function metaEdge( name: string, options?: { transitive?: boolean; symmetric?: boolean; reflexive?: boolean; inverse?: string; inference?: InferenceType; description?: string; }, ): MetaEdge; ``` ### Type Registry API The type registry is available via `store.registry` and provides methods to query the ontology at runtime. #### `isSubClassOf(child, parent)` Checks if a type is a subclass of another. ```typescript registry.isSubClassOf(child: string, parent: string): boolean; registry.isSubClassOf("Podcast", "Media"); // true ``` #### `expandSubClasses(type)` Returns a type and all its subclasses. ```typescript registry.expandSubClasses(type: string): readonly string[]; registry.expandSubClasses("Media"); // ["Media", "Podcast", "Article"] ``` #### `areDisjoint(a, b)` Checks if two types are disjoint. ```typescript registry.areDisjoint(a: string, b: string): boolean; registry.areDisjoint("Person", "Organization"); // true ``` #### `getDisjointKinds(type)` Returns all types disjoint with the given type. ```typescript registry.getDisjointKinds(type: string): readonly string[]; registry.getDisjointKinds("Person"); // ["Organization", "Media", ...] ``` #### `expandNarrower(type)` Returns all types narrower than the given type (via `broader` relationships). ```typescript registry.expandNarrower(type: string): readonly string[]; registry.expandNarrower("Technology"); // ["AI", "ML", "DeepLearning", ...] ``` #### `expandBroader(type)` Returns all types broader than the given type. ```typescript registry.expandBroader(type: string): readonly string[]; registry.expandBroader("DeepLearning"); // ["MachineLearning", "AI", "Technology"] ``` #### `getInverseEdge(edgeType)` Returns the inverse of an edge type. ```typescript registry.getInverseEdge(edgeType: string): string | undefined; registry.getInverseEdge("manages"); // "managedBy" ``` #### `getImpliedEdges(edgeType)` Returns edges implied by an edge type. ```typescript registry.getImpliedEdges(edgeType: string): readonly string[]; registry.getImpliedEdges("marriedTo"); // ["knows"] ``` #### `getImplyingEdges(edgeType)` Returns edges that imply an edge type. ```typescript registry.getImplyingEdges(edgeType: string): readonly string[]; registry.getImplyingEdges("knows"); // ["marriedTo", "bestFriends", "friends"] ``` #### `expandImplyingEdges(edgeType)` Returns an edge type and all edges that imply it. ```typescript registry.expandImplyingEdges(edgeType: string): readonly string[]; registry.expandImplyingEdges("knows"); // ["knows", "marriedTo", "bestFriends", "friends"] ``` # Indexes > Define and create indexes for TypeGraph queries TypeGraph stores node and edge properties in a JSON `props` column. When you filter or order by JSON properties at scale, you typically need **expression indexes** on those JSON paths. TypeGraph includes built-in indexes for common access patterns (lookups by ID, edge traversals, temporal filtering), but application-specific indexes are up to you. The `@nicia-ai/typegraph/indexes` entrypoint provides: - **Type-safe index definitions** for node and edge schemas - **Dialect-specific DDL generation** for PostgreSQL and SQLite - **Drizzle schema integration** so drizzle-kit can generate migrations - **Profiler integration** so recommendations account for indexes you already have ## Quick Start (Drizzle / drizzle-kit) Define your indexes once and pass them into the Drizzle schema factories: ```ts import { defineEdge, defineNode } from "@nicia-ai/typegraph"; import { createPostgresTables } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; import { andWhere, defineEdgeIndex, defineNodeIndex } from "@nicia-ai/typegraph/indexes"; import { z } from "zod"; const Person = defineNode("Person", { schema: z.object({ email: z.string().email(), name: z.string(), createdAt: z.date(), isActive: z.boolean().optional(), }), }); const worksAt = defineEdge("worksAt", { schema: z.object({ role: z.string(), }), }); export const personEmail = defineNodeIndex(Person, { fields: ["email"], unique: true, coveringFields: ["name"], where: (w) => andWhere(w.deletedAt.isNull(), w.isActive.eq(true)), }); export const worksAtRoleOut = defineEdgeIndex(worksAt, { fields: ["role"], direction: "out", where: (w) => w.deletedAt.isNull(), }); // drizzle-kit will include these indexes in generated migrations export const typegraphTables = createPostgresTables( {}, { indexes: [personEmail, worksAtRoleOut], }, ); ``` For SQLite, use `createSqliteTables`: ```ts import { createSqliteTables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; export const typegraphTables = createSqliteTables( {}, { indexes: [personEmail, worksAtRoleOut], }, ); ``` ## Node Indexes `defineNodeIndex(nodeType, config)` creates an index definition for node properties and TypeGraph system columns. **Key options:** - `keys`: ordered node B-tree keys with an explicit direction. Property and system-column keys can be interleaved to match a complete query order. This node-only option is mutually exclusive with `fields` and `keySystemColumns`, and does not define a `bulkFindByIndex` lookup key. - `fields`: JSON property paths used for filtering/ordering (B-tree expression keys). Optional if `keys`, `coveringFields`, or `keySystemColumns` supply the key instead. - `coveringFields`: additional properties frequently selected with the same filters. These become additional index keys to enable index-only reads when combined with smart select. - `keySystemColumns`: system columns (e.g. `"id"`) to include in the key, after the `scope` prefix and before `fields`. See [Keying on system columns](#keying-on-system-columns-keysystemcolumns). Rejects `"id"` combined with `unique: true` — every node's `id` is already unique per row, so a unique index keyed on `id` plus other columns can never enforce a meaningful constraint across those other columns. - `unique`: create a unique index. - `scope`: prefixes index keys with TypeGraph system columns (default is `"graphAndKind"`). - `where`: partial index predicate (portable DSL, compiled per dialect). :::note[Covering fields vs PostgreSQL INCLUDE] TypeGraph properties live inside `props`, so indexes are built on expressions. PostgreSQL `INCLUDE` does not support expressions, so `coveringFields` are implemented as additional index keys rather than an `INCLUDE (...)` clause. Because `coveringFields` become index keys, they: - Increase index size (more key data) - Affect index ordering (can help `ORDER BY`, but changes sort/range behavior) - Must be maintained on writes like any other key ::: Field ordering emits the database's native `NULLS FIRST` / `NULLS LAST` suffix. This keeps the primary sort expression aligned with a matching B-tree expression index. PostgreSQL and SQLite still choose plans from their statistics and the full query shape, so confirm important queries with `EXPLAIN`; an applicable index does not guarantee that every data distribution will use it. ### Directed node keys Use `keys` when the direction and interleaving of the B-tree keys must match a query's complete ordering. Scope columns remain first, keys retain declaration order, and `coveringFields` remain last: ```ts const newestPerson = defineNodeIndex(Person, { keys: [ { field: "createdAt", direction: "desc" }, { system: "id", direction: "asc" }, ], coveringFields: ["name"], }); const page = await store .query() .from("Person", "person") .select((fields) => ({ name: fields.person.name })) .orderBy("person", "createdAt", "desc") .orderBy("person", "id", "asc") .paginate({ first: 50 }); ``` The resulting key order is `(graph_id, kind, createdAt DESC, id ASC, name)`. In the first version, `keys` is supported for node B-tree indexes only and refuses `unique: true`. It intentionally does not participate in `bulkFindByIndex`: that operation probes the equality lookup fields declared through `fields`, while `keys` describes physical ordering. Use a separate legacy `fields` index when the same graph also needs keyed candidate lookup. ### Nested JSON Paths For top-level properties, use the field name: ```ts defineNodeIndex(Person, { fields: ["email"] }); ``` For nested properties inside `props`, use a JSON pointer: ```ts defineNodeIndex(Person, { fields: ["/metadata/priority"] }); ``` You can also pass pointer segments: ```ts defineNodeIndex(Person, { fields: [["metadata", "priority"] as const] }); ``` ### Index Scope Index `scope` controls which TypeGraph system columns are prefixed ahead of your JSON keys: - `"graphAndKind"` (default): prefixes with `(graph_id, kind)` to match most TypeGraph queries. - `"graph"`: prefixes with `graph_id` only (rare; useful for cross-kind queries within a graph). - `"none"`: no system prefix (rare; usually only correct for global queries). ### Keying on system columns (`keySystemColumns`) `fields` and `coveringFields` only ever reference your schema's own properties, and `scope` only ever prefixes `graph_id`/`kind` — neither can put a system column like `id` into the index **key**. Some query shapes need that. A reverse traversal — `.traverse("hasCreator", { direction: "in" }).to("Post", "post")` — compiles a join on the target node's own `id` (`n.id = e.from_id`). If you also filter, sort, or select a prop on that same node (say `creationDate`), a covering index needs `id` in its key to match that join — `graph_id`/`kind` and `creationDate` alone aren't enough, because the index can't be chosen for an `id`-equality join it doesn't cover. ```ts const postRecent = defineNodeIndex(Post, { keySystemColumns: ["id"], coveringFields: ["creationDate"], }); // -> CREATE INDEX ... ON typegraph_nodes (graph_id, kind, id, (props #>> ARRAY['creationDate'])) ``` **Validation:** - Node indexes only. Rejects edge-only system columns (`fromKind` / `fromId` / `toKind` / `toId`). - Rejects a column already implied by `scope` (e.g. `graph_id` when `scope: "graphAndKind"`). - Not supported with `method: "gin" | "trigram"` (same restriction as `coveringFields`). ## Edge Indexes `defineEdgeIndex(edgeType, config)` works the same way as node indexes, with one extra option: - `direction`: `"out" | "in" | "none"` (default `"none"`). When set, the index keys are prefixed with the join key used by traversal queries (`from_id` for `"out"`, `to_id` for `"in"`). This makes it easy to create indexes that match `.traverse()` patterns. **When to use `direction`:** - `"out"`: optimize outbound traversals that join on `from_id` (start node → edges). - `"in"`: optimize inbound traversals that join on `to_id` (end node → edges). - `"none"`: for edge queries not anchored by a traversal join key (less common). ## Partial Indexes (WHERE) Use `where` to create partial indexes with a small, typed predicate DSL. System columns are available (e.g. `deletedAt`, `createdAt`, `fromId`), as well as your schema properties (e.g. `email`, `role`). ```ts import { andWhere, defineNodeIndex } from "@nicia-ai/typegraph/indexes"; const activeEmail = defineNodeIndex(Person, { fields: ["email"], where: (w) => andWhere(w.deletedAt.isNull(), w.isActive.eq(true)), }); ``` ## Covering Indexes To maximize the benefit of [smart select optimization](/performance/overview#smart-select), create indexes that include both the filter columns and selected columns. This enables index-only scans where the database satisfies the entire query from the index. ```ts // Index covers email filter AND name selection const personEmailWithName = defineNodeIndex(Person, { fields: ["email"], coveringFields: ["name"], where: (w) => w.deletedAt.isNull(), }); ``` **Generated PostgreSQL:** ```sql CREATE INDEX idx_person_email_name ON typegraph_nodes (graph_id, kind, ((props #>> ARRAY['email'])), ((props #>> ARRAY['name']))) WHERE deleted_at IS NULL; ``` **Generated SQLite:** ```sql CREATE INDEX idx_person_email_name ON typegraph_nodes (graph_id, kind, json_extract(props, '$.email'), json_extract(props, '$.name')) WHERE deleted_at IS NULL; ``` :::caution[PostgreSQL: JSONB expression indexes don't get a true Index Only Scan] A covering index like the one above lets PostgreSQL avoid a full **table scan**, but not necessarily a **heap fetch** per matching row. PostgreSQL's `Index Only Scan` optimization — skip the heap entirely when the index already has everything the query needs — does not extend to JSONB extraction expressions (`props #>> ARRAY[...]`), only to real stored columns. Even a correctly-shaped `coveringFields` index shows as a plain `Index Scan` in `EXPLAIN`, not `Index Only Scan`, and PostgreSQL still visits the heap row for every match. This matters most for high-fan-in ranked reads — e.g. "each of N friends' most recent post, top 10 overall" — where a query visits many candidate rows and discards most of them after sorting. If `EXPLAIN (ANALYZE, BUFFERS)` shows most of a query's cost coming from a repeated `Index Scan` with high `Buffers: shared hit` relative to the rows actually returned, this is likely why — confirmed against a real workload at real scale, not a theoretical concern. TypeGraph doesn't currently offer a built-in way to materialize a hot prop as a real column (this is an active area of investigation). Until then, if this is a genuine hot-path bottleneck on PostgreSQL, the workaround is maintaining your own [stored generated column](https://www.postgresql.org/docs/current/ddl-generated-columns.html) alongside `props` with a plain index over it — a real column does get `Index Only Scan`, confirmed via `EXPLAIN (ANALYZE, BUFFERS)` showing `Heap Fetches: 0` (run `VACUUM ANALYZE`, not just `ANALYZE`, after backfilling — the visibility map needs to be current before PostgreSQL will prove it immediately). We haven't specifically verified whether SQLite's JSON-extraction covering indexes have the same limitation for this query shape. ::: ## Batched Index Lookup (`bulkFindByIndex`) `store.nodes..bulkFindByIndex(indexName, items, options?)` takes many in-memory records and returns the live nodes that share each record's declared **index key** — batched candidate retrieval for import reconciliation, dedup-candidate discovery, and joining incoming records against the graph by a declared composite key. ```ts const candidates = await store.nodes.Person.bulkFindByIndex("person_active_name", [ { props: { isActive: true, name: "Ana" } }, { props: { name: "Bo" } }, // missing isActive → matches stored null ]); // readonly Node[][] — one bucket per input, ordered by node id ``` Semantics: - **One bucket per input**, in input order; empty input returns `[]`. The index may be non-unique, so each bucket is a (possibly empty) array — this is candidate retrieval, not a uniqueness guarantee. For unique lookups prefer `bulkFindByConstraint` (backed by the uniqueness side-table). - TypeGraph computes the lookup key from **`index.fields` only** (JSON-pointer extraction, reusing the index's own extraction expressions). `keys`, `coveringFields`, and `keySystemColumns` are not part of the probe key. An index declared without `fields` has nothing to probe by and throws `ConfigurationError`. - The index's partial `where` is applied in SQL to **stored** rows only; probes carry index-field values, nothing else. Only the indexed fields are validated — full records are not required. - A missing/`undefined` indexed field matches stored `NULL` (null-safe equality). Live, non-soft-deleted nodes only. - `options.limitPerInput` caps each bucket (ordered by node id); unbounded by default — no silent truncation. A non-positive value throws `ValidationError`. An unknown index name throws `NodeIndexNotFoundError`. A non-scalar probe value throws `ValidationError`. On backends that support SQL window functions the cap is applied in-database (`ROW_NUMBER()`); on backends without them (`capabilities.windowFunctions: false`) it degrades to an in-memory cap after fetching the matching ids — same result, but it transfers all matching ids for low-selectivity keys. - **Key field types:** string, number, and boolean keys are supported. **Date-typed key fields are not** — they throw `ConfigurationError`, because SQLite compares stored ISO text byte-wise while PostgreSQL compares `timestamptz` instants, so the same instant in different ISO forms would match on one backend but not the other. Use a string-encoded key, or `store.query(...).where(...)` for date predicates. The lookup is correct whether or not the physical index has been materialized; materialize it (see below) for the query planner to actually use it. Null-safe predicates may be less reliably index-accelerated than plain equality. ## Choosing the Right Index Type TypeGraph's `defineNodeIndex` / `defineEdgeIndex` generate **B-tree expression indexes** — the right choice for scalar equality, range, and ordering queries. But JSON properties can also hold arrays and objects, which need different index strategies. | Data shape | Query pattern | Index type | TypeGraph utility? | | -------------------------------------- | ---------------------------------------------- | ------------------------ | ------------------------------------ | | Scalar (`string`, `number`, `boolean`) | `eq()`, `gt()`, `in()`, `orderBy()` | B-tree expression | Yes — `defineNodeIndex` | | Array of scalars | `contains()`, `containsAll()`, `containsAny()` | GIN (PostgreSQL) | No — use raw SQL | | Nested object | `hasKey()`, `pathEquals()`, `pathContains()` | GIN or B-tree expression | Partially — B-tree on specific paths | ### B-tree expression indexes (scalar properties) Best for equality, range, sorting, and prefix matching on individual JSON fields. This is what `defineNodeIndex` and `defineEdgeIndex` generate. ```ts // Good for: .whereNode("p", (p) => p.email.eq("...")) defineNodeIndex(Person, { fields: ["email"] }); // Good for: .orderBy("p", "createdScore", "desc") defineNodeIndex(Person, { fields: ["createdScore"] }); ``` ### GIN indexes (array containment — PostgreSQL only) TypeGraph compiles array predicates to PostgreSQL's JSONB containment operator over the field's extraction expression — `(props #> ARRAY['tags']) @> $1`. Declare a containment index with `method: "gin"` and TypeGraph emits the matching **expression GIN** (`jsonb_path_ops`): ```typescript const personTags = defineNodeIndex(Person, { fields: ["tags"], method: "gin", }); // materialized via store.materializeIndexes(): // CREATE INDEX ... USING GIN (("props" #> ARRAY['tags']) jsonb_path_ops); ``` The index accelerates all containment predicates on that field: ```typescript // contains: does the tags array include "typescript"? .whereNode("p", (p) => p.tags.contains("typescript")) // containsAll: does it include BOTH "typescript" AND "graphql"? .whereNode("p", (p) => p.tags.containsAll(["typescript", "graphql"])) // containsAny: does it include "typescript" OR "graphql"? .whereNode("p", (p) => p.tags.containsAny(["typescript", "graphql"])) ``` :::caution[Whole-column `GIN (props)` does not work] PostgreSQL matches expression indexes structurally. A whole-column `CREATE INDEX ... USING GIN (props)` serves `props @> …` — **not** the per-field `(props #> ARRAY['tags']) @> …` expressions TypeGraph compiles, so such an index is never used. Declare `method: "gin"` per field (or hand-write the same expression form) instead. ::: :::note[SQLite] SQLite has no GIN equivalent; `materializeIndexes()` reports gin/trigram declarations as `skipped` there. Array containment on SQLite uses `json_each()` scans, which can't be indexed. ::: ### Trigram indexes (substring and case-insensitive matching — PostgreSQL only) `contains` / `startsWith` / `endsWith` / `ilike` on string fields compile to `ILIKE` on PostgreSQL, which a B-tree can never serve for infix patterns. Declare `method: "trigram"` and TypeGraph emits an expression GIN with `gin_trgm_ops` (installing the `pg_trgm` extension on first materialization): ```typescript const personName = defineNodeIndex(Person, { fields: ["name"], method: "trigram", }); // CREATE INDEX ... USING GIN (("props" #>> ARRAY['name']) gin_trgm_ops); .whereNode("p", (p) => p.name.contains("smith")) // served by the index .whereNode("p", (p) => p.name.ilike("%SMITH%")) // also served ``` On SQLite these declarations are `skipped` — SQLite's substring-search story is [FTS5 fulltext](/fulltext-search) via `searchable()` fields. GIN-family methods take exactly one field and don't support `unique`, `coveringFields`, or `where`; the query's `graph_id` / `kind` filters apply as residual conditions over the index's candidate rows. ### Combining B-tree and GIN For kinds where you filter on both scalar fields (equality, range) and array or substring predicates, declare both index types: ```typescript const personEmail = defineNodeIndex(Person, { fields: ["email"] }); const personTags = defineNodeIndex(Person, { fields: ["tags"], method: "gin" }); ``` PostgreSQL's query planner can use both indexes together via a BitmapAnd scan when a query filters on both a scalar field and an array field. ## Generating SQL (No drizzle-kit) If you manage migrations yourself, generate DDL snippets: ```ts import { generateIndexDDL } from "@nicia-ai/typegraph/indexes"; const sql = generateIndexDDL(personEmail, "postgres"); // → CREATE INDEX ...; ``` ## Verifying Index Usage Use `EXPLAIN ANALYZE` to verify your indexes are being used: ```sql -- PostgreSQL EXPLAIN ANALYZE SELECT props #>> ARRAY['email'], props #>> ARRAY['name'] FROM typegraph_nodes WHERE graph_id = 'my_graph' AND kind = 'Person' AND deleted_at IS NULL AND (props #>> ARRAY['email']) = 'alice@example.com'; -- SQLite EXPLAIN QUERY PLAN SELECT json_extract(props, '$.email'), json_extract(props, '$.name') FROM typegraph_nodes WHERE graph_id = 'my_graph' AND kind = 'Person' AND deleted_at IS NULL AND json_extract(props, '$.email') = 'alice@example.com'; ``` Look for "Index Scan" or "Index Only Scan" (PostgreSQL) or "USING INDEX" (SQLite) in the output. ## Profiler Integration Pass your existing indexes to the [Query Profiler](/performance/profiler) so recommendations focus on what you *don't* have: ```ts import { QueryProfiler } from "@nicia-ai/typegraph/profiler"; import { toDeclaredIndexes } from "@nicia-ai/typegraph/indexes"; const profiler = new QueryProfiler({ declaredIndexes: toDeclaredIndexes([personEmail, worksAtRoleOut]), }); ``` ## Limitations - The default `defineNodeIndex` / `defineEdgeIndex` method generates B-tree expression indexes for **scalar** properties (`string`, `number`, `boolean`, `Date`). Array containment and substring matching are served by [`method: "gin"`](#gin-indexes-array-containment--postgresql-only) and [`method: "trigram"`](#trigram-indexes-substring-and-case-insensitive-matching--postgresql-only) declarations. - GIN-family methods are PostgreSQL-only; `materializeIndexes()` reports them as `skipped` on SQLite, which has no equivalent for JSON containment acceleration (substring search on SQLite is served by FTS5 fulltext). - Embedding fields live in per-`(graphId, kind, field)` vector tables (`tg_vec_*`) and are indexed through `store.materializeIndexes()` (pgvector builds an HNSW / IVFFlat ANN index; sqlite-vec and libSQL report `skipped`/build their own). See [Semantic Search](/semantic-search). ## System Indexes TypeGraph ships a set of **system indexes** on its own relations (nodes, edges, and the recorded history tables) — the traversal, listing, temporal-validity, and bare-id access paths every compiled query relies on. They are declared once (`SYSTEM_INDEX_DECLARATIONS`, exported for inspection), and both dialects' schemas derive from that single list, so SQLite and PostgreSQL always carry the same set. You normally never manage them: fresh databases get them at bootstrap, and `createStoreWithSchema()` brings an already-initialized database up to the running library version's set on boot (with `CREATE INDEX CONCURRENTLY` on PostgreSQL). Deployments that boot without `createStoreWithSchema` can run `store.materializeSystemIndexes()` once after a library upgrade; it shares `materializeIndexes()`'s status tracking, drift signatures, and concurrent-build claim protocol, and settles with no DDL when everything already exists. ## Next Steps - [Performance Overview](/performance/overview) — Best practices, N+1 prevention, batch patterns - [Query Profiler](/performance/profiler) — Automatic index recommendations # Performance Overview > Understanding the performance characteristics of TypeGraph TypeGraph is designed to be a high-performance, low-overhead layer on top of your relational database. By leveraging the power of modern SQL engines (SQLite and PostgreSQL) and precomputing complex relationships, TypeGraph ensures that your knowledge graph scales with your application. ## Performance Philosophy 1. **One Fluent Query, One Statement**: Every fluent query — including multi-hop traversals — compiles to a single SQL statement, so its statement count never grows with the size of the graph. This prevents compiler-generated N+1 work inside that query; application code can still create an N+1 by issuing separate reads in a loop. (Compilation, not execution: a query whose selective-field mapping falls back re-runs as a full fetch, costing a second statement. See [Batch reads](#batch-reads).) 2. **Precomputed Ontology**: Transitive closures, subclass hierarchies, and edge implications are computed once at schema initialization, not during every query. 3. **Batching & Transactions**: Bulk collection APIs minimize round-trips for writes. On the read side that job belongs to the query compiler — `store.batch()` only caps concurrency at one query in flight, it does not reduce round trips and it is not a snapshot. 4. **Zero-Cost Abstractions**: Type safety and ontological reasoning add no measurable runtime overhead. ## N+1 Prevention A common performance problem in ORMs is the N+1 query: you fetch N entities, then issue one query per entity to load related data. TypeGraph's fluent query compiler eliminates that pattern inside one graph-shaped query; it cannot eliminate separate collection reads issued by application code. Every query — regardless of how many traversals it chains — compiles to a **single SQL statement** using Common Table Expressions (CTEs). Each traversal step becomes a CTE that joins against the previous one: ```typescript // This compiles to ONE SQL statement, not 3 separate queries const results = await store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq("Alice")) .traverse("worksAt", "employment") .to("Company", "c") .traverse("locatedIn", "location") .to("City", "city") .select((ctx) => ({ person: ctx.p.name, company: ctx.c.name, city: ctx.city.name, })) .execute(); ``` The generated SQL looks like: ```sql WITH cte_p AS ( SELECT ... FROM typegraph_nodes WHERE graph_id = ? AND kind IN ('Person') AND ... ), cte_employment AS ( SELECT ... FROM typegraph_edges e JOIN typegraph_nodes n ON ... WHERE e.graph_id = ? AND ... ), cte_location AS ( SELECT ... FROM typegraph_edges e JOIN typegraph_nodes n ON ... WHERE e.graph_id = ? AND ... ) SELECT ... FROM cte_p JOIN cte_employment ON ... JOIN cte_location ON ... ``` This holds for all query types: - Multi-hop traversals (N CTEs, 1 statement) - [Recursive traversals](/queries/recursive) (WITH RECURSIVE, 1 statement) - Aggregations with traversals (CTEs + GROUP BY, 1 statement) - [Set operations](/queries/combine) (UNION/INTERSECT/EXCEPT of CTEs, 1 statement) The fluent query needs no dataloader for that joined read because the database handles its entire join graph in one execution. Separate reads can still form an N+1; use a traversal, `batchOnce()`, `neighbors()` / `countNeighbors()`, or `subgraph()`. Inside `batchOnce()`, its batch-scoped `read` builder creates composable versions of the set-oriented reads when unlike result shapes must share one statement. Chunked collection reads remain useful for homogeneous ID and endpoint sets. ## Batch Write Patterns ### Remote edge convergence For a latency-sensitive `getOrCreateByEndpoints()` path, declare the canonical identity on the edge registration instead of supplying an ad hoc `matchOn` list at each call: ```typescript const graph = defineGraph({ id: "work", nodes: { Person: { type: Person }, Company: { type: Company } }, edges: { worksAt: { type: worksAt, from: [Person], to: [Company], cardinality: "many", matchIdentity: { name: "employment", fields: ["role"] }, }, }, }); await store.edges.worksAt.getOrCreateByEndpoints(alice, acme, { role: "engineer", }); ``` On a schema-managed bundled root backend (including D1 and neon-http), with claims, sidecars, history, and revision work absent, the default `ifExists: "return"` path combines the schema fence, endpoint validation, unique arbitration, and created/found result into one statement. A typical Neon WebSocket miss therefore falls from roughly five sequential requests to one. The found path is also one request, but the PostgreSQL implementation performs a no-op conflict update: it takes a row lock and can create write amplification, so it is not a substitute for a hot read cache. Dynamic call-level `matchOn`, constrained single-edge writes, history/revision stores, caller-owned transactions, and custom backends without the matching semantic program retain the transactional path required by their additional contracts. In particular, writes made through `store.transaction()` remain on the interactive path and do not receive the root-path exchange-count reduction. Eligible durable bulk endpoint convergence now submits one closed native atomic exchange: the durable identity arbiter, endpoint validation, and ordered created/found results are all resolved by the program. This removes the outside probe, the transaction open/commit, and the per-item write legs for the eligible shape. The fallback bulk path still discovers exact directed endpoint pairs in set-oriented bind-budget chunks and retains its transactional contract. See [`getOrCreateByEndpoints`](/schemas-stores#getorcreatebyendpointsfrom-to-props-options) for field restrictions, migration rules, and PostgreSQL retry guidance. The one-request path is an authoritative command, not a general Store batch: the backend statement owns endpoint validation, durable-key arbitration, and the created/found result. Static adapter batches (including multi-row inserts) are separate internal optimizations and do not turn a sequence of public Store calls into one atomic operation. Use `store.transaction(...)` when several operations—including claims, Operational Identity, history, or revision sidecars—must commit together. Undeclared dynamic `matchOn` convergence keeps that interactive-transaction requirement; only a schema-declared durable `matchIdentity` can qualify for the one-statement root command. Bulk endpoint convergence has a narrower native envelope than direct edge inserts. A schema-declared durable `matchIdentity` with `cardinality: "many"`, the declaration's match fields, default `ifExists: "return"`, and no temporal mutation qualifies on an exact bundled root. The libSQL transport inventory records one client `batch` submission and zero client `execute` calls for a multi-item eligible call; this is a submission-count measurement, not a wall-clock benchmark. Dynamic `matchOn`, `ifExists: "update"`, constrained cardinality, temporal options, caller transactions, derived backends, custom backends without a registered durable-convergence family, and history/revision stores intentionally retain the fallback path. Outside the native envelope, an all-live `ifExists: "return"` batch is the read-only exception: every backend may return that result from its single set-oriented root read without opening a confirmation transaction. Inside the native envelope, the authoritative upsert program runs first. An all-live call still returns `"found"` in one exchange, but the conflict-update mechanism may take incumbent-row locks and produce write amplification. Any batch outside that envelope that may create, resurrect, or update requires the complete transactional fallback. If an otherwise eligible batch resolves a tombstoned identity, the native attempt rolls back and transactionless convergence refuses with the typed `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` (`edgeMatchKeyConvergence`) error. Use a transaction-capable backend when resurrection must merge partial properties through the graph's Zod update schema. Cloudflare D1's 100-parameter budget admits at most seven **unique durable identities** in the native convergence program; duplicate inputs reuse their first identity and do not consume another program entry. Above that ceiling, an all-live return completes through the read-only one-read path described above. A batch that needs a write uses the portable fallback on a transaction-capable backend and refuses on a transactionless D1 root rather than splitting one atomic convergence contract across multiple submissions. Bundled PostgreSQL roots using a recognized session-capable driver, Neon HTTP, Cloudflare D1, and libSQL also expose native write programs for eligible ingestion calls. Schema-managed `nodes.bulkInsert(items)` and `nodes.bulkCreate(items)` run one schema-fenced atomic program when the node has no Operational Identity, history, or revision work. The program composes each member's complete advertised uniqueness/disjointness claim set with its fulltext/vector projection transitions. Same-kind and hierarchy-wide uniqueness, generated and caller IDs, and mixed claim families share this one program boundary. Neon HTTP, D1, and libSQL submit that program as one transport batch. Session-capable PostgreSQL runs its statements on one pinned Drizzle transaction, with one SQL statement per bind-budget chunk. The batch may use generated IDs, caller-supplied IDs, or a mixture of both; `bulkCreate()` restores its rows to input order. Claim work is chunked by member inside the same atomic submission, so Cloudflare D1 no longer has a batch-wide claimed-member ceiling. Its 100-parameter budget leaves 87 claim-input binds per member after the row and fence: a canonical claim costs six, each legacy hierarchy-wide uniqueness probe costs nine, and each legacy disjointness probe costs six. Custom executors should call the exported `atomicNodeClaimInputCost()` owner rather than reproduce this formula. A member beyond that complexity, identity-enabled and history/revision-tracked shapes, and other unsupported node work retain the existing transaction or fallback path. This per-member budget is distinct from the seven unique durable edge identities admitted by the convergence program above. A successful program needs no diagnostic reads. A refused claim first rolls the entire native batch back, then uses committed-state fence and claim reads to recover the same typed error as the portable path. Bundled backends diagnose the complete refused input with set-oriented node and uniqueness reads rather than one probe per member. Custom backends without those batch reads advance through 32-member concurrency windows until the earliest refusal can be selected in input order; when no claim explains the rollback, every input is covered before the honest terminal error. These failure-only reads are not part of the successful-write RTT count. Both `edges.bulkInsert(items)` and `edges.bulkCreate(items)` use the same schema-fenced atomic program when the store has no history or revision capture. The program validates live endpoints, arbitrates declared durable `matchIdentity`, and maintains `one`, `unique`, and `oneActive` cardinality claims at the write boundary. It rolls back the whole call when any bind-budget chunk or constraint sidecar fails and restores `bulkCreate()` results to input order. Eligible `bulkDelete()` calls use the same mutation-program boundary. Direct edge batches on bundled roots submit one schema-fenced atomic program; the statement refuses an ID owned by another edge collection and rolls back every chunk. Restricted node batches also submit one exchange and release uniqueness and disjointness claims owned by the tombstoned rows in that program. The node statement rechecks connected live edges at the write boundary, so a restricted delete cannot race an earlier application-side probe. Cascade, disconnect, projection, identity, captured, derived-backend, unregistered custom-backend, and caller-transaction shapes keep the interactive path. The same exact-root programs serve eligible singleton `update()` and `delete()` calls without changing their per-operation hook contract. On an interactive PostgreSQL root, the guarded mutation owns a short transaction; on the single-submission transports it remains one batch. A node update with no unique/identity sidecars (disjointness has no update-side transition) may carry its fulltext/vector replacements in the same program. A `cardinality: "many"` edge update with no durable match identity, performs one authoritative preimage read, validates and merges properties in TypeGraph, then submits one guarded atomic update. Every direct edge delete, and a restricted node delete with supported claim cleanup, performs its existing live-row gate and then submits one guarded atomic delete. Missing or tombstoned deletes remain hook-free read-only no-ops. Temporal mutations, node cascade/disconnect, history/revision capture, ordinary derived backends, and unregistered custom families retain the complete portable transaction path. The guarded update converges optimistically rather than holding a transaction lock across its read and write. A one-row update gets four attempts; sustained same-row contention can still end in `DatabaseOperationError`, while larger resolved batches retain their two-attempt budget to bound retry cost. These operations are registered through an exact-resource mutation execution profile. Create and delete are **closed programs**: validation and arbitration can be expressed by the submitted SQL itself. Node `bulkReplaceById()` is also a closed program: every item supplies a complete document, so eligible bundled roots submit missing-row creation, live replacement, tombstone resurrection, claim release/acquisition, and fulltext/vector transitions in one read-free atomic exchange. Live rows preserve their validity windows; resurrected rows receive a fresh stamped window. Operational Identity and history/revision capture retain the portable path. `bulkUpsertById()` is deliberately different. It first reads authoritative stored properties, then merges and validates them before its write set is known. On bundled serverless roots, a distinct-ID batch with no claims, Operational Identity, durable edge match identity, temporal mutation, history, or revision capture submits its resolved mutation set, including node fulltext/vector replacements, as one atomic exchange. An eligible set on bundled session-capable PostgreSQL stays inside its exact open transaction and dispatches the same reviewed program through a separate registration bound to that pinned transaction. Its `applied | unsupported` result is explicit: `unsupported` is returned only before any program SQL runs, after which the collection enters the complete portable path. Update-only sets use one guarded set update. A set containing both fresh creates and live updates carries both legs plus a zero-write terminal postimage assertion in the same native batch; an incomplete preimage deliberately aborts the batch before any create can commit. Repeated IDs, resurrections, temporal changes, claims, edge sidecars, and unregistered transaction sessions retain the consolidated interactive path. The exact session may be collection-opened, supplied by `store.transaction()`, or adopted from the caller. A session program uses a savepoint so a deliberate database refusal can be rolled back and diagnosed without poisoning the caller's surrounding PostgreSQL transaction. Measured at the libSQL transport boundary, eligible plain node `bulkInsert()` and `bulkCreate()` calls each submit one exchange for generated, caller-supplied, and mixed ID batches. A one-chunk unconstrained edge batch remains 1 exchange, a durable-match batch drops from 6 transport submissions to 1 atomic exchange, and a cardinality-constrained batch drops from 8 transport submissions to 1 atomic exchange. Eligible one-chunk node and edge `bulkDelete()` calls likewise submit 1 atomic exchange instead of a transaction plus per-row probes/writes. On the portable edge path, one batched authoritative read plus one set-based soft delete replaces the former two statements per input. Eligible `nodes.bulkReplaceById()` calls submit 1 atomic exchange with no preimage read, including claim and projection sidecars. Eligible update-only and mixed create/update node and edge `bulkUpsertById()` calls whose preimages fit one bind-budget read submit 2 exchanges: one batched preimage read and one atomic mutation exchange. Mixed sets previously required separate create and update submissions after the read, so they fall from 3 exchanges to 2 (33%). D1's 100-parameter budget admits 17 node mutations or 6 edge mutations per mutation statement. One D1 native submission accepts at most 512 node members or 187 edge members; larger sets fail closed to the portable path rather than constructing an unbounded transport request. Within that ceiling, the mutation program chunks statements inside one atomic batch, and a terminal postimage assertion for every chunk rolls the complete submission back if any guarded member moved. When the preimage read also exceeds its bind budget, it costs one read exchange per read chunk plus the single atomic mutation submission. Other backends derive their per-statement chunk size from their declared parameter budget and retain an absolute 512-member submission ceiling. The native exchange still contains the SQL statements needed for inserts and node projection sidecars; it groups fulltext and per-vector-slot transitions into set statements and submits them as one transaction so they do not each pay network latency. The exact previous count varies by driver and endpoint shape. The program also proves the exact durable contribution-marker identity and strategy signature in that submission. A newly constructed backend therefore pays no separate cold marker read before an eligible projected write. Missing, stale, failed, or unmaterialized evidence aborts the whole submission; the failure path then reads committed marker state to recover the existing typed contribution diagnostic. The marker proof is an additional SQL statement inside that atomic submission, so the optimization removes a network exchange rather than all server-side proof work. Schemas with more marker identities may require more than one proof statement within the same submission. On session-capable PostgreSQL, the preimage read and mutation program remain in one collection-owned transaction. The program has a bounded number of statements independent of row count within its bind ceiling; it is not described as a single network exchange because wire-protocol drivers execute those statements on the pinned session. The gain is removal of the portable per-family/per-member write-plan fan-out while retaining whole-call rollback. These are internal execution optimizations, not a public Store batch API. History/revision capture, ordinary derived or custom backends, dynamic get-or-create convergence, and other unsupported shapes retain their transaction or fallback behavior. Eligible mixed sets inside a bundled PostgreSQL transaction are the narrow session-bound exception. ### Single vs bulk operations For small numbers of writes, individual `create()` calls inside a transaction are fine. For larger volumes, use the bulk collection APIs — they use multi-row INSERTs and handle parameter chunking internally. | Method | Returns results | Use case | | ----------------------------------------- | --------------- | ---------------------------------------------------- | | `bulkCreate(items)` | Yes | Need created nodes back | | `bulkInsert(items)` | No | Maximum throughput ingestion | | `bulkUpsertById(items)` | Yes | Idempotent import (create or update by ID) | | `bulkReplaceById(items)` | Yes | Idempotent complete-document replacement by ID | | `bulkDelete(ids)` | No | Mass soft-delete | | `trustedImportGraphStream(store, chunks)` | No | Fastest initial load into a fresh dedicated database | The collection APIs remain the default: they validate data and maintain every configured constraint and sidecar. For a one-time initial load whose producer already guarantees those invariants, the distinct [`trustedImportGraphStream`](/interchange#trusted-initial-import) surface uses a single transaction, engine-native inserts, and deferred secondary-index builds. It intentionally rejects non-empty databases and graph features it cannot yet maintain. ### PostgreSQL parameter limits PostgreSQL's protocol can encode 65,535 bind parameters, while TypeGraph uses a portable 65,533-parameter budget across its bundled drivers. Bulk operations are automatically chunked to stay within that budget: - Node inserts: ~7,200 per chunk (9 params per node) - Edge inserts: ~4,680 per chunk (budgeted at 14 params per durable edge) You don't need to chunk manually — pass arrays of any size and TypeGraph handles the rest. ### Transaction wrapping On a transaction-capable backend, each bulk method call is atomic across all of its bind-budget chunks. Eligible plain `nodes.bulkInsert()` and `nodes.bulkCreate()` calls, eligible plain node `bulkDelete()` calls, and direct edge `bulkInsert()` / `bulkCreate()` / `bulkDelete()` calls also provide whole-call atomicity on bundled transactionless roots through one native atomic exchange, including durable-match and cardinality-constrained edge batches. Other bulk shapes on a transactionless root either refuse when their contract requires a fence or use their documented non-atomic path. A certified atomic SQL program is available only to operations whose closed statement contract has been proven by the backend conformance runner; it does not make arbitrary Store calls atomic. `store.transaction()` refuses before invoking its callback on a transactionless root; it never presents sequential writes as atomic. To commit several bulk calls as one unit on a transaction-capable backend, wrap them in a transaction: ```typescript // Atomic: all-or-nothing for the entire import await store.transaction(async (tx) => { await tx.nodes.Person.bulkCreate(people); await tx.nodes.Company.bulkCreate(companies); await tx.edges.worksAt.bulkCreate(employments); }); ``` Without the wrapping transaction, a failure in a later bulk call leaves earlier calls committed. ### Choosing the right pattern ```typescript // Small batch (< 100 items): individual creates in a transaction are fine await store.transaction(async (tx) => { for (const person of people) { await tx.nodes.Person.create(person); } }); // Medium batch (100–10,000 items): bulkCreate const created = await store.nodes.Person.bulkCreate(people); // Large batch (10,000+ items): bulkInsert (no result allocation) await store.nodes.Person.bulkInsert(people); // Idempotent import: bulkUpsertById (creates or updates by ID) await store.nodes.Person.bulkUpsertById(itemsWithIds); // Fresh dedicated database + already-validated producer: await trustedImportGraphStream(store, interchangeChunks); ``` ### Batch sizing for large multi-call imports For a dataset too large for a single `bulkInsert`/`bulkCreate` call (e.g., streaming rows from a file in a loop), the *size* of each call matters, not just the total row count. Each call is its own transaction, and — per the default [`autoRefreshStatistics`](/backend-setup#refreshing-planner-statistics-after-bulk-loads) — can trigger a planner-statistics refresh on its own. In a large-scale bulk-load benchmark, batches of ~2,000 rows per call were consistently ~25-30% slower per row than batches of ~20,000+: fewer, larger calls amortize both the per-call transaction commit and the statistics refresh across more rows. Prefer batch sizes in the tens of thousands when looping over many calls for a large import, and consider `autoRefreshStatistics: false` plus one `store.refreshStatistics()` call after the loop if per-call refreshes still dominate. ### Batch reads `getByIds()` on node and edge collections uses `SELECT ... WHERE id IN (...)` — one statement per bind-limit chunk, so a single statement for id counts under the limit — instead of N individual queries. Results are returned in input order with `undefined` for missing entries. ```typescript const [alice, bob] = await store.nodes.Person.getByIds([aliceId, bobId]); ``` For multiple independent embeddable reads with different shapes and filters, use [`store.batchOnce()`](/schemas-stores#batch-query-execution) to execute exactly one statement: ```typescript const [activeUsers, recentOrders] = await store.batchOnce(() => [ store .query() .from("User", "u") .whereNode("u", (u) => u.status.eq("active")) .select((ctx) => ({ id: ctx.u.id, name: ctx.u.name })), store .query() .from("Order", "o") .select((ctx) => ({ id: ctx.o.id, total: ctx.o.total })) .orderBy("o", "createdAt", "desc") .limit(20), ]); ``` Runtime arrays of compatible `read.subgraph()` calls can opt into shared traversal and hydration with `store.batchOnce(build, { shareSubgraphs: true })`. This helps overlapping, payload-heavy neighborhoods; it adds overhead for membership and per-request reconstruction, so the independent one-statement plan remains the default. Measure the real root overlap and projection rather than enabling sharing universally. Compare the [concrete sharing examples](#choosing-shared-subgraphs) before enabling the option. The callback's batch-scoped builder composes set-oriented reads in the same call: ```typescript const [latest, versionCount, detail] = await store.batchOnce((read) => [ read.neighbors(document, { edges: ["hasVersion"], orderBy: { by: "node", field: "sequence", direction: "desc" }, limit: 1, }), read.countNeighbors(document, { edges: ["hasVersion"] }), read.subgraph(document.id, { edges: ["hasSection"], maxDepth: 2 }), ]); ``` The returned collection may be a runtime-sized array, including `roots.map(...)`. Empty arrays execute zero statements; every nonempty array, including a singleton, executes exactly one or is refused before execution. The portable ceiling is 500 reads and the final statement must fit the backend's bind-parameter budget. Since all member rows are returned in materialized JSON envelopes, use bounded projections and limits; `batchOnce()` does not stream, predict payload size, or impose a response-byte cap. When a request needs several independent neighborhoods, prefer one runtime batch over awaiting `store.subgraph()` in a loop, especially when the database is remote: ```typescript const neighborhoods = await store.batchOnce((read) => roots.map((root) => read.subgraph(root.id, { edges: ["knows", "worksAt"], maxDepth: 2, project: { nodes: { Person: ["name"], Company: ["name"] }, edges: { knows: [], worksAt: ["role"] }, }, }), ), ); ``` This changes several database round trips into one. Each subgraph still has its own recursive CTE and hydration work: By default, `batchOnce()` does not merge roots, share traversal, or guarantee less database CPU. For one large closure, the direct backend-tuned `store.subgraph()` path can be faster. Use `batchOnce()` when round-trip latency across several independent, bounded subgraphs is the cost to remove, and measure both forms when database work dominates. When the response needs a list of matching child records per parent, use a relation grouped by the parent key and [`expr.collect()` with named scalar fields](/queries/relations#ordered-collections). That aggregates the selected fields into ordered records in one query. Apply [`topPerPartition()`](/queries/relations#top-n-per-parent) before grouping when each parent needs only its highest-priority or most recent N children. The database chooses winners before returning results; it may still scan and sort all candidates. For several independent, bounded subgraphs, keep using the `batchOnce(read => roots.map(root => read.subgraph(...)))` pattern above; record collection does not replace subgraph hydration. Use `store.batch()` when queued edge collection reads must participate. It runs them in sequence. On a transactional backend it still issues at least one statement per query plus `begin`/`commit`, so N queries are N+2 round trips at best; without transactions there is no framing. It buys a connection profile that never peaks at N — not lower latency, and not a snapshot (PostgreSQL's default read-committed isolation lets a later query see a newer commit). Edge collection `batchFind*` methods (`batchFindFrom`, `batchFindTo`, `batchFindByEndpoints`) also participate in `store.batch()`. On a transactional backend they move N `findFrom`/`findTo` calls into one transaction — the statement count is unchanged either way. If the round trips are what hurt, replace the calls with `store.neighbors()` / `store.countNeighbors()` or a traversal (one statement), or compose the batch-scoped `read.neighbors()`, `read.countNeighbors()`, and `read.subgraph()` forms in `batchOnce()`. The same read family is available on `TransactionContext`. Use `tx.neighbors()`, `tx.countNeighbors()`, or `tx.subgraph()` for a direct read that must see earlier writes in the callback. Use `tx.batchOnce()` to combine independent transaction-bound reads into exactly one statement on the held connection; fluent items in that batch start from `tx.query()`. Direct `tx.subgraph()` also uses its one-statement plan so it never submits concurrent statements to the held transaction connection. Direct `store.subgraph()` and batch-scoped `read.subgraph()` share one semantic planner and produce the same result, but intentionally use different physical plans. The direct form uses 2 statements on SQLite and 3 on PostgreSQL so each backend can hydrate a closure efficiently. The scoped form uses 1 statement everywhere to make cross-shape composition possible. On PostgreSQL, prefer the direct form for a standalone large closure; use the batch-scoped form when eliminating network round trips across several independent reads matters more than optimizing that closure in isolation. To read the edges of a *set* of endpoints, prefer `bulkFindFrom` / `bulkFindTo` (see [Edge Collections](/schemas-stores#edge-collections)). Where `store.batch()` runs N singleton reads over one connection, these widen the endpoint predicate itself to `from_id IN (...)` — one set-oriented statement per endpoint kind and bind-budget chunk, on the same index prefix seek the singleton read uses — and return the edges grouped per input: ```typescript const people = await store.nodes.Person.find({ limit: 50 }); const jobsPerPerson = await store.edges.worksAt.bulkFindFrom(people); // jobsPerPerson[i] holds the worksAt edges of people[i] ``` This is the fix for the "list view with relationship counts" N+1: statement count grows with endpoint kinds and bind-budget chunks instead of with every item on the page. Pass `limitPerInput` to bound each endpoint's fan-out. If a view spans several source kinds and edge kinds, use the Store-level `bulkFindEdgesFrom` operation instead of calling each licensed edge collection separately. It accepts heterogeneous source groups and edge kinds, then executes one set-oriented statement per bind-budget chunk. Round trips therefore grow with input size, not with the number of licensed `(source kind, edge kind)` combinations: ```typescript const edgesBySource = await store.bulkFindEdgesFrom({ sources: [ { kind: "Company", ids: companyIds }, { kind: "Person", ids: personIds }, ], edgeKinds: ["employs", "owns", "dependsOn"], }); // edgesBySource[i] identifies its source and contains that source's matching edges ``` :::note[Operation hooks] Bulk operations (`bulkCreate`, `bulkInsert`, `bulkUpsertById`, `bulkDelete`) skip per-item operation hooks for throughput, and the bulk hooks (`onBulkOperationStart` / `onBulkOperationEnd`) do not stand in for them — those fire only for node `updateWhere`, so a bulk method emits no hook events at all, neither per-item nor bulk. To observe every individual write, call the single-item method instead. Query hooks still fire normally. See [Schemas & Stores](/schemas-stores#observability-hooks) for details. ::: #### Choosing shared subgraphs Consider a `Person` graph with `name` and a long `biography` property, connected by outgoing `knows` edges. The following examples use the same Store and bounded depth. Sharing preserves a separate subgraph for each input root; it changes the database plan and response encoding, not the returned neighborhoods. **Good candidate: overlapping neighborhoods with substantial properties.** Ada and Bea both know Cara, and Cara knows Dev: ```text Ada ──knows──▶ Cara ──knows──▶ Dev Bea ──knows──▶ Cara ``` Both depth-two subgraphs contain Cara and Dev. If their biographies are several kilobytes each, independent plans repeat that property data. Enable sharing so the compatible reads hydrate those shared entities once: ```typescript const overlapping = await store.batchOnce( (read) => [ada.id, bea.id].map((rootId) => read.subgraph(rootId, { edges: ["knows"], maxDepth: 2, project: { nodes: { Person: ["name", "biography"] } }, }), ), { shareSubgraphs: true }, ); // overlapping[0]: Ada, Cara, Dev // overlapping[1]: Bea, Cara, Dev // Each result owns independent projected values, including nested objects. ``` **Keep the default: disjoint neighborhoods.** Suppose Erin knows Finn, Finn knows Gia, Hana knows Ivan, and Ivan knows Jules, with no connections between the groups: ```text Erin ──knows──▶ Finn ──knows──▶ Gia Hana ──knows──▶ Ivan ──knows──▶ Jules ``` No entities are reused across roots. Sharing adds membership information without removing duplicate biographies. Default `batchOnce()` still combines the two reads into one statement: ```typescript const disjoint = await store.batchOnce((read) => [erin.id, hana.id].map((rootId) => read.subgraph(rootId, { edges: ["knows"], maxDepth: 2, project: { nodes: { Person: ["name", "biography"] } }, }), ), ); ``` **Keep the default initially: overlapping roots, identity-only results.** A graph preview might need only connectivity, even when the stored biographies are large. Use empty property selections to retain identities and edge endpoints without transferring biographies: ```typescript const connectivity = await store.batchOnce((read) => [ada.id, bea.id].map((rootId) => read.subgraph(rootId, { edges: ["knows"], maxDepth: 2, project: { nodes: { Person: [] }, edges: { knows: [] } }, }), ), ); ``` Overlap alone does not make sharing a payload optimization here: little repeated property data remains to remove. Sharing may still improve latency, so measure both response size and elapsed time before choosing it. | Eight-root benchmark shape | Shared vs default batch encoded bytes | Starting choice | | --- | --- | --- | | 75% overlap, full 2,048-byte payload | About 27–29% fewer | Try sharing | | Disjoint roots, full 256-byte payload | About 21–25% more | Default batching | | 75% overlap, identities only | About 10–18% more | Default batching; measure latency | These response-size differences appeared in small local SQLite and PostgreSQL runs and a same-region remote Neon PostgreSQL run; they are not universal thresholds. Bytes measure JSON encoding at the backend boundary, not protocol traffic. In the remote run, shared batches were faster at the median in all three shapes, including the disjoint and identity-only shapes whose encoded responses grew. Default `batchOnce()` was faster at the median than concurrent direct subgraph calls in all three shapes. The remote client egress was observed in Bend, Oregon, and the pooled database endpoint was in Oregon; a simple pooled `SELECT 1` round trip measured 24 ms at the median. Tail latency varied, so compare elapsed time and response size on the deployment route that matters to your application. See the [SQLite report](https://github.com/nicia-ai/typegraph/blob/6196354c/packages/benchmarks/reports/subgraph-batch-sqlite-2026-09-14.md) and [PostgreSQL report](https://github.com/nicia-ai/typegraph/blob/6196354c/packages/benchmarks/reports/subgraph-batch-postgres-2026-09-14.md) for local timings and methodology, and the [remote Neon report](https://github.com/nicia-ai/typegraph/blob/cffd0082906bffd5f5e993dfeede1e01e6e6f300/packages/benchmarks/reports/subgraph-batch-neon-oregon-2026-09-15.md) for 40 raw samples per shape and mode, placement, and reproduction commands. The older PostgreSQL report also includes a separately labeled delay simulation; the Neon measurements used actual network calls without injected delay. Sharing also requires compatible options. Different edge sets, depths, temporal coordinates, projections, or edge windows can keep reads in separate groups even when their results overlap. For example, a social `knows` neighborhood and an employment `worksAt` neighborhood still fit in one `batchOnce()` call, but enabling sharing does not fuse those incompatible plans. ## Connection Management Managed local Store and backend factories own and close their SQLite or PGlite resources. Bring-your-own adapter integrations leave the supplied connection or pool under application control. See [Backend Setup](/backend-setup#connection-management) for the ownership matrix and shutdown examples. ### PostgreSQL pooling Always use a connection pool in production. An individual query holds a connection only while each statement runs. Most queries issue a single statement; a query whose selective-field mapping falls back issues a second. `store.transaction()` holds one connection for the whole callback, and `store.batch()` does the same for its implicit transaction. ```typescript import { Pool } from "pg"; const pool = new Pool({ connectionString: process.env.DATABASE_URL, max: 20, // Size based on your concurrency needs idleTimeoutMillis: 30_000, connectionTimeoutMillis: 2_000, }); pool.on("error", (err) => { console.error("Unexpected pool error", err); }); ``` **Sizing guidance:** Each concurrent query holds one connection for as long as its statement runs. A pool of 10–20 connections handles most workloads. If you're running bulk imports in parallel, size up accordingly. **Reducing pool pressure with `batch()`:** When loading multiple independent queries (e.g., a detail page with several relationship types), `Promise.all` can acquire up to N connections simultaneously — fewer if the pool is undersized or saturated, in which case it queues instead. [`store.batch()`](/schemas-stores#batch-query-execution) keeps at most one query in flight, so peak connection use is 1 — on a transactional backend that is literally one checked-out connection for the implicit transaction; elsewhere it is one at a time, and whether the adapter reuses the same client is its own business. It does not reduce the statement count, and read-committed isolation means it is not a snapshot. ### SQLite concurrency SQLite is single-writer. For best throughput: - Use WAL with `synchronous=NORMAL`. `createLocalSqliteBackend` applies both (plus a 5s `busy_timeout`) automatically; on a bring-your-own connection set them yourself: `sqlite.pragma("journal_mode = WAL")`, `sqlite.pragma("synchronous = NORMAL")`. On file databases this makes single-operation writes roughly 5× faster than the driver defaults. - Batch writes in transactions rather than issuing many small commits. (One nuance: pure bulk appends of fresh pages can run marginally faster under the rollback journal than WAL, since WAL writes pages twice — the per-commit wins dominate everywhere else.) - For read-heavy workloads, SQLite performs well without pooling since `better-sqlite3` is synchronous ### Transaction isolation PostgreSQL transactions accept an optional isolation level: ```typescript await store.transaction( async (tx) => { // Serializable isolation for strict consistency const snapshot = await tx.nodes.Account.getById(accountId); // ... }, { isolationLevel: "serializable" }, ); ``` Available levels: `read_uncommitted`, `read_committed` (default), `repeatable_read`, `serializable`. Schema-managed Stores fence writes against concurrent schema-version commits. That includes Stores opened by `createStoreWithSchema`, `createAdapterStoreWithSchema`, `createVerifiedStore`, or `createVerifiedAdapterStore`; an adapter Store constructed with a cached `{ reconciled }` snapshot; and Stores returned by `evolve()` or rebound from one of those Stores. `store.introspect().schemaVersion !== undefined` is the runtime test. PostgreSQL reacquires and validates the active-schema row lock at every managed write. The lock is normally reentrant and remains held to transaction end, but the repeated check is required because rolling back to a caller-created savepoint releases row locks acquired after that savepoint. At `repeatable_read` or `serializable`, a concurrent schema commit can raise PostgreSQL's normal serialization failure; retry the whole transaction. Graph-merge commits already retry those failures automatically. Raw `createStore` / `createAdapterStore` instances without a reconciled snapshot, and writes issued directly through a backend, do not carry schema metadata and remain outside this guarantee. `store.clear()` also resets the cleared Store to that raw state. SQLite always operates at `serializable` isolation. ## Query Optimization Features ### Precomputed Closures When you define an ontology (e.g., `subClassOf`, `implies`), TypeGraph precomputes the full transitive closure at store initialization. Queries like `.from("Parent", "p", { includeSubClasses: true })` use a pre-calculated list of kinds rather than recursive lookups at runtime. ### Smart Select For explicit SQL field selection, prefer [`project()`](/queries/expressions#projection-and-mapping). Use `map()` afterward for JavaScript transformations. This avoids legacy selector probing and makes the database projection explicit. TypeGraph automatically optimizes queries based on which fields your `select()` callback accesses. When you select specific fields, TypeGraph generates SQL that only extracts those fields using `json_extract()` (SQLite) or JSONB path extraction (PostgreSQL), rather than fetching the entire `props` blob. ```typescript // Optimized: Only fetches email and name from the database const results = await store .query() .from("Person", "p") .whereNode("p", (p) => p.email.eq("alice@example.com")) .select((ctx) => ({ email: ctx.p.email, name: ctx.p.name, })) .execute(); // SQL: SELECT json_extract(props, '$.email'), json_extract(props, '$.name') ... ``` This optimization pairs well with [covering indexes](/performance/indexes#covering-indexes): if your index contains both the filter keys and the selected keys, the database can serve the query straight from the index instead of scanning the whole table — though on PostgreSQL specifically, this stops short of a true `Index Only Scan` for JSONB-extracted fields; see the [covering indexes](/performance/indexes#covering-indexes) section for the concrete limitation and a workaround. **When optimization applies:** | Pattern | Optimized? | Reason | | ----------------------------------------------- | ---------- | ---------------------------------- | | `ctx => ({ email: ctx.p.email })` | Yes | Simple field extraction | | `ctx => [ctx.p.id, ctx.p.name]` | Yes | Multiple fields in array | | `ctx => ctx.p` | No | Whole node returned | | `ctx => ({ upper: ctx.p.email.toUpperCase() })` | Yes | Field extracted; method runs in JS | | `ctx => ({ ...ctx.p })` | No | Spread requires full node | The optimization is transparent — if your callback can't be optimized, TypeGraph automatically falls back to fetching the full node data. For data-dependent callbacks, TypeGraph first plans with representative values, including a high-value pass that covers common numeric threshold branches. If an unobserved branch accesses an additional field at execution time, the first miss may require a second statement that fetches the full row. Prepared queries remember that missing-field failure and use the full-row plan directly on later executions. Comparisons against arbitrary string values can still take an unobserved branch; the high-value pass does not guarantee that every possible callback path is planned in advance. :::note[Select callback purity] Smart select applies to `.execute()`, `.paginate()`, and `.stream()`. The `select()` callback may be evaluated multiple times during planning/optimization, so it should be pure (no side effects). ::: :::note[Known limitations] Smart select is not currently applied to queries that include variable-length traversals (recursive CTEs), even when the select callback is otherwise optimizable. ::: ### Built-in Indexes The default TypeGraph schema includes optimized indexes for the most common access patterns: - **Graph + Kind + ID**: Primary key for node lookups - **Graph + From/To ID**: Optimized for edge traversals - **Temporal columns**: Indexes on `valid_from`, `valid_to`, and `deleted_at` For application-specific indexes on JSON properties, see [Indexes](/performance/indexes). ### SQL Compilation Each builder method (`.where()`, `.limit()`, `.orderBy()`, etc.) returns a new immutable instance. A reused query instance compiles **once**. The first `.execute()` builds a cached template and every later call reuses it — for standard queries, aggregate queries, set-operation queries (`union`, `intersect`, `except`), and prepared queries alike. Explicit `.toSQL()` / `.compile()` calls are the exception: they compile on demand every time, because producing the statement is the thing the caller asked for. The subtlety a cache like that has to survive is freshness. A "current" (live) read filters on temporal validity as of the instant it runs, so a template with a concrete "now" baked into it would freeze that instant for the query instance's whole lifetime, hiding every row created afterward. The template therefore reserves the read instant as a **placeholder** rather than a value, and each execution fills it with a fresh instant alongside that call's bindings. Nothing in the statement's text depends on either, so reuse costs no freshness. ```typescript const activeUsers = store .query() .from("User", "u") .whereNode("u", (u) => u.status.eq("active")) .select((ctx) => ctx.u); // One compilation, two executions. The read instant is bound per call, so a // user created between these two is visible to the second one. await activeUsers.execute(); await activeUsers.execute(); ``` Two things fall back to compiling on every call: - **Backends that cannot execute pre-compiled SQL text** — a custom or async backend, i.e. one without `executeRaw`. - **Statements whose execution semantics ride on the compiled SQL object rather than its text**, even on PostgreSQL with `executeRaw` fully available. Two query shapes do: **approximate vector search** (`similarTo(..., { approximate: true })`, which carries the pgvector / `sqlite-vec` iterative-scan wrapper) and **`store.subgraph()` on PostgreSQL**, whose id-array fetches are marked to force a custom plan so the planner sizes them against the actual array rather than reusing a generic one. Flattening either to cacheable text would silently drop the behavior it depends on, so they are excluded deliberately — the trade is a template hit against correct execution, and correctness wins. Compilation is pure, in-memory string-building with no I/O, so both fallbacks are cheap; the query's database round-trip dominates either way. Worth knowing if you are profiling a vector query and expecting the compile-once behavior described above — that is the one shape where it does not apply. ### Prepared Queries For hot paths that execute the same query shape with different values, `.prepare()` builds and structurally validates the query AST once — a malformed query fails fast, before the first `.execute()`, instead of on first use — and compiles the statement once into a cached template. Each `.execute(bindings)` fills that template's placeholders (a fresh read instant plus the call's own parameter values) and runs the cached text directly through `executeRaw`. Because arity never reaches the SQL text, a list-valued parameter reuses the same template no matter how long the list is: ```typescript const byIds = store .query() .from("Person", "p") .whereNode("p", (p) => p.id.in(param("ids"))) .select((ctx) => ctx.p) .prepare(); await byIds.execute({ ids: ["a", "b", "c"] }); await byIds.execute({ ids: ["d"] }); // same compiled statement ``` Best for: validating a query shape once, then reusing it with different parameter values. The saved compilation is real but small — the database round-trip still dominates. See [Prepared Queries](/queries/execute#prepared-queries) for usage details. ### Subgraph extraction For the "load entity with all relationships" pattern, [`store.subgraph()`](/schemas-stores#subgraph-extraction) is a backend-tuned option for a single bounded neighborhood. It compiles to a recursive CTE that fans out across all specified edge types in a fixed 2 statements on SQLite and 3 on PostgreSQL — no matter how many relationship kinds are involved, or how much it returns. See [Choosing a query strategy](/schemas-stores#choosing-a-query-strategy) for guidance on when to use `subgraph()` vs the fluent query builder vs manual `findFrom` calls. The [`project` option](/schemas-stores#subgraph-projection) further reduces overhead by extracting only the specified fields per kind at the SQL level via `json_extract()` / JSONB paths, skipping full `props` blob transfer and metadata columns for projected kinds. ## Best Practices ### Filter early Use `.whereNode()` and `.whereEdge()` for match constraints that should restrict expansion. The compiler applies them at the matching stage regardless of their position in the chain. Use scoped `.where()` when the condition must filter completed rows; moving a completed-row condition into an optional match or recursive hop can change its meaning. ### Select specific fields When you only need certain fields, use `project()` to make the SQL projection explicit. Legacy `select()` also supports [smart select optimization](#smart-select). Smaller projections reduce transferred data and may benefit from covering indexes, subject to the engine limitations described above. ```typescript // Preferred: Only fetches what you need .project((e) => ({ name: e.p.name, email: e.p.email })) // Avoid when possible: Fetches entire props blob .select((ctx) => ctx.p) ``` ### Use specific kinds Unless you specifically need to query across a hierarchy, avoid `includeSubClasses: true`. Being specific about the node kind allows the SQL engine to use more restrictive index scans. ### Use cursor pagination For large datasets, prefer `.paginate()` over `.limit()` and `.offset()`. Keyset pagination (using cursors) avoids the `O(N)` cost of skipping rows in standard SQL offsets. ### Index your filter and sort properties TypeGraph's built-in indexes cover structural lookups (by ID, by edge endpoints). Properties you filter or sort on in `whereNode()`, `whereEdge()`, and `orderBy()` need application-specific [expression indexes](/performance/indexes). Use the [Query Profiler](/performance/profiler) to identify which properties need coverage. ## Profile Your Queries Use the [Query Profiler](/performance/profiler) to identify missing indexes and understand query patterns in your application. The profiler captures property access patterns and generates prioritized index recommendations. ```typescript import { QueryProfiler } from "@nicia-ai/typegraph/profiler"; const profiler = new QueryProfiler(); const profiledStore = profiler.attachToStore(store); // Run your application or test suite... const report = profiler.getReport(); console.log(report.recommendations); ``` ## Benchmarks TypeGraph uses a deterministic performance sanity suite as its benchmark and regression gate. The suite seeds a realistic graph shape and measures end-to-end query latency across: - forward and reverse traversals - inverse/symmetric traversal (`expand: "inverse"` / `expand: "all"`) - 2-hop and 3-hop traversals - aggregate queries - cached execute vs prepared execute - deep traversals (`10`/`100`/`1000` hop recursive with `cyclePolicy: "allow"`) Guardrail thresholds enforce expected behavior in CI (for example, traversal latency caps and ratio checks such as reverse/forward and deep-hop scaling). Deep-recursive benchmark probes explicitly set `cyclePolicy: "allow"` to isolate recursive CTE expansion cost; the default `cyclePolicy: "prevent"` prioritizes cycle-safe semantics and is expected to be slower on long traversals. *Note: Real-world performance varies by hardware, database driver, network latency (for PostgreSQL), and schema/data shape.*
Benchmark configuration and guardrails Current suite configuration: | Setting | Value | | ----------------------------------- | ----- | | Seed users | 1200 | | Follows per user | 10 | | Posts per user | 5 | | Batch size | 250 | | Warmup iterations | 2 | | Sample iterations (median reported) | 15 | Default guardrails: | Check | Threshold | | ------------------------------------------ | --------- | | reverse/forward ratio | <= 6x | | inverse traversal latency | <= 500ms | | inverse/forward ratio | <= 10x | | 3-hop latency | <= 500ms | | 3-hop/2-hop ratio | <= 8x | | aggregate latency | <= 500ms | | aggregate distinct latency | <= 700ms | | aggregateDistinct/aggregate ratio | <= 4x | | cached execute latency | <= 500ms | | prepared execute latency | <= 500ms | | prepared/cached ratio | <= 2x | | 10-hop recursive latency | <= 250ms | | 100-hop recursive latency | <= 1000ms | | 100-hop-recursive/10-hop-recursive ratio | <= 30x | | 1000-hop recursive latency | <= 5000ms | | 1000-hop-recursive/100-hop-recursive ratio | <= 20x | Backend-specific overrides: | Backend | Check | Threshold | | ---------- | -------------------------- | --------- | | SQLite | 1000-hop recursive latency | <= 7000ms | | PostgreSQL | inverse traversal latency | <= 1000ms | | PostgreSQL | inverse/forward ratio | <= 30x | | PostgreSQL | 3-hop latency | <= 1000ms | | PostgreSQL | aggregate distinct latency | <= 1200ms | | PostgreSQL | prepared execute latency | <= 700ms |
### Real-world workload validation Beyond the synthetic guardrail suite above, TypeGraph is also exercised against the [LDBC Social Network Benchmark (SNB) Interactive](https://github.com/ldbc/ldbc_snb_interactive_v1) workload — a standard, independently-defined graph benchmark, not a TypeGraph-specific one — at SF1 scale (~10k persons, ~1M posts, ~2M comments). This surfaced and fixed two real scaling bugs in the library: an unbounded `ANALYZE` cost on bulk SQLite loads, and an N+1 endpoint-existence check in batched edge creation. It also directly produced the `keySystemColumns` guidance and the PostgreSQL index-only-scan caveat in [Indexes](/performance/indexes#covering-indexes). The benchmark source lives in `packages/benchmarks/src/real/` in the repository. ### Running benchmarks locally ```bash pnpm bench ``` For guardrail mode (fails on regression thresholds): ```bash pnpm --filter @nicia-ai/typegraph-benchmarks perf:check ``` Run the same guardrailed suite against PostgreSQL: ```bash POSTGRES_URL=postgresql://typegraph:typegraph@127.0.0.1:5432/typegraph_test \ pnpm --filter @nicia-ai/typegraph-benchmarks perf:check:postgres ``` By default the SQLite suite runs against an in-memory database, which measures engine and compile cost but not WAL/fsync behavior. Add `--storage=file` (or use the `perf:file` / `perf:check:file` scripts) to run against a temporary on-disk database — the lane that reflects real local deployments. A separate write-throughput bench measures single-op creates, transaction-amortized creates, `bulkCreate`, search-indexed creates (fulltext + vector sync), and `importGraph`, normalized to milliseconds per operation: ```bash pnpm --filter @nicia-ai/typegraph-benchmarks bench:write # sqlite, in-memory pnpm --filter @nicia-ai/typegraph-benchmarks bench:write:file # sqlite, on-disk POSTGRES_URL=... pnpm --filter @nicia-ai/typegraph-benchmarks bench:write:postgres ``` The write bench is report-only (no guardrails): write latency is dominated by fsync behavior on the file lane and needs per-machine calibration. The benchmark source code is located in `packages/benchmarks/src/`. ## Next Steps - [Indexes](/performance/indexes) — Define custom indexes for your schema - [Query Profiler](/performance/profiler) — Identify missing indexes automatically - [Backend Setup](/backend-setup) — Connection setup, pooling, and lifecycle # Query Profiler > Capture query patterns and generate index recommendations The Query Profiler captures property access patterns from your queries and generates index recommendations. Use it during development or in test suites to identify missing indexes. ## Quick Start ```typescript import { QueryProfiler } from "@nicia-ai/typegraph/profiler"; // Create a profiler and attach it to your store const profiler = new QueryProfiler(); const profiledStore = profiler.attachToStore(store); // Run queries as normal - they're automatically tracked await profiledStore .query() .from("Person", "p") .whereNode("p", (p) => p.email.eq("alice@example.com")) .select((ctx) => ({ name: ctx.p.name })) .execute(); // Get recommendations const report = profiler.getReport(); for (const rec of report.recommendations) { console.log( `[${rec.priority}] ${rec.entityType}:${rec.kind} ${rec.fields.join(", ")}`, ); console.log(` ${rec.reason}`); } ``` ## How It Works The profiler uses JavaScript Proxy to transparently wrap your store and query builders. When queries execute, it extracts property access patterns from the query AST: - **Filter patterns**: Properties used in `.whereNode()` and `.whereEdge()` predicates - **Sort patterns**: Properties used in `.orderBy()` - **Select patterns**: Properties accessed in `.select()` callbacks - **Group patterns**: Properties used in `.groupBy()` The profiler then compares these patterns against your declared indexes and generates recommendations for missing coverage. ## Kinds and `includeSubClasses` When you query with `includeSubClasses: true`, a single alias can represent multiple kinds. When the profiler is attached to a store, it uses the graph schema to attribute a property access only to kinds where that JSON path exists. This avoids recommending indexes for unrelated subclasses. ## Attaching to a Store ```typescript const profiler = new QueryProfiler(); const profiledStore = profiler.attachToStore(store); // The profiled store behaves exactly like the original await profiledStore.nodes.Person.create({ email: "bob@example.com", name: "Bob" }); // Queries are tracked automatically await profiledStore.query().from("Person", "p").select((ctx) => ctx.p).execute(); // Access the profiler from the store profiledStore.profiler.getReport(); ``` The profiled store exposes a `profiler` property for convenient access. ## Declaring Existing Indexes Pass your existing indexes so the profiler doesn't recommend indexes you already have: ```typescript import { QueryProfiler } from "@nicia-ai/typegraph/profiler"; import { toDeclaredIndexes } from "@nicia-ai/typegraph/indexes"; import { personEmail, worksAtRole } from "./indexes"; const profiler = new QueryProfiler({ declaredIndexes: toDeclaredIndexes([personEmail, worksAtRole]), }); ``` You can also declare indexes manually: ```typescript const profiler = new QueryProfiler({ declaredIndexes: [ { entityType: "node", kind: "Person", fields: ["/email"], unique: true, name: "idx_person_email", }, { entityType: "node", kind: "Person", fields: ["/name"], unique: false, name: "idx_person_name", }, ], }); ``` ## Understanding the Report ```typescript const report = profiler.getReport(); ``` The report contains: ### `recommendations` Prioritized index recommendations sorted by importance: ```typescript for (const rec of report.recommendations) { console.log( `[${rec.priority}] ${rec.entityType}:${rec.kind} ${rec.fields.join(", ")}`, ); console.log(` Reason: ${rec.reason}`); console.log(` Frequency: ${rec.frequency}`); } ``` **Priority levels:** - `high`: Property accessed 10+ times in filters/sorts (configurable) - `medium`: Property accessed 5-9 times (configurable) - `low`: Property accessed 3-4 times (configurable) ### `unindexedFilters` Properties used in filter predicates that lack index coverage: ```typescript for (const path of report.unindexedFilters) { const target = path.target.__type === "prop" ? path.target.pointer : path.target.field; console.log(`Unindexed filter: ${path.entityType}:${path.kind} ${target}`); } ``` ### `patterns` Raw property access statistics: ```typescript for (const [key, stats] of report.patterns) { console.log(`${key}: ${stats.count} accesses`); console.log(` Contexts: ${[...stats.contexts].join(", ")}`); console.log(` Predicates: ${[...stats.predicateTypes].join(", ")}`); } ``` ### `summary` Session statistics: ```typescript console.log(`Total queries: ${report.summary.totalQueries}`); console.log(`Unique patterns: ${report.summary.uniquePatterns}`); console.log(`Duration: ${report.summary.durationMs}ms`); ``` ## Test Assertions Use `assertIndexCoverage()` to fail tests when queries filter on unindexed properties: ```typescript import { describe, it, beforeAll, afterAll } from "vitest"; import { QueryProfiler } from "@nicia-ai/typegraph/profiler"; describe("Query Performance", () => { let profiler: QueryProfiler; let profiledStore: ProfiledStore; beforeAll(() => { profiler = new QueryProfiler({ declaredIndexes: toDeclaredIndexes([personEmail, personName]), }); profiledStore = profiler.attachToStore(store); }); // Run your test suite against profiledStore... it("all filtered properties should be indexed", () => { // Throws if any filter property lacks an index profiler.assertIndexCoverage(); }); }); ``` ## Configuration ```typescript const profiler = new QueryProfiler({ // Indexes you already have declaredIndexes: [...], // Minimum frequency to generate a recommendation (default: 3) minFrequencyForRecommendation: 5, // Optional priority thresholds (defaults: 5 and 10) mediumFrequencyThreshold: 8, highFrequencyThreshold: 20, }); ``` ## Lifecycle Methods ```typescript // Reset collected data (keeps configuration) profiler.reset(); // Detach from store (allows reattachment) profiler.detach(); // Check attachment status if (profiler.isAttached) { console.log("Profiler is attached to a store"); } ``` ## Manual Recording For custom integrations, record queries directly from their AST: ```typescript const query = store .query() .from("Person", "p") .whereNode("p", (p) => p.email.eq("test@example.com")) .select((ctx) => ctx.p); // Record without executing profiler.recordQuery(query.toAst()); ``` ## Composite Index Detection The profiler understands composite index prefix matching. If you have an index on `["email", "name"]`, queries filtering on just `email` are considered covered: ```typescript const profiler = new QueryProfiler({ declaredIndexes: [ { entityType: "node", kind: "Person", fields: ["/email", "/name"], unique: false, name: "idx_email_name", }, ], }); // This query IS covered (uses the email prefix of the composite index) await profiledStore .query() .from("Person", "p") .whereNode("p", (p) => p.email.eq("test@example.com")) .execute(); // No recommendation generated for email ``` ## Best Practices 1. **Profile realistic workloads**: Run your actual queries or test suite, not synthetic benchmarks. 2. **Profile before optimizing**: Don't guess which indexes you need - let the profiler tell you. 3. **Use in CI**: Add `assertIndexCoverage()` to your test suite to catch regressions. 4. **Declare all indexes**: Pass your existing indexes so recommendations are accurate. 5. **Review frequency**: High-frequency patterns are most important to index. ## Next Steps - [Indexes](/performance/indexes) - Create the indexes the profiler recommends - [Performance Overview](/performance/overview) - Best practices and smart select # Project Structure > Recommended patterns for organizing TypeGraph in your codebase How you organize your TypeGraph code depends on your project's size and complexity. This guide covers recommended patterns from simple single-file setups to large multi-domain graphs. ## Small Projects For projects with a handful of node and edge types, keep everything in two files: ```text src/ graph.ts # Node/edge definitions + graph graph-store.ts # Store instantiation ``` ### graph.ts Contains all definitions and exports the graph: ```typescript import { z } from "zod"; import { defineNode, defineEdge, defineGraph, disjointWith } from "@nicia-ai/typegraph"; // Node definitions export const Person = defineNode("Person", { schema: z.object({ name: z.string(), email: z.string().email().optional(), }), }); export const Company = defineNode("Company", { schema: z.object({ name: z.string(), industry: z.string().optional(), }), }); // Edge definitions export const worksAt = defineEdge("worksAt", { schema: z.object({ role: z.string().optional(), since: z.string().optional(), }), }); // Graph definition export const graph = defineGraph({ id: "my_app", nodes: { Person: { type: Person }, Company: { type: Company }, }, edges: { worksAt: { type: worksAt, from: [Person], to: [Company] }, }, ontology: [disjointWith(Person, Company)], }); ``` ### graph-store.ts Instantiates and exports the store: ```typescript import { createStore } from "@nicia-ai/typegraph"; import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; import { graph } from "./graph"; const { backend } = createLocalSqliteBackend({ path: "./data.db" }); export const store = createStore(graph, backend); ``` This separation keeps the schema definition (which is static) separate from store instantiation (which involves runtime configuration like database paths). ## Medium Projects When your graph grows to 10+ node types or you want better organization, split definitions into separate files: ```text src/graph/ index.ts # Re-exports + defineGraph nodes.ts # All node definitions edges.ts # All edge definitions ontology.ts # Ontological relations store.ts # Store instantiation ``` ### nodes.ts ```typescript import { z } from "zod"; import { defineNode } from "@nicia-ai/typegraph"; export const Person = defineNode("Person", { schema: z.object({ name: z.string(), email: z.string().email().optional(), role: z.string().optional(), }), }); export const Company = defineNode("Company", { schema: z.object({ name: z.string(), industry: z.string().optional(), founded: z.number().optional(), }), }); export const Project = defineNode("Project", { schema: z.object({ name: z.string(), status: z.enum(["planning", "active", "completed"]), }), }); // ... more node definitions ``` ### edges.ts ```typescript import { z } from "zod"; import { defineEdge } from "@nicia-ai/typegraph"; export const worksAt = defineEdge("worksAt", { schema: z.object({ role: z.string().optional(), since: z.string().optional(), }), }); export const manages = defineEdge("manages"); export const assignedTo = defineEdge("assignedTo", { schema: z.object({ assignedAt: z.string().optional(), }), }); // ... more edge definitions ``` ### ontology.ts ```typescript import { subClassOf, disjointWith, inverseOf } from "@nicia-ai/typegraph"; import { Person, Company, Project } from "./nodes"; import { manages } from "./edges"; export const ontology = [ disjointWith(Person, Company), disjointWith(Person, Project), disjointWith(Company, Project), // inverseOf(manages, reportsTo), ]; ``` ### index.ts Combines everything into the graph definition: ```typescript import { defineGraph } from "@nicia-ai/typegraph"; import { Person, Company, Project } from "./nodes"; import { worksAt, manages, assignedTo } from "./edges"; import { ontology } from "./ontology"; export const graph = defineGraph({ id: "my_app", nodes: { Person: { type: Person }, Company: { type: Company }, Project: { type: Project }, }, edges: { worksAt: { type: worksAt, from: [Person], to: [Company] }, manages: { type: manages, from: [Person], to: [Person] }, assignedTo: { type: assignedTo, from: [Project], to: [Person] }, }, ontology, }); // Re-export for convenience export * from "./nodes"; export * from "./edges"; export { store } from "./store"; ``` ### store.ts ```typescript import { createStore } from "@nicia-ai/typegraph"; import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; import { graph } from "./index"; const { backend } = createLocalSqliteBackend({ path: "./data.db" }); export const store = createStore(graph, backend); ``` ## Large Projects For large graphs with distinct domains, group related nodes and edges together: ```text src/graph/ index.ts # Combines all domains store.ts # Store instantiation domains/ users.ts # User, Profile, Team + related edges content.ts # Document, Comment, Tag + related edges projects.ts # Project, Task, Milestone + related edges ``` ### domains/users.ts ```typescript import { z } from "zod"; import { defineNode, defineEdge, subClassOf, disjointWith } from "@nicia-ai/typegraph"; // Nodes export const User = defineNode("User", { schema: z.object({ email: z.string().email(), name: z.string(), role: z.enum(["admin", "member", "guest"]), }), }); export const Profile = defineNode("Profile", { schema: z.object({ bio: z.string().optional(), avatarUrl: z.string().optional(), }), }); export const Team = defineNode("Team", { schema: z.object({ name: z.string(), description: z.string().optional(), }), }); // Edges export const hasProfile = defineEdge("hasProfile"); export const memberOf = defineEdge("memberOf", { schema: z.object({ joinedAt: z.string().optional() }), }); export const leads = defineEdge("leads"); // Domain-specific ontology export const usersOntology = [ disjointWith(User, Team), disjointWith(User, Profile), ]; // Export for graph assembly export const usersNodes = { User: { type: User }, Profile: { type: Profile }, Team: { type: Team }, }; export const usersEdges = { hasProfile: { type: hasProfile, from: [User], to: [Profile] }, memberOf: { type: memberOf, from: [User], to: [Team] }, leads: { type: leads, from: [User], to: [Team] }, }; ``` ### index.ts Assembles domains into the final graph: ```typescript import { defineGraph } from "@nicia-ai/typegraph"; import { usersNodes, usersEdges, usersOntology } from "./domains/users"; import { contentNodes, contentEdges, contentOntology } from "./domains/content"; import { projectsNodes, projectsEdges, projectsOntology } from "./domains/projects"; export const graph = defineGraph({ id: "my_app", nodes: { ...usersNodes, ...contentNodes, ...projectsNodes, }, edges: { ...usersEdges, ...contentEdges, ...projectsEdges, }, ontology: [ ...usersOntology, ...contentOntology, ...projectsOntology, ], }); // Re-export types for convenience export * from "./domains/users"; export * from "./domains/content"; export * from "./domains/projects"; export { store } from "./store"; ``` ## Cross-Domain Edges When edges connect nodes from different domains, define them at the graph level: ```typescript // index.ts import { defineEdge } from "@nicia-ai/typegraph"; import { User } from "./domains/users"; import { Document } from "./domains/content"; import { Project } from "./domains/projects"; // Cross-domain edges const authored = defineEdge("authored"); const assignedTo = defineEdge("assignedTo"); export const graph = defineGraph({ // ... edges: { ...usersEdges, ...contentEdges, ...projectsEdges, // Cross-domain authored: { type: authored, from: [User], to: [Document] }, assignedTo: { type: assignedTo, from: [User], to: [Project] }, }, }); ``` ## Naming Conventions | Element | Convention | Example | |---------|------------|---------| | Node definitions | PascalCase | `Person`, `Company` | | Edge definitions | camelCase | `worksAt`, `hasAuthor` | | Graph IDs | snake_case | `my_app`, `content_graph` | | Files | kebab-case | `graph-store.ts`, `project-structure.ts` | | Query aliases | short lowercase | `p`, `c`, `e1` | ## Type Exports Export types alongside definitions for use in your application: ```typescript // graph/nodes.ts import { type Node, type NodeProps, type NodeId } from "@nicia-ai/typegraph"; export const Person = defineNode("Person", { /* ... */ }); // Convenience type exports export type PersonNode = Node; export type PersonProps = NodeProps; export type PersonId = NodeId; ``` This lets consumers import types directly: ```typescript import { type PersonNode, type PersonProps } from "./graph"; function displayPerson(person: PersonNode) { console.log(person.name); } function validatePersonInput(data: unknown): PersonProps { return Person.schema.parse(data); } ``` ## Framework Integration ### Next.js / React Server Components Keep the store in a server-only module: ```text src/ graph/ index.ts store.server.ts # Server-only store ``` ```typescript // store.server.ts import "server-only"; import { createStore } from "@nicia-ai/typegraph"; import { graph } from "./index"; // ... ``` ### Edge Runtimes (Cloudflare Workers, Vercel Edge) Use the Drizzle backend with edge-compatible drivers: ```typescript // graph/store.ts import { createStore } from "@nicia-ai/typegraph"; import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite"; import { drizzle } from "drizzle-orm/d1"; import { graph } from "./index"; export function createGraphStore(env: { DB: D1Database }) { const db = drizzle(env.DB); const backend = createSqliteBackend(db); return createStore(graph, backend); } ``` ## Next Steps - [Getting Started](/getting-started) - Build your first graph - [Schemas & Types](/core-concepts) - Deep dive into node and edge definitions - [Integration](/integration) - Database setup and Drizzle integration # Provenance and Retraction > Track source lineage for derived facts, retract bad sources, and use recorded time to replay what the graph believed before and after the transition. Provenance and Retraction is the TypeGraph subpath for source lineage and belief transitions. It maps your ordinary graph kinds onto four roles: - one or more retractable source node kinds with a boolean `retracted` flag - a justification node that represents an AND support rule - one or more derived fact node kinds - two typed edges: premises point to justifications, and justifications derive facts The API lives at `@nicia-ai/typegraph/provenance`: ```typescript import { createRetractionCapability } from "@nicia-ai/typegraph/provenance"; const provenance = createRetractionCapability(store, { source: { kind: "Source" }, justification: { kind: "Justification" }, fact: { kinds: ["Fact"] }, premiseOf: { kind: "premiseOf" }, derives: { kind: "derives" }, }); ``` Use `source: { kinds: [...] }` when different source node kinds share the same boolean retraction field: ```typescript const provenance = createRetractionCapability(store, { source: { kinds: ["ScannerSource", "VendorSource"] }, justification: { kind: "Justification" }, fact: { kinds: ["Vulnerability", "DeployDecision"] }, premiseOf: { kind: "premiseOf" }, derives: { kind: "derives" }, }); ``` `store` must be created with `{ history: true }`. Retraction mutates graph row currency, so TypeGraph-managed recorded capture is required: ```typescript const [store] = await createStoreWithSchema(graph, backend, { history: true, }); ``` For a complete runnable version, see [Provenance Retraction](/examples/provenance-retraction). ## Graph shape Define the roles as normal TypeGraph nodes and edges. ```typescript const Source = defineNode("Source", { schema: z.object({ label: z.string(), retracted: z.boolean().default(false), }), }); const Fact = defineNode("Fact", { schema: z.object({ label: z.string() }), }); const TerminalFact = defineNode("TerminalFact", { schema: z.object({ label: z.string() }), }); const Justification = defineNode("Justification", { schema: z.object({ label: z.string() }), }); const premiseOf = defineEdge("premiseOf"); const derives = defineEdge("derives"); const graph = defineGraph({ id: "claims", nodes: { Source: { type: Source }, Fact: { type: Fact }, TerminalFact: { type: TerminalFact }, Justification: { type: Justification }, }, edges: { premiseOf: { type: premiseOf, from: [Source, Fact], to: [Justification] }, derives: { type: derives, from: [Justification], to: [Fact, TerminalFact] }, }, }); ``` A justification fires when all of its premise nodes are in the well-founded support set. Sources are in support unless their `retracted` flag is true. Facts enter support when at least one firing justification derives them. Fact kinds only need to appear in `premiseOf.from` if they can support another justification. Terminal facts can be listed in `fact.kinds` and `derives.to` without being valid premise endpoints. ## Retraction `retract(source)` sets the source flag, recomputes support from the current provenance graph, and makes unsupported facts non-current. A transition only touches facts reachable from the flipped sources, and closing a fact is a belief-status change, not a domain delete: none of the fact's edges are deleted (its `onDelete` behavior is not enforced), so `unRetract` restores the fact exactly as it was. ```typescript const before = await store.recordedNow(); const report = await provenance.retract({ kind: "Source", id: sourceId }); const after = await store.recordedNow(); const previous = before ? store.asOfRecorded(before) : undefined; const current = after ? store.asOfRecorded(after) : undefined; ``` The report partitions facts relative to the retracted source: - `died`: facts that were believed before and lost grounded support - `survivedVia`: affected facts that still have a firing justification - `unaffected`: previously believed facts outside the source's provenance `unRetract(source)` clears the source flag, recomputes support, and reopens facts that regain support. Use `retractMany(sources)` or `unRetractMany(sources)` to change several source flags in one recorded transaction: ```typescript const report = await provenance.retractMany([ { kind: "ScannerSource", id: scannerId }, { kind: "VendorSource", id: vendorId }, ]); ``` ## Recorded time Retraction uses TypeGraph-managed writes, so before and after states are visible through recorded-time reads. On PostgreSQL, provenance transitions serialize with TypeGraph-managed history writes on the same graph before computing and applying fact currency. Capture is scoped to TypeGraph-managed writes; it does not claim to observe out-of-band database mutations. ```typescript const factBefore = before ? await store.asOfRecorded(before).nodes.Fact.getById(factId) : undefined; const factAfter = after ? await store.asOfRecorded(after).nodes.Fact.getById(factId) : undefined; ``` Use `holding()` when you only need the current well-founded believed facts: ```typescript const facts = await provenance.holding(); ``` # Subqueries > EXISTS, IN, and correlated subqueries for complex filtering Subqueries let you filter based on conditions that depend on related data—check if related records exist, or if values appear in another query's results. ## EXISTS Check if related records exist: ```typescript import { exists, fieldRef } from "@nicia-ai/typegraph"; // Find people who have authored at least one PR const authors = await store .query() .from("Person", "p") .whereNode("p", () => exists( store .query() .from("PullRequest", "pr") .traverse("author", "e", { direction: "in" }) .to("Person", "author") .whereNode("author", (a) => a.id.eq(fieldRef("p", ["id"]))) .select((ctx) => ({ id: ctx.pr.id })) .toAst() ) ) .select((ctx) => ctx.p) .execute(); ``` ## NOT EXISTS Find records without related records: ```typescript import { notExists, fieldRef } from "@nicia-ai/typegraph"; // Find people with no pull requests const nonContributors = await store .query() .from("Person", "p") .whereNode("p", () => notExists( store .query() .from("PullRequest", "pr") .traverse("author", "e", { direction: "in" }) .to("Person", "author") .whereNode("author", (a) => a.id.eq(fieldRef("p", ["id"]))) .select((ctx) => ({ id: ctx.pr.id })) .toAst() ) ) .select((ctx) => ctx.p) .execute(); ``` ## IN Check if a value is in a subquery result set: ```typescript import { inSubquery, fieldRef } from "@nicia-ai/typegraph"; // Find people who work at tech companies const techWorkers = await store .query() .from("Person", "p") .whereNode("p", () => inSubquery( fieldRef("p", ["companyId"]), store .query() .from("Company", "c") .whereNode("c", (c) => c.industry.eq("Technology")) .aggregate({ id: fieldRef("c", ["id"], { valueType: "string" }), }) .toAst() ) ) .select((ctx) => ctx.p) .execute(); ``` ## NOT IN Exclude values that appear in a subquery: ```typescript import { notInSubquery, fieldRef } from "@nicia-ai/typegraph"; // Find people not in the blocklist const allowedUsers = await store .query() .from("Person", "p") .whereNode("p", () => notInSubquery( fieldRef("p", ["id"]), store .query() .from("BlockedUser", "b") .aggregate({ userId: fieldRef("b", ["props", "userId"], { valueType: "string" }), }) .toAst() ) ) .select((ctx) => ctx.p) .execute(); ``` ## fieldRef() The `fieldRef()` function creates a reference to a field in the outer query for use in subquery predicates: ```typescript import { fieldRef } from "@nicia-ai/typegraph"; fieldRef("alias", ["field"]) // Reference a single field fieldRef("alias", ["nested", "path"]) // Reference a nested field ``` **Parameters:** | Parameter | Type | Description | |-----------|------|-------------| | `alias` | `string` | The alias of the node/edge in the outer query | | `path` | `string[]` | Path to the field (array for nested access) | ## Helpers Reference | Function | Description | |----------|-------------| | `exists(subqueryAst)` | True if subquery returns any rows | | `notExists(subqueryAst)` | True if subquery returns no rows | | `inSubquery(fieldRef, subqueryAst)` | True if field value is in subquery results | | `notInSubquery(fieldRef, subqueryAst)` | True if field value is not in subquery results | For `inSubquery()` and `notInSubquery()`, the subquery must project exactly one scalar column. Prefer `aggregate({ ... })` with a single field. ## Real-World Examples ### Users with Recent Activity ```typescript // Find users who logged in within the last 7 days const activeUsers = await store .query() .from("User", "u") .whereNode("u", () => exists( store .query() .from("LoginEvent", "e") .whereNode("e", (e) => e.userId.eq(fieldRef("u", ["id"])) .and(e.timestamp.gte(sevenDaysAgo)) ) .select((ctx) => ({ id: ctx.e.id })) .toAst() ) ) .select((ctx) => ctx.u) .execute(); ``` ### Products Not in Any Cart ```typescript // Find products that haven't been added to any cart const unpopularProducts = await store .query() .from("Product", "p") .whereNode("p", () => notExists( store .query() .from("CartItem", "ci") .whereNode("ci", (ci) => ci.productId.eq(fieldRef("p", ["id"]))) .select((ctx) => ({ id: ctx.ci.id })) .toAst() ) ) .select((ctx) => ctx.p) .execute(); ``` ### Users in Specific Teams ```typescript // Find users who are members of either the engineering or design team const targetTeamIds = ["team-eng", "team-design"]; const teamMembers = await store .query() .from("User", "u") .whereNode("u", () => inSubquery( fieldRef("u", ["id"]), store .query() .from("TeamMembership", "tm") .whereNode("tm", (tm) => tm.teamId.in(targetTeamIds)) .aggregate({ userId: fieldRef("tm", ["props", "userId"], { valueType: "string", }), }) .toAst() ) ) .select((ctx) => ctx.u) .execute(); ``` ## Query Debugging For debugging or advanced use cases, you can inspect the query AST or generated SQL. ### View the AST ```typescript const query = store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ctx.p); const ast = query.toAst(); console.log(JSON.stringify(ast, null, 2)); ``` ### View Generated SQL `toSQL()` returns the SQL text and bound parameters for the current backend dialect: ```typescript const { sql, params } = query.toSQL(); console.log("SQL:", sql); console.log("Parameters:", params); ``` This is useful for: - Debugging query behavior - Understanding performance characteristics - Logging queries in production - Running the query with a custom executor ## Next Steps - [Filter](/queries/filter) - Basic filtering with predicates - [Combine](/queries/combine) - Set operations - [Execute](/queries/execute) - Running queries # Aggregate > GROUP BY, aggregate functions, and HAVING clauses TypeGraph supports SQL-style aggregations for analytics and reporting. Group nodes by properties, compute aggregates like COUNT and SUM, and filter groups with HAVING clauses. Call [`asRelation()`](/queries/relations/) on an aggregate query to filter its completed output or aggregate those results again. Both expression aggregates and compatibility aggregates support prepared execution and one-statement batching through the shared relation API. Typed expression callbacks are recommended for new queries. They provide schema-checked operands, computed aggregate arguments, and inferred nullable result types: ```typescript import { expr } from "@nicia-ai/typegraph"; const companySizes = await store .query() .from("Person", "p") .groupBy((e) => [e.p.department]) .having((e) => expr.gt(expr.count(e.p.id), expr.literal(5))) .aggregate((e) => ({ department: e.p.department, employees: expr.count(e.p.id), payroll: expr.sum(e.p.salary), })) .execute(); ``` The string helpers below remain supported as compatibility adapters. See [Database Expressions](/queries/expressions) for arithmetic, conditions, projection, and scope safety. ## When to Use Aggregations Aggregations are useful for: - **Analytics dashboards**: Employee counts by department, revenue by region - **Reporting**: Average order value, total sales by product category - **Data exploration**: Find groups meeting certain criteria - **Metrics**: Count active users, sum transaction amounts ## Basic Aggregation Use `groupBy()` and `aggregate()` with aggregate helper functions: ```typescript import { count, field } from "@nicia-ai/typegraph"; const companySizes = await store .query() .from("Person", "p") .traverse("worksAt", "e") .to("Company", "c") .groupBy("c", "name") // Group by company name .aggregate({ companyName: field("c", "name"), // Include the grouped field employeeCount: count("p"), // Count people in each group }) .execute(); // Result: [{ companyName: "Acme Corp", employeeCount: 42 }, ...] ``` ## Aggregate Functions Import aggregate functions from `@nicia-ai/typegraph`: ```typescript import { count, countDistinct, sum, avg, min, max, field } from "@nicia-ai/typegraph"; ``` ### count Count rows in each group: ```typescript count("p") // COUNT(p.id) - count all nodes count("p", "department") // COUNT(p.props.department) - count non-null values ``` ### countDistinct Count unique values: ```typescript countDistinct("p") // COUNT(DISTINCT p.id) countDistinct("p", "department") // COUNT(DISTINCT p.props.department) ``` Distinct counts support string, number, Boolean, and date fields. Structured JSON, array, embedding, and unresolved dynamic fields are refused so SQLite and PostgreSQL cannot disagree about value equality. ### sum Sum numeric values: ```typescript sum("p", "salary") // SUM(p.props.salary) ``` ### avg Average of numeric values: ```typescript avg("p", "age") // AVG(p.props.age) ``` ### min / max Minimum and maximum values: ```typescript min("p", "hireDate") // MIN(p.props.hireDate) max("p", "salary") // MAX(p.props.salary) ``` Aggregate results follow JavaScript value conventions. `count()` and `countDistinct()` always return a number, including `0` for an empty input. `sum()`, `avg()`, `min()`, and `max()` return `undefined` when SQL produces `NULL`, such as an aggregate over an empty input. Minimum and maximum preserve the schema field's scalar type, so string results remain strings and date results are decoded as `Date` values. Minimum and maximum support string, number, and date fields; schema-known boolean and structured operands are refused. Sum and average require numeric fields. ### field Include a grouped field in the output: ```typescript field("p", "department") // The grouped field value field("c", "id") // Node ID field("c", "name") // Property value ``` ## Multiple Aggregations Combine multiple aggregates in one query: ```typescript import { count, countDistinct, sum, avg, min, max, field } from "@nicia-ai/typegraph"; const departmentStats = await store .query() .from("Employee", "e") .groupBy("e", "department") .aggregate({ department: field("e", "department"), headcount: count("e"), uniqueRoles: countDistinct("e", "role"), avgSalary: avg("e", "salary"), minSalary: min("e", "salary"), maxSalary: max("e", "salary"), totalPayroll: sum("e", "salary"), }) .execute(); ``` ## Grouping by Multiple Fields Chain `groupBy()` calls for multi-column grouping: ```typescript const breakdown = await store .query() .from("Employee", "e") .groupBy("e", "department") .groupBy("e", "level") .aggregate({ department: field("e", "department"), level: field("e", "level"), count: count("e"), avgSalary: avg("e", "salary"), }) .execute(); // Result: [ // { department: "Engineering", level: "Senior", count: 15, avgSalary: 150000 }, // { department: "Engineering", level: "Junior", count: 8, avgSalary: 80000 }, // { department: "Sales", level: "Senior", count: 5, avgSalary: 120000 }, // ... // ] ``` ## Grouping by Node Use `groupByNode()` to group by unique nodes (by ID): ```typescript const projectContributions = await store .query() .from("Commit", "c") .traverse("author", "e") .to("Developer", "d") .groupByNode("d") // Group by developer node .aggregate({ developerId: field("d", "id"), developerName: field("d", "name"), commitCount: count("c"), }) .execute(); ``` ## Filtering Groups with HAVING Use `having()` to filter groups based on aggregate values (SQL's HAVING clause): ```typescript import { count, havingGt } from "@nicia-ai/typegraph"; // Only departments with more than 5 employees const largeDepartments = await store .query() .from("Employee", "e") .groupBy("e", "department") .having(havingGt(count("e"), 5)) // HAVING COUNT(e) > 5 .aggregate({ department: field("e", "department"), headcount: count("e"), }) .execute(); ``` ### Available HAVING Helpers ```typescript import { having, havingGt, havingGte, havingLt, havingLte, havingEq, } from "@nicia-ai/typegraph"; // Comparison helpers havingGt(aggregate, value) // > havingGte(aggregate, value) // >= havingLt(aggregate, value) // < havingLte(aggregate, value) // <= havingEq(aggregate, value) // = // Generic comparison (for custom operators) having(aggregate, "gt", value) ``` ### Multiple HAVING Conditions Chain multiple having conditions: ```typescript const qualifiedDepartments = await store .query() .from("Employee", "e") .groupBy("e", "department") .having(havingGte(count("e"), 5)) // At least 5 employees .having(havingGte(avg("e", "salary"), 100000)) // Average salary >= 100k .aggregate({ department: field("e", "department"), headcount: count("e"), avgSalary: avg("e", "salary"), }) .execute(); ``` ## Aggregations with Traversals Combine graph traversals with aggregations: ```typescript const topContributors = await store .query() .from("PullRequest", "pr") .whereNode("pr", (pr) => pr.state.eq("merged")) .traverse("targetsRepo", "e1") .to("Repository", "repo") .traverse("author", "e2", { direction: "in" }) .to("Developer", "dev") .groupBy("repo", "name") .groupBy("dev", "name") .aggregate({ repository: field("repo", "name"), developer: field("dev", "name"), prCount: count("pr"), linesChanged: sum("pr", "linesAdded"), }) .limit(50) .execute(); ``` ## Ordering Aggregated Results Aggregate queries have their own `orderBy(key, direction?)`, called after `.aggregate({...})`. `key` is any output name from the fields object — a grouped field or an aggregate alias — so `limit()` finally means "top N", not "an arbitrary N": ```typescript const topDepartments = await store .query() .from("Employee", "e") .groupBy("e", "department") .aggregate({ department: field("e", "department"), headcount: count("e"), totalSalary: sum("e", "salary"), }) .orderBy("totalSalary", "desc") .limit(10) .execute(); ``` ## Real-World Example: Team Analytics ```typescript import { count, countDistinct, sum, avg, field, havingGt } from "@nicia-ai/typegraph"; // 1. Productivity by department const departmentMetrics = await store .query() .from("Developer", "dev") .traverse("authored", "e") .to("PullRequest", "pr") .whereNode("pr", (pr) => pr.state.eq("merged")) .groupBy("dev", "department") .aggregate({ department: field("dev", "department"), developerCount: countDistinct("dev"), totalPRs: count("pr"), totalLinesAdded: sum("pr", "linesAdded"), avgLinesPerPR: avg("pr", "linesAdded"), }) .execute(); // 2. Active reviewers (reviewed > 10 PRs) const activeReviewers = await store .query() .from("Developer", "d") .traverse("reviewed", "r") .to("PullRequest", "pr") .groupByNode("d") .having(havingGt(count("pr"), 10)) .aggregate({ developer: field("d", "name"), reviewCount: count("pr"), }) .orderBy("reviewCount", "desc") .execute(); // 3. Repository health const repoHealth = await store .query() .from("Repository", "r") .traverse("contains", "e") .to("PullRequest", "pr") .groupByNode("r") .aggregate({ repo: field("r", "name"), openPRs: count("pr"), avgAge: avg("pr", "daysOpen"), }) .execute(); ``` ## Next Steps - [Shape](/queries/shape) - Output transformation with `select()` - [Order](/queries/order) - Ordering and limiting results - [Traverse](/queries/traverse) - Graph traversals # Combine > Set operations with union(), intersect(), and except() Combine operations merge results from multiple queries using set operations. Use `union()` to combine results, `intersect()` to find common results, and `except()` to exclude results. For new SQL projections, use [`project().asRelation()`](/queries/relations/) to combine visible output columns, then order, filter, prepare, or batch the combined relation. The `select()` examples below describe the compatibility API; its JavaScript result mapper does not define SQL row equality. ## Set Operations Overview | Operation | Description | Duplicates | |-----------|-------------|------------| | `union()` | Combine results from both queries | Removed | | `unionAll()` | Combine results from both queries | Kept | | `intersect()` | Results that appear in both queries | Removed | | `except()` | Results in first query but not second | Removed | ## union() Combine results from multiple queries, removing duplicates: ```typescript const activeOrAdmin = await store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })) .union( store .query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("admin")) .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })) ) .execute(); ``` This returns all active users PLUS all admins, with duplicates removed (active admins appear once). ### Selection Shape Must Match Both queries must have the same selection shape: ```typescript // Valid: Same shape query1.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })) .union( query2.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })) ) // Invalid: Different shapes - will cause an error query1.select((ctx) => ({ id: ctx.p.id })) .union( query2.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })) ) ``` ## unionAll() Combine results keeping duplicates: ```typescript const allMentions = await store .query() .from("Comment", "c") .whereNode("c", (c) => c.mentions.contains(userId)) .select((ctx) => ({ id: ctx.c.id, text: ctx.c.text })) .unionAll( store .query() .from("Post", "p") .whereNode("p", (p) => p.mentions.contains(userId)) .select((ctx) => ({ id: ctx.p.id, text: ctx.p.content })) ) .execute(); ``` Use `unionAll()` when: - You want to preserve duplicates - Performance matters (no deduplication overhead) - You're counting occurrences ## intersect() Find results that appear in both queries: ```typescript const activeAdmins = await store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ({ id: ctx.p.id })) .intersect( store .query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("admin")) .select((ctx) => ({ id: ctx.p.id })) ) .execute(); ``` This returns only users who are BOTH active AND admins. ### Equivalent to AND `intersect()` can often be replaced with combined predicates: ```typescript // Using intersect query1.intersect(query2) // Often equivalent to .whereNode("p", (p) => p.status.eq("active").and(p.role.eq("admin")) ) ``` Use `intersect()` when the queries are complex or involve different traversal paths. ## except() Find results in the first query but not the second (set difference): ```typescript const nonAdminActive = await store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ({ id: ctx.p.id })) .except( store .query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("admin")) .select((ctx) => ({ id: ctx.p.id })) ) .execute(); ``` This returns active users who are NOT admins. ### Order Matters Unlike `union()` and `intersect()`, the order of queries in `except()` matters: ```typescript // Active users who are NOT admins activeUsers.except(admins) // Admins who are NOT active (different result!) admins.except(activeUsers) ``` ## Chaining Set Operations Chain multiple set operations: ```typescript const complexSet = await store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ({ id: ctx.p.id })) .union( store.query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("admin")) .select((ctx) => ({ id: ctx.p.id })) ) .except( store.query() .from("Person", "p") .whereNode("p", (p) => p.suspended.eq(true)) .select((ctx) => ({ id: ctx.p.id })) ) .execute(); // (active OR admin) AND NOT suspended ``` ## Ordering and Limiting Combined Results Apply ordering and limits after set operations: ```typescript const results = await query1 .union(query2) .orderBy("name", "asc") .limit(100) .execute(); ``` ## Real-World Examples ### Multi-Source Search Search across different node types: ```typescript import { expr } from "@nicia-ai/typegraph"; async function globalSearch(term: string) { const people = store .query() .from("Person", "p") .whereNode("p", (p) => p.name.ilike(`%${term}%`)) .project((fields) => ({ id: fields.p.id, type: expr.literal("person" as string), title: fields.p.name, })) .asRelation(); const companies = store .query() .from("Company", "c") .whereNode("c", (c) => c.name.ilike(`%${term}%`)) .project((fields) => ({ id: fields.c.id, type: expr.literal("company" as string), title: fields.c.name, })) .asRelation(); return people .union(companies) .limit(20) .execute(); } ``` ### Exclude Blocklist ```typescript const eligibleUsers = await store .query() .from("User", "u") .whereNode("u", (u) => u.status.eq("active")) .select((ctx) => ({ id: ctx.u.id, email: ctx.u.email })) .except( store .query() .from("BlockedUser", "b") .traverse("blockedUser", "e") .to("User", "u") .select((ctx) => ({ id: ctx.u.id, email: ctx.u.email })) ) .execute(); ``` ### Find Common Connections ```typescript async function mutualFriends(userId1: string, userId2: string) { const user1Friends = store .query() .from("Person", "p") .whereNode("p", (p) => p.id.eq(userId1)) .traverse("follows", "e") .to("Person", "friend") .select((ctx) => ({ id: ctx.friend.id, name: ctx.friend.name })); const user2Friends = store .query() .from("Person", "p") .whereNode("p", (p) => p.id.eq(userId2)) .traverse("follows", "e") .to("Person", "friend") .select((ctx) => ({ id: ctx.friend.id, name: ctx.friend.name })); return user1Friends .intersect(user2Friends) .execute(); } ``` ### Deduplicate Recursive Results Remove duplicate nodes from recursive traversals by projecting proven node identities: ```typescript // Get unique reachable nodes (recursive may return duplicates via different paths) const uniqueNodes = await store .query() .from("Node", "start") .traverse("linkedTo", "e") .recursive() .to("Node", "reachable") .project((fields) => ({ kind: fields.reachable.kind, id: fields.reachable.id, })) .asRelation() .distinctNodes({ kind: "kind", id: "id" }) .execute(); ``` `distinctNodes()` accepts only the `kind` and `id` columns from one proven node binding. This avoids choosing an arbitrary edge, path, or payload value when several matches reach the same node. ## Using Set Operations with batch() Set operation queries implement the `BatchableQuery` interface, so you can include them in `store.batch()` alongside regular queries: ```typescript const [adminOrOwner, companies] = await store.batch( store .query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("admin")) .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })) .union( store .query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("owner")) .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })), ), store .query() .from("Company", "c") .select((ctx) => ({ id: ctx.c.id, name: ctx.c.name })), ); ``` See [Batch Query Execution](/schemas-stores#batch-query-execution) for details. ## Next Steps - [Advanced](/queries/advanced) - Subqueries with `exists()` and `inSubquery()` - [Execute](/queries/execute) - Running queries - [Compose](/queries/compose) - Reusable query fragments # Compose > Reusable query transformations with pipe() and fragment composition Compose operations let you create reusable query transformations. Use `pipe()` to apply transformations and `createFragment()` to build typed, composable query parts. ## The pipe() Method Apply a transformation function to a query builder: ```typescript const results = await store .query() .from("User", "u") .pipe((q) => q.whereNode("u", ({ status }) => status.eq("active"))) .pipe((q) => q.orderBy("u", "createdAt", "desc")) .select((ctx) => ctx.u) .execute(); ``` Each `pipe()` receives the current builder and returns a modified builder, enabling chained transformations. ## Defining Reusable Fragments Extract common patterns into reusable functions: ```typescript // Define reusable fragments const activeOnly = (q) => q.whereNode("u", ({ status }) => status.eq("active")); const recentFirst = (q) => q.orderBy("u", "createdAt", "desc"); const first10 = (q) => q.limit(10); // Use in queries const results = await store .query() .from("User", "u") .pipe(activeOnly) .pipe(recentFirst) .pipe(first10) .select((ctx) => ctx.u) .execute(); ``` ## Typed Fragments with createFragment() For full type safety, use the `createFragment()` factory: ```typescript import { createFragment } from "@nicia-ai/typegraph"; // Create a typed fragment factory for your graph const fragment = createFragment(); // Define typed fragments const activeUsers = fragment((q) => q.whereNode("u", ({ status }) => status.eq("active")) ); const withRecentPosts = fragment((q) => q.traverse("authored", "a") .to("Post", "p") .whereNode("p", ({ createdAt }) => createdAt.gte("2024-01-01")) ); // Compose into queries const results = await store .query() .from("User", "u") .pipe(activeUsers) .pipe(withRecentPosts) .select((ctx) => ({ user: ctx.u, post: ctx.p, })) .execute(); ``` ## Composing Fragments Use `composeFragments()` to combine multiple fragments into one: ```typescript import { composeFragments, limitFragment, orderByFragment } from "@nicia-ai/typegraph"; // Compose multiple fragments into one const paginatedActiveUsers = composeFragments( (q) => q.whereNode("u", ({ status }) => status.eq("active")), (q) => q.orderBy("u", "createdAt", "desc"), (q) => q.limit(20) ); // Apply as a single transformation const results = await store .query() .from("User", "u") .pipe(paginatedActiveUsers) .select((ctx) => ctx.u) .execute(); ``` ## Helper Fragments TypeGraph provides pre-built helper fragments: ```typescript import { limitFragment, offsetFragment, orderByFragment, composeFragments } from "@nicia-ai/typegraph"; // Pre-built fragments const paginated = composeFragments( orderByFragment("u", "createdAt", "desc"), limitFragment(20), offsetFragment(40) ); const results = await store .query() .from("User", "u") .pipe(paginated) .select((ctx) => ctx.u) .execute(); ``` ### Available Helpers | Helper | Description | |--------|-------------| | `limitFragment(n)` | Limits results to n rows | | `offsetFragment(n)` | Skips the first n rows | | `orderByFragment(alias, field, direction)` | Orders by a field | ## Fragments with Traversals Fragments can include traversals: ```typescript // Fragment that adds a manager traversal const withManager = fragment((q) => q.traverse("reportsTo", "r").to("User", "manager") ); // Fragment that adds department info const withDepartment = fragment((q) => q.traverse("belongsTo", "b").to("Department", "dept") ); // Compose for a complete employee view const employeeDetails = composeFragments(withManager, withDepartment); const results = await store .query() .from("User", "u") .pipe(employeeDetails) .select((ctx) => ({ employee: ctx.u, manager: ctx.manager, department: ctx.dept, })) .execute(); ``` ## Post-Select Fragments `pipe()` is also available on `ExecutableQuery`: ```typescript // Define a pagination fragment for executable queries const paginate = (q) => q.orderBy("u", "name", "asc").limit(10).offset(20); const results = await store .query() .from("User", "u") .select((ctx) => ({ name: ctx.u.name, email: ctx.u.email })) .pipe(paginate) .execute(); ``` ## Real-World Patterns ### Search with Conditional Filters ```typescript function searchUsers(filters: { status?: string; role?: string; search?: string; }) { let query = store.query().from("User", "u"); // Apply filters conditionally using pipe if (filters.status) { query = query.pipe((q) => q.whereNode("u", ({ status }) => status.eq(filters.status)) ); } if (filters.role) { query = query.pipe((q) => q.whereNode("u", ({ role }) => role.eq(filters.role)) ); } if (filters.search) { query = query.pipe((q) => q.whereNode("u", ({ name }) => name.ilike(`%${filters.search}%`)) ); } return query.select((ctx) => ctx.u).execute(); } ``` ### Configurable Pagination ```typescript function createPaginationFragment(options: { sortField: string; sortDir: "asc" | "desc"; page: number; pageSize: number; }) { return composeFragments( orderByFragment("u", options.sortField, options.sortDir), limitFragment(options.pageSize), offsetFragment((options.page - 1) * options.pageSize) ); } // Use with any query const pagination = createPaginationFragment({ sortField: "createdAt", sortDir: "desc", page: 2, pageSize: 25, }); const results = await store .query() .from("User", "u") .pipe(pagination) .select((ctx) => ctx.u) .execute(); ``` ### Domain-Specific Query Helpers ```typescript // Create domain-specific query helpers const userQueries = { active: (q) => q.whereNode("u", ({ status }) => status.eq("active")), verified: (q) => q.whereNode("u", ({ emailVerified }) => emailVerified.eq(true)), withRole: (role: string) => (q) => q.whereNode("u", ({ role: r }) => r.eq(role)), withPosts: (q) => q.traverse("authored", "a").to("Post", "p"), recentlyActive: (q) => q.whereNode("u", ({ lastLogin }) => lastLogin.gte(new Date(Date.now() - 7 * 24 * 60 * 60 * 1000).toISOString()) ), }; // Compose for specific use cases const activeAdmins = await store .query() .from("User", "u") .pipe(userQueries.active) .pipe(userQueries.verified) .pipe(userQueries.withRole("admin")) .select((ctx) => ctx.u) .execute(); ``` ## Type Definitions For advanced use cases, TypeGraph exports fragment type definitions: ```typescript import type { QueryFragment, FlexibleQueryFragment, TraversalFragment } from "@nicia-ai/typegraph"; ``` - **`QueryFragment`** - A typed fragment transformation - **`FlexibleQueryFragment`** - A fragment that works with any compatible builder - **`TraversalFragment`** - A fragment for transforming TraversalBuilder instances ## Next Steps - [Filter](/queries/filter) - Filtering with predicates - [Traverse](/queries/traverse) - Graph traversals - [Combine](/queries/combine) - Set operations # Execute > Running queries with execute(), paginate(), and stream() Execute operations run your query and retrieve results. Use `execute()` for simple queries, `paginate()` for cursor-based pagination, and `stream()` for processing large datasets. ## execute() Run the query and return all results: ```typescript const results = await store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ctx.p) .execute(); // results: readonly Person[] ``` ### Return Type Returns a readonly array of the selected type: ```typescript // TypeScript infers the shape from your selection const results = await store .query() .from("Person", "p") .select((ctx) => ({ name: ctx.p.name, email: ctx.p.email, })) .execute(); // results: readonly { name: string; email: string | undefined }[] ``` ## executeChecked(expectedSchemaVersion) Check a cached reconciled schema while reading data in one SQL statement: ```typescript const rows = await store.query() .from("Person", "person") .select((ctx) => ctx.person) .executeChecked(store.reconciledSchema.version); ``` The active schema version and data come from the same statement snapshot. A mismatch throws `SchemaChangedError` with `details.graphId`, `details.expected`, and `details.actual` before calling the selector, even when the data query matches no rows. `undefined` means no active schema; it is distinct from version zero. On mismatch, reload the reconciled schema, rebuild the query against the reopened store, and retry. Retry in a new transaction if the old one holds a repeatable-read snapshot. This is an explicit alternative to a standalone `getCommittedSchemaVersion` probe on the first relational query. It checks that statement only; subsequent request reads can observe later commits. It neither locks the schema nor replaces write fences, and does not alter store-open or application cache policies. Checked reads fetch full rows and support relational traversals, ordering, offsets, and limits. Recursive and relevance-ranked queries are refused with `ConfigurationError`; use a separate probe for those. Named parameters must be bound as ordinary values before building the query. Bundled SQLite and PostgreSQL backends provide the required `tableNames.schemaVersions` binding. A custom backend without it is refused before executing SQL. A custom binding must name a relation with the standard `graph_id`, `version`, and `is_active` columns and one active row per graph, consistent with `getActiveSchema`. ## first() Get the first selected result or `undefined`. An existing `limit(0)` remains empty; `offset()` is preserved. Add `orderBy()` when the choice of first row must be deterministic. Only the returned row is passed to the selector: ```typescript const alice = await store .query() .from("Person", "p") .whereNode("p", (p) => p.email.eq("alice@example.com")) .select((ctx) => ctx.p) .first(); if (alice) { console.log(alice.name); } ``` ## count() Count matching SQL rows without fetching their data or running a `select()` callback. `count()` and `exists()` are available before and after `select()`. They preserve `groupBy()`, `having()`, `limit()`, and `offset()`: a grouped query counts groups, and `limit(0).count()` returns zero. Traversals count match rows, so multiple relationships can count the same node more than once. Bind named parameters as concrete values before using these terminals: ```typescript const activeCount = await store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .count(); // activeCount: number ``` ## exists() Check if any results exist: ```typescript const hasActiveUsers = await store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .exists(); // hasActiveUsers: boolean ``` ## Cursor Pagination Use `first`/`after` for forward pages or `last`/`before` for backward pages. Page sizes must be positive safe integers. Do not combine directions or add query-level `limit()`/`offset()` to a paginated or streamed query; those bounds are refused rather than silently discarded. Use `limit()`/`offset()` with `execute()` for offset pagination. For large datasets, cursor-based pagination is more efficient than `limit`/`offset`. It uses keyset pagination which doesn't degrade as you go deeper. ### paginate() Nullable sort values follow the same ordering across page boundaries as in `execute()`: ascending order places missing values last, and descending order places them first. Forward and backward cursors retain rows in both the missing-value and non-missing-value groups. Sort fields do not have to appear in the selected result; pagination retains them internally for its cursors. ```typescript const firstPage = await store .query() .from("Person", "p") .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name, })) .orderBy("p", "name", "asc") // ORDER BY required .paginate({ first: 20 }); ``` ### Pagination Result Shape ```typescript { data: readonly T[], // The actual results hasNextPage: boolean, // More results available forward hasPrevPage: boolean, // More results available backward nextCursor: string | undefined, // Opaque cursor for next page prevCursor: string | undefined, // Opaque cursor for previous page } ``` ### Forward Pagination Use `first` and `after` to paginate forward: ```typescript // Get first page const page1 = await query.paginate({ first: 20 }); // Get next page using the cursor if (page1.hasNextPage && page1.nextCursor) { const page2 = await query.paginate({ first: 20, after: page1.nextCursor, }); } ``` ### Backward Pagination Use `last` and `before` to paginate backward: ```typescript // Get last page const lastPage = await query.paginate({ last: 20 }); // Get previous page if (lastPage.hasPrevPage && lastPage.prevCursor) { const prevPage = await query.paginate({ last: 20, before: lastPage.prevCursor, }); } ``` ### Batch cursor pages with `page()` `paginate()` executes immediately. Use `page()` to build the same cursor-bounded read without executing it, so the page can compose with other independent reads in one `batchOnce()` statement: ```typescript const people = store .query() .from("Person", "person") .orderBy("person", "name") .select((fields) => fields.person); const [page, companies] = await store.batchOnce(() => [ people.page({ first: 20, after: cursor }), store .query() .from("Company", "company") .orderBy("company", "name") .select((fields) => fields.company), ]); ``` The returned page has the same `PaginatedResult` shape as `paginate()`. A page read also has an `execute()` method for independent execution. As with every `batchOnce()` member, all reads must belong to the same graph and execution target, and the combined statement must fit the backend's bind-parameter budget. ### Pagination Parameters | Parameter | Type | Description | | --------- | -------- | -------------------------------------------- | | `first` | `number` | Number of results from the start | | `after` | `string` | Cursor to start after (forward pagination) | | `last` | `number` | Number of results from the end | | `before` | `string` | Cursor to start before (backward pagination) | ### Pagination with Traversals Pagination works with graph traversals: ```typescript const employeesPage = await store .query() .from("Company", "c") .whereNode("c", (c) => c.name.eq("Acme Corp")) .traverse("worksAt", "e", { direction: "in" }) .to("Person", "p") .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name, role: ctx.e.role, })) .orderBy("p", "name", "asc") .paginate({ first: 50 }); ``` ## Streaming For very large datasets, use streaming to process results without loading everything into memory. ### stream() ```typescript const stream = store .query() .from("Event", "e") .select((ctx) => ctx.e) .orderBy("e", "createdAt", "desc") // ORDER BY required .stream({ batchSize: 1000 }); // Process results as they arrive for await (const event of stream) { console.log(event.title); await processEvent(event); } ``` ### Batch Size The `batchSize` option controls how many records are fetched per database query: ```typescript // Smaller batches: Lower memory usage, more database queries .stream({ batchSize: 100 }) // Larger batches: Higher memory usage, fewer database queries .stream({ batchSize: 5000 }) // Default is 1000 .stream() ``` ### Streaming with Processing ```typescript async function exportAllUsers(): Promise { const stream = store .query() .from("User", "u") .whereNode("u", (u) => u.status.eq("active")) .select((ctx) => ({ id: ctx.u.id, email: ctx.u.email, name: ctx.u.name, })) .orderBy("u", "id", "asc") .stream({ batchSize: 500 }); let count = 0; for await (const user of stream) { await exportToExternalSystem(user); count++; if (count % 1000 === 0) { console.log(`Exported ${count} users...`); } } console.log(`Export complete: ${count} users`); } ``` ## Batch Execution When independent reads must share one database round trip, use `store.batchOnce()`. It embeds each read as a CTE and returns the independently typed results in input order. Fluent queries preserve explicit ordering even when the sort field is not selected. The callback's scoped builder creates batch-scoped composable graph reads without adding parallel `*Query` methods to the executing Store API: ```typescript const [people, neighbors, neighborhood] = await store.batchOnce((read) => [ store.query().from("Person", "p").select((ctx) => ctx.p), read.neighbors(person, { edges: ["knows"], limit: 5 }), read.subgraph(person.id, { edges: ["knows"], maxDepth: 2 }), ]); ``` The callback can also return `roots.map(...)`, a singleton, or an empty array. A nonempty batch is one statement with no sequential fallback; an empty batch executes no SQL. At most 500 reads may be planned, and the combined statement must fit the backend's bind-parameter budget. Every response is materialized as JSON rather than streamed, so bound each member's result explicitly. Response size is data-dependent; TypeGraph neither estimates it nor imposes a response-byte cap before execution. For several independent subgraphs, use the runtime-array form to collapse their database round trips into one statement: ```typescript const subgraphs = await store.batchOnce((read) => roots.map((root) => read.subgraph(root.id, { edges: ["knows"], maxDepth: 2, project: { nodes: { Person: ["name"] } }, }), ), ); ``` When compatible subgraphs have substantially overlapping neighborhoods and project meaningful payloads, opt into shared traversal and hydration: ```typescript const subgraphs = await store.batchOnce( (read) => roots.map((root) => read.subgraph(root.id, { edges: ["knows"], maxDepth: 2, project: { nodes: { Person: ["name", "profile"] } }, }), ), { shareSubgraphs: true }, ); ``` The option groups only compatible subgraph reads and hydrates a shared entity once while preserving an independent result object for every request. The default remains independent subgraph plans in the same one statement. Sharing adds membership and reconstruction overhead, so enable it for measured overlap and payload shapes rather than assuming it is universally faster. See [shared subgraph examples](/performance/overview#choosing-shared-subgraphs) for overlapping biographies, disjoint neighborhoods, and identity-only results with different tradeoffs. Tuple members may use different roots, edge sets, depths, windows, and projections when a page needs heterogeneous neighborhoods: ```typescript const [social, employment] = await store.batchOnce((read) => [ read.subgraph(person.id, { edges: ["knows"], maxDepth: 2, edgeWindows: { knows: { limit: 20 } }, }), read.subgraph(person.id, { edges: ["worksAt"], maxDepth: 1, project: { nodes: { Company: ["name", "industry"] }, edges: { worksAt: ["role"] }, }, }), ]); ``` The one-statement guarantee reduces round trips, which is often valuable for remote databases. By default it does not combine recursive plans or share hydration between overlapping subgraphs, and it does not promise less database work than direct `store.subgraph()` calls. The explicit `shareSubgraphs` option changes that planning choice for compatible subgraph members only. Use `store.batch()` when the batch includes queued edge collection `batchFind*` reads or when sequential execution is the intended connection profile. `batch()` does not batch round trips. The portable guarantee is that at most one query is in flight at a time — at least one statement each, and two for a query whose selective-field mapping falls back after its statement has already run. On a SQL backend with transactions it frames them with `begin`/`commit`, putting a networked one at N+2 round trips **at best**; Durable Objects use an ambient storage transaction with no framing, and without transactions there is no framing at all. Connection reuse is the adapter's business either way. Whole-node, whole-edge, and spread selections detected during planning use a full fetch from the start. A selector branch that depends on actual row values can still trigger the fallback. It will not merge arbitrary promises or collection calls. Use fluent queries or the callback's `read.neighbors()`, `read.countNeighbors()`, and `read.subgraph()` methods when independent result shapes must share its one statement. Other alternatives are a `.traverse()` chain, `store.neighbors()` or `store.countNeighbors()` (one statement each), `store.subgraph()` (2 statements on SQLite, 3 on PostgreSQL), or `getByIds()` / `bulkFindByIndex()`, which are chunked rather than fixed-cost. Direct `store.subgraph()` and batch-scoped `read.subgraph()` share validation, traversal, projection, and result semantics. The call context selects the physical execution contract: the direct read uses backend-tuned hydration, while the batch-scoped read is embedded into the batch's single statement. It is also **not** a snapshot: PostgreSQL defaults to read-committed isolation, so a later query can observe a commit the earlier ones did not. When several reads need one stable snapshot, run `tx.query()`, `tx.neighbors()`, `tx.countNeighbors()`, `tx.subgraph()`, or `tx.batchOnce()` inside `store.transaction(fn, { isolationLevel: "repeatable_read" })`. These reads are bound to the open transaction and see its earlier uncommitted writes. Transactions require a backend with interactive transaction support; a history-enabled store on PostgreSQL additionally requires `accessMode: "read_only"` for a read-only transaction. ```typescript const [people, companies] = await store.batch( store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })), store .query() .from("Company", "c") .select((ctx) => ({ id: ctx.c.id, name: ctx.c.name })) .orderBy("c", "name", "asc") .limit(5), ); // people: readonly { id: string; name: string }[] // companies: readonly { id: string; name: string }[] ``` Each query preserves its own projection, filtering, sorting, and pagination. Results are returned as a typed tuple matching the input order. Edge collection `batchFind*` methods also return `BatchableQuery` and can be mixed freely with fluent queries — each still costs its own statement: ```typescript const [skills, employer] = await store.batch( store.edges.hasSkill.batchFindFrom(alice), store.edges.worksAt.batchFindFrom(alice), ); ``` **vs `Promise.all`**: workload- and adapter-dependent in both directions. `Promise.all` overlaps its queries against a pool with idle capacity, but it does not necessarily hold N connections, and against a single client or a saturated pool it queues. `batch()` keeps at most one query in flight, so it pays the sum of their latencies — but it can still come out ahead where connection acquisition dominates. Measure rather than assume. **vs `transaction()`**: `batch()` may open an internal transaction only to serialize its statements. Use `transaction()` when reads must share an explicit isolation level or see writes made earlier in the callback. Its context supports fluent and set-oriented reads; `tx.batchOnce()` still emits exactly one statement. See [Batch Query Execution](/schemas-stores#batch-query-execution) for full API reference. ## Prepared Queries Prepared queries let you build and structurally validate a query's AST once — so a malformed query fails fast, before the first `.execute()` — and execute it many times with different parameter values. ### `param(name)` Use `param()` to declare a named placeholder inside any predicate position: ```typescript import { param } from "@nicia-ai/typegraph"; ``` ### `prepare()` Call `.prepare()` on an executable query to build and validate the AST once. Returns a `PreparedQuery` that can be executed with different bindings. The statement is compiled once into a cached template and reused by every `.execute()` call — see [Prepared query SQL compilation](#prepared-query-sql-compilation) below for how that stays fresh. ```typescript const findByName = store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq(param("name"))) .select((ctx) => ctx.p) .prepare(); // Execute with different bindings const alices = await findByName.execute({ name: "Alice" }); const bobs = await findByName.execute({ name: "Bob" }); ``` ### Parameterized Bounds Parameters work anywhere a scalar value is accepted: ```typescript const findByAge = store .query() .from("Person", "p") .whereNode("p", (p) => p.age.between(param("minAge"), param("maxAge"))) .select((ctx) => ctx.p) .prepare(); const youngAdults = await findByAge.execute({ minAge: 18, maxAge: 25 }); const seniors = await findByAge.execute({ minAge: 65, maxAge: 120 }); ``` `prepared.execute(bindings)` validates bindings strictly: all declared parameters must be provided, and unknown binding keys are rejected. ### Supported Positions `param()` works with any scalar predicate: | Predicate | Example | | --------------------------- | ----------------------------------------- | | `eq` / `neq` | `p.name.eq(param("name"))` | | `gt` / `gte` / `lt` / `lte` | `p.age.gt(param("minAge"))` | | `between` | `p.age.between(param("lo"), param("hi"))` | | `contains` | `p.name.contains(param("substr"))` | | `startsWith` / `endsWith` | `p.name.startsWith(param("prefix"))` | | `like` / `ilike` | `p.email.like(param("pattern"))` | `in()` and `notIn()` take a list-valued parameter — the **whole** list, not individual elements: ```typescript const byIds = store .query() .from("Person", "p") .whereNode("p", (p) => p.id.in(param("ids"))) .select((ctx) => ctx.p) .prepare(); await byIds.execute({ ids: ["a", "b", "c"] }); await byIds.execute({ ids: ["d"] }); ``` The list is bound as a single parameter that the database unpacks, so the compiled SQL text does not depend on the list's length: one statement serves every arity, and a list of ten thousand ids still costs one bound parameter rather than blowing past the engine's bind limit. An empty list is valid — `in([])` matches nothing, `notIn([])` matches everything. Every element must be the field's type, and numbers must be finite. A mixed list — `[1, "a"]` bound against a number field — is rejected with a `ConfigurationError` before it reaches the database, on every backend, as is `NaN` or `Infinity`. This matches the literal form, which already refuses a mixed list, and it is what keeps the two backends in step: left unchecked, PostgreSQL would fail casting while SQLite silently matched nothing. :::caution A `param()` sitting among the **elements** of a literal list — `p.name.in(["Alice", param("other")])` — is rejected with an `UnsupportedPredicateError`. Bind the whole list instead. ::: ### Prepared Query SQL Compilation `.prepare()` builds and validates the AST once. On a backend that can compile and run raw SQL text (both the SQLite and PostgreSQL backends can), the statement is then compiled **once** into a cached template and reused by every `.execute()` call. The subtlety a cache like that has to survive is freshness: a "current" (live) read filters on temporal validity as of the instant it runs, so caching a compiled statement that had a concrete "now" baked into it would freeze that instant for the prepared query's entire lifetime — hiding every row created after `.prepare()` from every subsequent call. The template therefore reserves the read instant as a **placeholder** rather than a value, and each `.execute()` fills it with a fresh instant alongside the call's own bindings. Nothing about the statement's text depends on either. Two cases fall back to substituting parameters into the AST and compiling through the standard path on every call — same results and the same freshness guarantee, without the cached-template fast path: - `executeRaw` is unavailable (a custom or async backend). - The statement's execution semantics ride on the compiled SQL object rather than its text, which no amount of `executeRaw` support changes. **Approximate vector search** (`similarTo(..., { approximate: true })`) carries the engine's iterative-scan wrapper, and `store.subgraph()` on PostgreSQL forces a custom plan for its id-array fetches. Flattening either to cacheable text would drop the behavior it depends on, so both are excluded deliberately. ## Query Debugging ### toAst() Get the query AST for inspection: ```typescript const builder = store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ctx.p); const ast = builder.toAst(); console.log(JSON.stringify(ast, null, 2)); ``` ### compile() Use `toSQL()` to render SQL for the Store's configured dialect without executing it: ```typescript const compiled = builder.toSQL(); console.log("SQL:", compiled.sql); console.log("Parameters:", compiled.params); ``` For adapter and tooling authors, `builder.compile()` returns TypeGraph's database-independent `CompiledSelectSql` fragment. It can be passed to a `GraphBackend` or rendered explicitly with `renderSqlite()` or `renderPostgres()`. It is intentionally not a Drizzle `SQL` object. Useful for: - Debugging query behavior - Understanding performance characteristics - Building custom query executors ## Ordering Requirements Both `paginate()` and `stream()` require an `orderBy()` clause: ```typescript // Required for pagination .orderBy("p", "name", "asc") .paginate({ first: 20 }); // Required for streaming .orderBy("e", "createdAt", "desc") .stream(); ``` ### Stable Ordering Cursor pagination and streaming automatically append missing start-node identity keys: `id ASC` for a single kind, or `kind ASC` and `id ASC` for a multi-kind source. Existing caller-specified identity ordering is preserved. Offset pagination needs an explicit total ordering. For deterministic offset pagination of one kind, include `id` in your ordering: ```typescript .orderBy("p", "name", "asc") .orderBy("p", "id", "asc") // Ensures stable ordering ``` ## Real-World Examples ### Paginated API Endpoint ```typescript async function listUsers(cursor?: string, limit = 20) { const query = store .query() .from("User", "u") .whereNode("u", (u) => u.status.eq("active")) .select((ctx) => ({ id: ctx.u.id, name: ctx.u.name, email: ctx.u.email, })) .orderBy("u", "createdAt", "desc") .orderBy("u", "id", "desc"); const result = cursor ? await query.paginate({ first: limit, after: cursor }) : await query.paginate({ first: limit }); return { users: result.data, nextCursor: result.nextCursor, hasMore: result.hasNextPage, }; } ``` ### Batch Processing ```typescript async function processAllOrders() { const stream = store .query() .from("Order", "o") .whereNode("o", (o) => o.status.eq("pending")) .select((ctx) => ctx.o) .orderBy("o", "createdAt", "asc") .stream({ batchSize: 100 }); for await (const order of stream) { try { await fulfillOrder(order); await store.nodes.Order.update(order.id, { status: "fulfilled" }); } catch (error) { console.error(`Failed to process order ${order.id}:`, error); } } } ``` ### Infinite Scroll ```typescript function useInfiniteUsers() { const [users, setUsers] = useState([]); const [cursor, setCursor] = useState(); const [hasMore, setHasMore] = useState(true); async function loadMore() { const result = await store .query() .from("User", "u") .select((ctx) => ctx.u) .orderBy("u", "name", "asc") .paginate({ first: 20, after: cursor }); setUsers((prev) => [...prev, ...result.data]); setCursor(result.nextCursor); setHasMore(result.hasNextPage); } return { users, loadMore, hasMore }; } ``` ## Next Steps - [Order](/queries/order) - Ordering and limiting results - [Shape](/queries/shape) - Output transformation - [Overview](/queries/overview) - Query categories reference # Database Expressions > Type-safe filtering, calculation, projection, grouping, and ordering in SQL Database expressions describe work that TypeGraph sends to the database. They carry a value type, SQL nullability, and query scope, so TypeScript can reject incompatible operands and references to aliases outside the current query. ```typescript import { expr } from "@nicia-ai/typegraph"; const adults = await store .query() .from("Person", "p") .whereNode("p", (_person, e) => expr.gte(e.p.age, expr.literal(18)), ) .project((e) => ({ name: e.p.name, ageNextYear: expr.add(e.p.age, expr.literal(1)), })) .orderBy((e) => e.p.age, "desc") .execute(); ``` Expression callbacks are evaluated once while TypeGraph builds the query. Their expression trees are compiled to SQL; they do not run once per result row. ## Fields, literals, and parameters The callback context exposes every query alias. Node and edge schema properties appear directly on their alias, including nested object properties. System metadata lives under `$meta`. ```typescript .project((e) => ({ name: e.person.name, author: e.person.metadata.author, // `$get` reaches schema keys that share a name with expression metadata. nestedNode: e.person.metadata.$get("node"), validFrom: e.person.$meta.validFrom, })) ``` Object keys named `node`, `nullable`, `valueType`, `scopeIdentity`, `__type`, `__value`, or `__scope` share names with the expression wrapper. Access those declared JSON properties with `$get("key")`. Other declared keys support ordinary dot access. `$get` also accepts dynamic keys on record schemas. Use `expr.literal(value)` for a fixed value. Use `expr.param(name, valueType)` for a prepared value: ```typescript const query = store .query() .from("Person", "p") .whereNode("p", (_person, e) => expr.gt(e.p.age, expr.param("minimumAge", "number")), ) .project((e) => ({ name: e.p.name })); const prepared = query.prepare(); const rows = await prepared.execute({ minimumAge: 21 }); ``` ## Comparisons and Boolean expressions `expr.eq`, `neq`, `gt`, `gte`, `lt`, and `lte` require compatible scalar operands. Compose Boolean results with `expr.and`, `or`, and `not`. Use `isNull` and `isNotNull` for optional values. ```typescript .whereNode("p", (_person, e) => expr.and( expr.eq(e.p.active, expr.literal(true)), expr.or( expr.gte(e.p.score, expr.literal(90)), expr.isNull(e.p.score), ), ), ) ``` Comparisons with a nullable operand produce `boolean | undefined`, matching SQL's three-valued logic. Null checks always produce a Boolean. ## Array membership expressions `expr.arrayContains(array, element)` tests an array expression against another database expression. It is useful when the element comes from the candidate row or a correlated outer row; unlike the field accessor `tags.contains("value")`, it does not encode the element as a fixed JSON literal. Missing arrays, stored JSON `null`, and non-array JSON values simply do not match. Array elements must be scalar; structured elements are rejected because portable structural equality is not defined. Adapters that omit expression array membership support refuse this expression with a configuration error instead of silently changing its meaning. ```typescript .where((e) => expr.arrayContains(e.document.tags, e.person.id)) .where((e) => expr.arrayContains(e.document.tags, expr.literal("reference")), ) ``` ## Arithmetic and conversion Use `add`, `subtract`, `multiply`, and `divide` with numeric expressions. Division produces `undefined` when the divisor is zero. `toNumber` accepts numbers or strings in JSON number syntax: an optional minus sign, an integer (`0` or digits without a leading zero), an optional fraction with at least one digit, and an optional exponent with one to three digits. ASCII whitespace around the value is ignored. The trimmed input may contain at most 400 characters. Finite values use ordinary IEEE-754 rounding, including subnormal values; malformed input or a result outside the finite double range produces `undefined`. ```typescript .project((e) => ({ gross: expr.multiply(e.invoice.unitPrice, e.invoice.quantity), ratio: expr.divide(e.invoice.used, e.invoice.capacity), importedAmount: expr.toNumber(e.invoice.rawAmount), })) ``` These semantics are the same on SQLite and PostgreSQL. ## Coalescing and conditions `coalesce` returns the first non-null value. `when` selects between compatible result expressions. ```typescript .project((e) => ({ displayName: expr.coalesce(e.p.nickname, e.p.name), segment: expr.when( expr.gte(e.p.score, expr.literal(90)), expr.literal("priority"), expr.literal("standard"), ), })) ``` ## Projection and mapping Use `project()` to select database columns and computed values. Use `map()` for JavaScript work after each row is decoded. ```typescript const labels = await store .query() .from("Person", "p") .project((e) => ({ name: e.p.name, age: e.p.age })) .map((row) => `${row.name} (${row.age})`) .execute(); ``` `select()` remains the compatibility result-mapping API. Existing selectors keep their current execution behavior: compatibility planning may probe or retry a selector, so legacy selectors must be pure. New code that should run in SQL should use `project()` explicitly. A `project()` callback runs exactly once when the query is built, and a `map()` callback runs exactly once for each decoded result row. Use [`asRelation()`](/queries/relations/) to compose a SQL projection with set operations, derived filters and aggregates, whole-row distinctness, and output-column ordering. Keep `map()` at the end of SQL composition; mapped relations cannot become new SQL projections or set operands. ## Grouping and aggregates Expression callbacks also work with grouping, aggregate projections, ordering, and HAVING: ```typescript const totals = await store .query() .from("Person", "p") .groupBy((e) => [e.p.department]) .having((e) => expr.gt(expr.count(e.p.id), expr.literal(2))) .aggregate((e) => ({ department: e.p.department, people: expr.count(e.p.id), totalSalary: expr.sum(e.p.salary), })) .orderBy((e) => e.p.department, "asc") .execute(); ``` `count` and `countDistinct` return a number. `sum`, `avg`, `min`, and `max` include `undefined` in their result type because an empty input produces SQL `NULL`. `countDistinct` accepts string, number, Boolean, and date expressions. Arrays, objects, embeddings, and unknown dynamic values are refused because SQLite text equality and PostgreSQL JSON equality do not define the same distinct groups for structured values. For ordered scalar lists or flat records, use [`expr.collect()` on a relation](/queries/relations#ordered-collections). Collection ordering is explicit and independent of result-row ordering. ## Scope safety Each field expression belongs to the query scope that created it. TypeGraph refuses expression trees that mix unrelated scopes at runtime, and callback types prevent aliases from another query from being passed accidentally. Correlated subquery helpers provide an explicit outer context when an inner query needs an outer field; ordinary callbacks cannot capture one by alias name. Use `$exists()` for a correlated Boolean and `$scalar()` for one nullable value: ```ts const people = await store .query() .from("Person", "person") .project((e) => ({ hasNamesake: e.$exists((subquery, outer) => subquery .from("Person", "candidate") .whereNode("candidate", (_candidate, inner) => expr.eq(inner.candidate.name, outer.person.name), ) .project((inner) => ({ id: inner.candidate.id })), ), peerAge: e.$scalar((subquery, outer) => subquery .from("Person", "candidate") .whereNode("candidate", (_candidate, inner) => expr.eq(inner.candidate.name, outer.person.name), ) .project((inner) => ({ age: inner.candidate.age })) .limit(1), ), name: e.person.name, })) .execute(); ``` Both callbacks must return an explicitly projected query on the same graph, execution target, and temporal coordinate as the enclosing query. `$exists()` accepts any nonempty projection. `$scalar()` requires exactly one projected field and either `limit(1)` or an ungrouped aggregate, so behavior does not depend on an engine's handling of multiple scalar rows. A scalar with no matching row is decoded as `undefined`. Parameters inside either subquery participate in the enclosing prepared query's bindings. Expression builders do not accept raw SQL. This keeps parameter binding, decoding, and SQLite / PostgreSQL behavior on the same compiler path. ## Collection expression nodes `expr.collect(value, options)` returns `DatabaseExpression`. Reusable helpers can import `CollectOptions` for its required, nonempty `orderBy` tuple and optional `filter?: DatabaseExpression`. SQL TRUE includes an element; false and SQL NULL exclude it. An included NULL value remains `undefined` in the result, so filtering does not narrow the operand or result type. Collection expressions remain distinct `node.kind: "collect"` nodes, with `operand`, `orderBy`, and optional `filter`. Ordinary `"aggregate"` nodes retain their operator and optional operand; collection-only options do not appear on them. The collection value is either a scalar expression or an explicit flat record of named scalar expressions. A record remains present when every admitted field is SQL NULL, with those fields decoded as `undefined`. Ordering is required, and `distinct` and aggregate-local `limit` are not collection options. # Filter > Match constraints and completed-result filters Filter operations reduce the result set based on property values. TypeGraph provides `whereNode()` and `whereEdge()` for match constraints, plus `where()` for completed match rows. ## whereNode() Filter nodes based on their properties: ```typescript const engineers = await store .query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("Engineer")) .select((ctx) => ctx.p) .execute(); ``` ### Parameters ```typescript .whereNode(alias, predicateFunction) ``` | Parameter | Type | Description | | ------------------- | ------------------------- | ---------------------------------------------- | | `alias` | `string` | The node alias to filter (must exist in query) | | `predicateFunction` | `(accessor) => Predicate` | Function that returns a predicate | The predicate function receives a typed accessor for the node's properties. ## whereEdge() Filter based on edge properties during traversals: ```typescript const highPaying = await store .query() .from("Person", "p") .traverse("worksAt", "e") .whereEdge("e", (e) => e.salary.gte(100000)) .to("Company", "c") .select((ctx) => ({ person: ctx.p.name, company: ctx.c.name, salary: ctx.e.salary, })) .execute(); ``` ### Parameters ```typescript .whereEdge(alias, predicateFunction) ``` | Parameter | Type | Description | | ------------------- | ------------------------- | ---------------------------------------------- | | `alias` | `string` | The edge alias to filter (must exist in query) | | `predicateFunction` | `(accessor) => Predicate` | Function that returns a predicate | ## Combining Predicates ### AND Both conditions must be true: ```typescript .whereNode("p", (p) => p.status.eq("active").and(p.role.eq("admin")) ) ``` ### OR Either condition can be true: ```typescript .whereNode("p", (p) => p.role.eq("admin").or(p.role.eq("moderator")) ) ``` ### NOT Negate a condition: ```typescript .whereNode("p", (p) => p.status.eq("deleted").not() ) ``` ### Complex Combinations Build complex logic with parenthetical grouping: ```typescript .whereNode("p", (p) => p.status .eq("active") .and(p.role.eq("admin").or(p.role.eq("moderator"))) ) ``` This evaluates as: `status = 'active' AND (role = 'admin' OR role = 'moderator')` ## Multiple Filters Chain multiple `whereNode()` calls for AND logic: ```typescript const activeManagers = await store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .whereNode("p", (p) => p.role.eq("Manager")) .select((ctx) => ctx.p) .execute(); ``` This is equivalent to: ```typescript .whereNode("p", (p) => p.status.eq("active").and(p.role.eq("Manager")) ) ``` ## Filtering After Traversal Filter nodes at any point in the query: ```typescript const techCompanyEngineers = await store .query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("Engineer")) .traverse("worksAt", "e") .to("Company", "c") .whereNode("c", (c) => c.industry.eq("Technology")) .select((ctx) => ({ person: ctx.p.name, company: ctx.c.name, })) .execute(); ``` `whereNode("c", ...)` constrains the traversal match itself. For a recursive traversal it applies at every hop and prunes a branch as soon as a target fails. Use scoped-expression `where()` when intermediate nodes may fail the condition but a later endpoint should still be returned: ```typescript const activeEndpoints = await store .query() .from("Person", "start") .traverse("knows", "edge") .recursive({ maxHops: 5 }) .to("Person", "person") .where((fields) => expr.eq(fields.person.active, expr.literal(true))) .select((ctx) => ctx.person) .execute(); ``` On an optional alias, an ordinary comparison removes rows where that alias is absent. Use an explicit null check when absent matches should remain. ## Common Predicates Here are the most commonly used predicates. For complete reference, see [Predicates](/queries/predicates/). ### Equality ```typescript p.name.eq("Alice"); // equals p.name.neq("Bob"); // not equals ``` ### Comparison ```typescript p.age.gt(21); // greater than p.age.gte(21); // greater than or equal p.age.lt(65); // less than p.age.lte(65); // less than or equal p.age.between(18, 65); // inclusive range ``` ### String Matching ```typescript p.name.contains("ali"); // substring match p.name.startsWith("A"); // prefix match p.name.endsWith("ice"); // suffix match p.email.like("%@example.com"); // SQL LIKE pattern p.name.ilike("alice"); // case-insensitive LIKE ``` ### Fulltext Search For nodes with at least one field declared with `searchable()`, use the node-level `$fulltext.matches()` for BM25-style ranked fulltext search. See [Fulltext Search](/fulltext-search) for the full guide. ```typescript d.$fulltext.matches("climate change", 20); // Top 20 by relevance d.$fulltext.matches("quarterly earnings", 10, { mode: "websearch", // Google-style syntax }); ``` Combine with any other predicate — fulltext composes with metadata filters, graph traversal, and vector search: ```typescript d.$fulltext.matches("climate", 20).and(d.tenantId.eq(tenant)).and(d.published.eq(true)); ``` ### Null Checks ```typescript p.deletedAt.isNull(); // is null/undefined p.email.isNotNull(); // is not null ``` ### List Membership ```typescript p.status.in(["active", "pending"]); p.status.notIn(["archived", "deleted"]); ``` ### Array Operations ```typescript p.tags.contains("typescript"); p.tags.containsAll(["typescript", "nodejs"]); p.tags.containsAny(["typescript", "rust", "go"]); p.tags.isEmpty(); p.tags.isNotEmpty(); ``` ## Predicate Types by Field The available predicates depend on the field type: | Field Type | Key Predicates | | -------------------------------- | --------------------------------------------------- | | String | `eq`, `contains`, `startsWith`, `like`, `ilike` | | Nodes with `searchable()` fields | `$fulltext.matches()` (node-level, not per-field) | | Number | `eq`, `gt`, `gte`, `lt`, `lte`, `between` | | Date | `eq`, `gt`, `gte`, `lt`, `lte`, `between` | | Array | `contains`, `containsAll`, `containsAny`, `isEmpty` | | Object | `get()`, `hasKey`, `pathEquals` | | Embedding | `similarTo()` | See [Predicates](/queries/predicates/) for complete documentation. ## Count and Existence Helpers ### Count Results ```typescript const count: number = await store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .count(); ``` ### Check Existence ```typescript const exists: boolean = await store .query() .from("Person", "p") .whereNode("p", (p) => p.email.eq("alice@example.com")) .exists(); ``` ### Get First Result ```typescript const alice = await store .query() .from("Person", "p") .whereNode("p", (p) => p.email.eq("alice@example.com")) .select((ctx) => ctx.p) .first(); if (alice) { console.log(alice.name); } ``` ## Next Steps - [Predicates](/queries/predicates/) - Complete predicate reference - [Traverse](/queries/traverse) - Navigate relationships - [Advanced](/queries/advanced) - Subqueries with `exists()` and `inSubquery()` # Order > Control result ordering with orderBy(), limit(), and offset() Order operations control how results are sorted and how many are returned. Use `orderBy()` for sorting, `limit()` to cap results, and `offset()` for simple pagination. ## orderBy() Sort results by one or more fields: ```typescript const sorted = await store .query() .from("Person", "p") .select((ctx) => ctx.p) .orderBy((ctx) => ctx.p.name, "asc") .execute(); ``` ### Parameters ```typescript .orderBy(fieldSelector, direction?) .orderBy(alias, field, direction?) ``` | Parameter | Type | Description | |-----------|------|-------------| | `fieldSelector` | `(ctx) => field` | Function that selects the field to sort by | | `alias` | `string` | Node/edge alias (alternative syntax) | | `field` | `string` | Field name (alternative syntax) | | `direction` | `"asc" \| "desc"` | Sort direction (default: `"asc"`) | ### Single Field ```typescript // Function syntax .orderBy((ctx) => ctx.p.name, "asc") // Alias syntax .orderBy("p", "name", "asc") ``` ### Multiple Fields Chain `orderBy()` for multi-field sorting: ```typescript const sorted = await store .query() .from("Task", "t") .select((ctx) => ctx.t) .orderBy("t", "priority", "desc") // Primary sort .orderBy("t", "createdAt", "asc") // Secondary sort .execute(); ``` Or use the array syntax: ```typescript .orderBy((ctx) => [ { field: ctx.t.priority, direction: "desc" }, { field: ctx.t.createdAt, direction: "asc" }, ]) ``` ### Temporal System Columns The alias form can order directly by TypeGraph's physical temporal metadata: `valid_from`, `valid_to`, `created_at`, `updated_at`, and `deleted_at`. When the queried node or edge kinds do not declare a property with that name, it resolves to the physical system column: ```typescript const oldestFirst = await store .query() .from("Task", "t") .orderBy("t", "valid_from", "asc") .select((ctx) => ctx.t) .execute(); ``` The same classification applies whether `orderBy()` appears before or after `select()`. A declared property takes precedence over same-named temporal metadata. For example, if `Task` declares its own `created_at` property, `whereNode("t", (field) => field.created_at...)` and `orderBy("t", "created_at")` both address `props.created_at`. Rename the property when a query must order by TypeGraph's physical `created_at` column. ### Null Handling Control where null values appear: ```typescript .orderBy((ctx) => ({ field: ctx.p.email, direction: "asc", nulls: "last", // or "first" })) ``` ### Ordering by Edge Properties Order by properties on traversed edges: ```typescript const employees = await store .query() .from("Company", "c") .traverse("worksAt", "e", { direction: "in" }) .to("Person", "p") .select((ctx) => ({ name: ctx.p.name, startDate: ctx.e.startDate, })) .orderBy("e", "startDate", "desc") // Most recent hires first .execute(); ``` ### Ordering Aggregated Results Aggregate queries have their own `orderBy(key, direction?)`, called after `.aggregate({...})`. `key` is any output name from the fields object — a grouped field or an aggregate alias: ```typescript import { count, field } from "@nicia-ai/typegraph"; const topDepartments = await store .query() .from("Employee", "e") .groupBy("e", "department") .aggregate({ department: field("e", "department"), headcount: count("e"), }) .orderBy("headcount", "desc") .execute(); ``` ## limit() Cap the number of results returned: ```typescript const top10 = await store .query() .from("Person", "p") .select((ctx) => ctx.p) .orderBy("p", "score", "desc") .limit(10) .execute(); ``` ### Parameters ```typescript .limit(n) ``` | Parameter | Type | Description | |-----------|------|-------------| | `n` | `number` | Maximum number of results to return | ## offset() Skip a number of results (useful for simple pagination): ```typescript const page2 = await store .query() .from("Person", "p") .select((ctx) => ctx.p) .orderBy("p", "name", "asc") .limit(10) .offset(10) // Skip first 10 results .execute(); ``` ### Parameters ```typescript .offset(n) ``` | Parameter | Type | Description | |-----------|------|-------------| | `n` | `number` | Number of results to skip | ## Simple Pagination with limit/offset ```typescript async function getPage(pageNumber: number, pageSize: number) { return store .query() .from("Person", "p") .select((ctx) => ctx.p) .orderBy("p", "name", "asc") .limit(pageSize) .offset((pageNumber - 1) * pageSize) .execute(); } // Usage const page1 = await getPage(1, 20); // Results 1-20 const page2 = await getPage(2, 20); // Results 21-40 ``` > **Note:** For large datasets, use [cursor pagination](/queries/execute#cursor-pagination) instead. > Offset-based pagination becomes slower as offset increases. ## Ordering Requirements ### For Pagination Both `paginate()` and `stream()` require an `orderBy()` clause: ```typescript // Required for pagination const page = await store .query() .from("Person", "p") .select((ctx) => ctx.p) .orderBy("p", "name", "asc") // Required .paginate({ first: 20 }); // Required for streaming const stream = store .query() .from("Event", "e") .select((ctx) => ctx.e) .orderBy("e", "createdAt", "desc") // Required .stream(); ``` ### Stable Ordering Cursor pagination and streaming automatically append missing start-node identity keys: `id ASC` for a single kind, or `kind ASC` and `id ASC` for a multi-kind source. Existing caller-specified identity ordering is preserved. Offset pagination needs an explicit total ordering. For deterministic offset pagination of one kind, include `id` in your ordering: ```typescript .orderBy("p", "name", "asc") .orderBy("p", "id", "asc") // Ensures stable ordering when names are equal ``` ## Real-World Examples ### Leaderboard ```typescript const leaderboard = await store .query() .from("Player", "p") .select((ctx) => ({ name: ctx.p.name, score: ctx.p.score, })) .orderBy("p", "score", "desc") .limit(100) .execute(); ``` ### Recent Activity Feed ```typescript const feed = await store .query() .from("Activity", "a") .whereNode("a", (a) => a.userId.eq(currentUserId)) .select((ctx) => ctx.a) .orderBy("a", "createdAt", "desc") .limit(50) .execute(); ``` ### Paginated Search Results ```typescript async function searchProducts(query: string, page: number) { const pageSize = 20; return store .query() .from("Product", "p") .whereNode("p", (p) => p.name.ilike(`%${query}%`)) .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name, price: ctx.p.price, })) .orderBy("p", "relevance", "desc") .orderBy("p", "id", "asc") .limit(pageSize) .offset((page - 1) * pageSize) .execute(); } ``` ## Next Steps - [Execute](/queries/execute) - Cursor pagination and streaming - [Shape](/queries/shape) - Output transformation - [Filter](/queries/filter) - Reducing results with predicates ## Newest target through an edge When the source node is already available, `store.neighbors()` is the direct one-statement read: ```typescript const [latest] = await store.neighbors(document, { edges: ["hasVersion"], orderBy: { by: "node", field: "sequence", direction: "desc" }, limit: 1, }); ``` Use `by: "edge"` (or omit `by`) for edge metadata and `by: "node"` for an adjacent-node property. Both forms sort nulls last and use the edge ID as the deterministic final tie-breaker. The equivalent fluent traversal is useful when additional target predicates or projections are needed: ```typescript const [newest] = await store.query() .from("Document", "document") .whereNode("document", (node) => node.id.eq(documentId)) .traverse("hasVersion", "edge").to("Version", "version") .orderBy("version", "sequence", "desc") .orderBy("version", "id", "desc") .select((ctx) => ctx.version) .limit(1) .execute(); ``` The ID order resolves ties deterministically. With no matching target, `newest` is `undefined`. # Predicates > Complete reference for filtering predicates by data type Predicates are the building blocks for filtering in TypeGraph queries. Each data type has its own set of predicates optimized for that type. ## How Predicates Work Predicates are accessed through property accessors in `whereNode()` and `whereEdge()`: ```typescript .whereNode("p", (p) => p.name.eq("Alice")) // ^accessor ^predicate ``` The accessor provides type-safe access to the field, and returns a predicate builder with methods appropriate for that field's type. Edge fields work the same way: ```typescript .whereEdge("e", (e) => e.role.eq("admin")) ``` New queries can also use the typed expression context passed as the second callback argument. This is the recommended form when a predicate combines fields, arithmetic, parameters, or subqueries: ```typescript import { expr } from "@nicia-ai/typegraph"; .whereNode("p", (_person, e) => expr.gt(e.p.age, expr.literal(18)), ) ``` The first argument preserves the existing accessor API. Both forms compile through the same predicate representation. See [Database Expressions](/queries/expressions) for composition and null semantics. ## Predicate Types | Type | Predicates | Section | |------|------------|---------| | All types | `eq`, `neq`, `in`, `notIn`, `isNull`, `isNotNull` | [Common](#common-predicates) | | String | `contains`, `startsWith`, `endsWith`, `like`, `ilike` | [String](#string) | | Searchable string | `matches` | [Searchable](#searchable) | | Number | `gt`, `gte`, `lt`, `lte`, `between` | [Number](#number) | | Boolean | *(common only)* | [Boolean](#boolean) | | Date | `gt`, `gte`, `lt`, `lte`, `between` | [Date](#date) | | Array | `contains`, `containsAll`, `containsAny`, `isEmpty`, `isNotEmpty`, `lengthEq/Gt/Gte/Lt/Lte` | [Array](#array) | | Object | `get`, `field`, `hasKey`, `hasPath`, `pathEquals`, `pathContains`, `pathIsNull`, `pathIsNotNull` | [Object](#object) | | Embedding | `similarTo` | [Embedding](#embedding) | | Subquery | `exists`, `notExists`, `inSubquery`, `notInSubquery` | [Subqueries](/queries/advanced) | ## Combining Predicates All predicates can be combined using logical operators: ### AND ```typescript p.status.eq("active").and(p.role.eq("admin")) ``` ### OR ```typescript p.role.eq("admin").or(p.role.eq("moderator")) ``` ### NOT ```typescript p.status.eq("deleted").not() ``` ### Complex Combinations ```typescript p.status .eq("active") .and(p.role.eq("admin").or(p.role.eq("moderator"))) ``` Parenthesization is handled automatically. Vector similarity predicates cannot be nested under `OR` or `NOT`. --- ## Common Predicates These predicates are available on **all** field types: | Predicate | Description | SQL | |-----------|-------------|-----| | `eq(value)` | Equals | `= value` | | `neq(value)` | Not equals | `!= value` | | `in(values[])` | Value is in array | `IN (...)` | | `notIn(values[])` | Value is not in array | `NOT IN (...)` | | `isNull()` | Is null/undefined | `IS NULL` | | `isNotNull()` | Is not null | `IS NOT NULL` | `eq` and `neq` accept `param()` references for [prepared queries](/queries/execute#prepared-queries). `in` and `notIn` accept one in place of the **whole** list — `p.id.in(param("ids"))`, bound with `execute({ ids: [...] })` — but not in place of an individual element. Equality and membership values follow the field's schema type. For example, a `number` field accepts numbers (or a parameter), while a `string` field accepts strings. `eq` and `neq` also accept a compatible explicit `fieldRef` when comparing two query fields. Values that enter through a dynamic accessor or another unchecked boundary are validated while the predicate is built and are refused before TypeGraph compiles SQL when their runtime type disagrees with the registered schema. --- ## String String predicates for text matching and pattern searches. ### Equality ```typescript p.name.eq("Alice") // Exact match p.name.neq("Bob") // Not equal ``` ### Substring Match ```typescript p.name.contains("ali") // Case-insensitive substring match ``` ### Prefix/Suffix ```typescript p.name.startsWith("A") // Case-insensitive prefix match p.name.endsWith("ice") // Case-insensitive suffix match ``` ### Pattern Matching ```typescript p.email.like("%@example.com") // SQL LIKE (case-sensitive) — % = any chars, _ = single char p.name.ilike("alice%") // Case-insensitive LIKE ``` ### List Membership ```typescript p.status.in(["active", "pending"]) p.status.notIn(["archived", "deleted"]) ``` ### Null Checks ```typescript p.email.isNull() p.email.isNotNull() ``` ### Reference | Predicate | Accepts | Description | SQL | Case | |-----------|---------|-------------|-----|------| | `eq(value)` | `string \| param()` | Exact match | `=` | sensitive | | `neq(value)` | `string \| param()` | Not equal | `!=` | sensitive | | `contains(str)` | `string \| param()` | Substring match | `ILIKE '%str%'` | insensitive | | `startsWith(str)` | `string \| param()` | Prefix match | `ILIKE 'str%'` | insensitive | | `endsWith(str)` | `string \| param()` | Suffix match | `ILIKE '%str'` | insensitive | | `like(pattern)` | `string \| param()` | SQL LIKE pattern | `LIKE` | sensitive | | `ilike(pattern)` | `string \| param()` | Case-insensitive LIKE | `ILIKE` | insensitive | | `in(values[])` | `string[]` | In array | `IN (...)` | sensitive | | `notIn(values[])` | `string[]` | Not in array | `NOT IN (...)` | sensitive | | `isNull()` | — | Is null | `IS NULL` | — | | `isNotNull()` | — | Is not null | `IS NOT NULL` | — | > **Wildcard escaping:** User input passed to `contains`, `startsWith`, and `endsWith` is > automatically escaped — `%` and `_` characters are treated as literals. Use `like` or `ilike` > when you need wildcard control. --- ## Number Number predicates for numeric comparisons and ranges. ### Equality ```typescript p.age.eq(30) p.age.neq(0) ``` ### Comparisons ```typescript p.salary.gt(50000) // Greater than p.salary.gte(50000) // Greater than or equal p.age.lt(65) // Less than p.age.lte(65) // Less than or equal ``` ### Range ```typescript p.age.between(18, 65) // Inclusive on both bounds ``` ### List Membership ```typescript p.priority.in([1, 2, 3]) p.priority.notIn([0]) ``` ### Null Checks ```typescript p.score.isNull() p.score.isNotNull() ``` ### Reference | Predicate | Accepts | Description | SQL | |-----------|---------|-------------|-----| | `eq(value)` | `number \| param()` | Equals | `=` | | `neq(value)` | `number \| param()` | Not equals | `!=` | | `gt(value)` | `number \| param()` | Greater than | `>` | | `gte(value)` | `number \| param()` | Greater than or equal | `>=` | | `lt(value)` | `number \| param()` | Less than | `<` | | `lte(value)` | `number \| param()` | Less than or equal | `<=` | | `between(lo, hi)` | `number \| param()` | Inclusive range | `BETWEEN lo AND hi` | | `in(values[])` | `number[]` | In array | `IN (...)` | | `notIn(values[])` | `number[]` | Not in array | `NOT IN (...)` | | `isNull()` | — | Is null | `IS NULL` | | `isNotNull()` | — | Is not null | `IS NOT NULL` | --- ## Boolean Boolean fields support only the [common predicates](#common-predicates): ```typescript p.isActive.eq(true) p.isActive.neq(false) p.isVerified.isNull() p.role.in(["admin", "moderator"]) // works on string enums too ``` No additional boolean-specific predicates are provided — `eq(true)` and `eq(false)` cover the typical cases. --- ## Date Date predicates for temporal comparisons. Accepts `Date` objects or ISO 8601 strings. ### Equality ```typescript p.createdAt.eq("2024-01-01") p.createdAt.neq(new Date("2024-01-01")) ``` ### Comparisons ```typescript p.createdAt.gt("2024-01-01") // After p.createdAt.gte("2024-01-01") // On or after p.createdAt.lt(new Date()) // Before now p.createdAt.lte("2024-12-31") // On or before ``` ### Range ```typescript p.createdAt.between("2024-01-01", "2024-12-31") ``` ### List Membership ```typescript p.birthday.in(["2024-01-01", "2024-07-04"]) ``` ### Null Checks ```typescript p.deletedAt.isNull() p.verifiedAt.isNotNull() ``` ### Reference | Predicate | Accepts | Description | SQL | |-----------|---------|-------------|-----| | `eq(value)` | `Date \| string \| param()` | Equals | `=` | | `neq(value)` | `Date \| string \| param()` | Not equals | `!=` | | `gt(value)` | `Date \| string \| param()` | After | `>` | | `gte(value)` | `Date \| string \| param()` | On or after | `>=` | | `lt(value)` | `Date \| string \| param()` | Before | `<` | | `lte(value)` | `Date \| string \| param()` | On or before | `<=` | | `between(lo, hi)` | `Date \| string \| param()` | Inclusive range | `BETWEEN lo AND hi` | | `in(values[])` | `(Date \| string)[]` | In array | `IN (...)` | | `notIn(values[])` | `(Date \| string)[]` | Not in array | `NOT IN (...)` | | `isNull()` | — | Is null | `IS NULL` | | `isNotNull()` | — | Is not null | `IS NOT NULL` | --- ## Array Array predicates for fields that contain arrays (e.g., `tags: z.array(z.string())`). ### Containment ```typescript p.tags.contains("typescript") // Has specific value p.tags.containsAll(["typescript", "nodejs"]) // Has ALL values p.tags.containsAny(["typescript", "rust"]) // Has ANY value ``` Containment predicates (`contains`, `containsAll`, `containsAny`) are only available when the array element type is a scalar — `string`, `number`, `boolean`, or `Date`. They will not type-check for arrays of objects or arrays. ### Empty Checks ```typescript p.tags.isEmpty() // Empty array OR null p.tags.isNotEmpty() // Has at least one element ``` ### Length Predicates ```typescript p.scores.lengthEq(3) // Exactly 3 elements p.scores.lengthGt(0) // More than 0 elements p.scores.lengthGte(3) // 3 or more elements p.scores.lengthLt(10) // Fewer than 10 elements p.scores.lengthLte(5) // 5 or fewer elements ``` ### Reference | Predicate | Accepts | Description | SQL | |-----------|---------|-------------|-----| | `contains(value)` | `T` | Has value | JSON array contains | | `containsAll(values[])` | `T[]` | Has all values | AND of contains | | `containsAny(values[])` | `T[]` | Has any value | OR of contains | | `isEmpty()` | — | Empty or null | `IS NULL OR length = 0` | | `isNotEmpty()` | — | Has elements | `IS NOT NULL AND length > 0` | | `lengthEq(n)` | `number` | Exactly n elements | `json_array_length(col) = n` | | `lengthGt(n)` | `number` | More than n | `json_array_length(col) > n` | | `lengthGte(n)` | `number` | n or more | `json_array_length(col) >= n` | | `lengthLt(n)` | `number` | Fewer than n | `json_array_length(col) < n` | | `lengthLte(n)` | `number` | n or fewer | `json_array_length(col) <= n` | > **Note:** `isEmpty()` matches both empty arrays (`[]`) and null/undefined values. Use `isNull()` > to check specifically for null. --- ## Object Object predicates for JSON/object fields. Supports both fluent chaining with `get()` and [JSON Pointer](https://www.rfc-editor.org/rfc/rfc6901) syntax for deep access. ### Nested Access with `get()` Type-safe chaining through known keys: ```typescript p.metadata.get("theme").eq("dark") p.settings.get("notifications").get("email").eq(true) ``` `get()` returns a typed field builder — if the nested field is a string you get string predicates, if it's a number you get number predicates, and so on. ### Nested Access with `field()` Access nested fields by JSON Pointer path: ```typescript p.config.field("/settings/theme").eq("dark") p.config.field(["settings", "theme"]).eq("dark") // Array form ``` Like `get()`, `field()` returns a typed field builder for the resolved path. Use `field()` when you need to reach deeply nested paths in a single call. ### Key Existence ```typescript p.metadata.hasKey("theme") // Has top-level key ``` ### Path Operations ```typescript p.config.hasPath("/nested/key") // Has nested path p.config.pathEquals("/settings/theme", "dark") // Value at path equals scalar p.config.pathContains("/tags", "featured") // Array at path contains value p.config.pathIsNull("/optional") // Value at path is null p.config.pathIsNotNull("/required") // Value at path is not null ``` ### Reference | Predicate | Accepts | Description | |-----------|---------|-------------| | `get(key)` | `string` (key name) | Access nested field, returns typed field builder | | `field(pointer)` | `string \| string[]` (JSON Pointer) | Access field by path, returns typed field builder | | `hasKey(key)` | `string` | Has top-level key | | `hasPath(pointer)` | `string \| string[]` | Has nested path | | `pathEquals(pointer, value)` | pointer + `string \| number \| boolean \| Date` | Value at path equals scalar | | `pathContains(pointer, value)` | pointer + `string \| number \| boolean \| Date` | Array at path contains value | | `pathIsNull(pointer)` | `string \| string[]` | Value at path is null | | `pathIsNotNull(pointer)` | `string \| string[]` | Value at path is not null | > **JSON Pointer syntax:** Use `/key/nested/value` string form or `["key", "nested", "value"]` > array form. `pathEquals` only works on scalar values (not objects or arrays). `pathContains` > requires the path to point to an array. --- ## Embedding Embedding predicates for vector similarity search on embedding fields. ### similarTo() Find similar vectors using distance metrics: ```typescript p.embedding.similarTo(queryEmbedding, 10) // Top 10 similar (cosine) ``` ### With Options ```typescript p.embedding.similarTo(queryEmbedding, 10, { metric: "cosine", // "cosine" | "l2" | "inner_product" minScore: 0.8, // Minimum similarity threshold }) ``` ### Reference | Predicate | Accepts | Description | |-----------|---------|-------------| | `similarTo(embedding, k)` | `number[], number` | Top k most similar vectors (cosine) | | `similarTo(embedding, k, opts)` | `number[], number, SimilarToOptions` | Top k with custom metric and threshold | ### Distance Metrics | Metric | Description | Range | Default | Best For | |--------|-------------|-------|---------|----------| | `cosine` | Cosine similarity | 0–1 (1 = identical) | Yes | Normalized embeddings, semantic similarity | | `l2` | Euclidean distance | 0–∞ (0 = identical) | | Absolute distances, unnormalized vectors | | `inner_product` | Inner product (PostgreSQL only) | -∞ to ∞ | | Maximum Inner Product Search (MIPS) | ### Example: Semantic Search ```typescript const similar = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 20, { metric: "cosine", minScore: 0.7, }) ) .select((ctx) => ({ id: ctx.d.id, title: ctx.d.title, content: ctx.d.content, })) .execute(); ``` > **Limitations:** Results are automatically ordered by similarity (most similar first). > `similarTo` cannot be nested under `OR` or `NOT`. Requires pgvector (Postgres) or > sqlite-vec (SQLite) to be loaded at runtime. --- ## Searchable Node-level fulltext match predicate, exposed via `n.$fulltext.matches()`. Only available on nodes whose schema has at least one `searchable()` field. See the [Fulltext Search guide](/fulltext-search) for the full story. ### $fulltext.matches() Find the top-k most relevant matches against the combined content of all `searchable()` fields on the node: ```typescript d.$fulltext.matches("climate change", 10) // Top 10 websearch matches ``` Only exposed on nodes with at least one `searchable()` field: ```typescript const Document = defineNode("Document", { schema: z.object({ title: searchable({ language: "english" }), body: searchable({ language: "english" }), authorId: z.string(), // Plain string — indexed content comes from title + body }), }); ``` ### With Options ```typescript d.$fulltext.matches("machine learning", 20, { mode: "websearch", // "websearch" | "phrase" | "plain" | "raw" language: "english", minScore: 0.01, }) ``` ### Reference | Predicate | Accepts | Description | |-----------|---------|-------------| | `$fulltext.matches(query, k)` | `string, number` | Top k websearch matches | | `$fulltext.matches(query, k, opts)` | `string, number, MatchesOptions` | Top k with parse mode and filters | ### Query Modes | Mode | Description | Example | |------|-------------|---------| | `websearch` (default) | Google-style: quoted phrases, `-excluded`, `OR` | `"New York" -Times` | | `phrase` | Treats the whole query as an exact phrase | `climate change` | | `plain` | ANDs all whitespace-separated terms (no special syntax) | `climate change` | | `raw` | Dialect-native syntax passed through unchanged | Postgres tsquery / FTS5 MATCH | ### Example: Multi-Tenant RAG Filter ```typescript const results = await store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext .matches("quarterly earnings", 20) .and(d.tenantId.eq(tenant.id)) .and(d.published.eq(true)) ) .select((ctx) => ctx.d) .execute(); ``` ### Example: Hybrid Search ($fulltext + similarTo) Combine fulltext and vector in one query — TypeGraph fuses the two ranked lists with Reciprocal Rank Fusion at the SQL layer: ```typescript const results = await store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext .matches("renewable energy", 50) .and(d.embedding.similarTo(queryVec, 50)) .and(d.tenantId.eq(tenant)) ) .fuseWith({ k: 60, weights: { vector: 1.0, fulltext: 1.5 } }) .select((ctx) => ctx.d) .limit(10) .execute(); ``` > **Limitations:** Results are ordered by relevance rank (highest first). > At most one `$fulltext.matches()` per query. Cannot be nested under > `OR` or `NOT`. --- ## Parameterized Predicates Use `param(name)` to create a named placeholder for [prepared queries](/queries/execute#prepared-queries). ```typescript import { param } from "@nicia-ai/typegraph"; const prepared = store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq(param("name"))) .select((ctx) => ctx.p) .prepare(); const results = await prepared.execute({ name: "Alice" }); ``` ### Supported Positions | Position | Supported | Example | |----------|-----------|---------| | Scalar comparisons (`eq`, `neq`, `gt`, `gte`, `lt`, `lte`) | Yes | `p.age.gt(param("minAge"))` | | `between` bounds | Yes | `p.age.between(param("lo"), param("hi"))` | | String operations (`contains`, `startsWith`, `endsWith`, `like`, `ilike`) | Yes | `p.name.contains(param("search"))` | | `in` / `notIn` (whole list) | Yes | `p.id.in(param("ids"))` with `execute({ ids: [...] })` | | `in` / `notIn` (one element) | No | Bind the whole list instead | | Array predicates | No | — | | Subquery predicates | No | — | See [Prepared Queries](/queries/execute#prepared-queries) for full usage and performance details. ## Next Steps - [Filter](/queries/filter) — Using predicates in queries - [Subqueries](/queries/advanced) — `exists()`, `notExists()`, `inSubquery()`, `notInSubquery()` - [Overview](/queries/overview) — Query builder categories # Recursive Traversals > Variable-length path traversals with recursive() Graph queries often need to follow edges to an unknown depth: find all ancestors in a hierarchy, all transitive dependencies of a package, or everyone reachable within six degrees of separation. In a relational database, each depth level requires another self-join — and you have to know the depth ahead of time. Recursive traversals solve this by walking edges until a stopping condition is met. When the backend advertises recursive traversal support, TypeGraph compiles `.recursive()` into a SQL `WITH RECURSIVE` CTE. The database engine handles the iteration, so you get the full performance of native recursive SQL without writing it by hand. ## Backend Support The bundled SQLite and PostgreSQL backends support recursive traversal. A custom backend that cannot compute a bounded transitive closure in one round trip must declare: ```typescript const capabilities = { recursiveTraversal: { supported: false, reason: "engine has no WITH RECURSIVE or graph-native equivalent", }, } satisfies Partial; ``` Calling `.recursive()` through that backend throws `ConfigurationError` with `details.code: "RECURSIVE_TRAVERSAL_UNSUPPORTED"`. `details.operation` identifies the refusing query path and `details.reason` contains the backend's declaration. This refusal also applies when the recursive query is nested inside `union()`, `intersect()`, or `except()`. An absent `recursiveTraversal` declaration means supported for backward compatibility with custom backends that already execute recursive SQL. See [Recursive traversal capability](/backend-setup#recursive-traversal-capability) for the complete backend-author contract and the other operations governed by this capability. ## How It Works A recursive traversal starts from a set of source nodes and repeatedly follows edges, accumulating results at each level: ```text Level 0: Alice │ reportsTo Level 1: Bob │ reportsTo Level 2: Carol │ reportsTo Level 3: Dana (CEO) ``` With `.recursive()`, a single query returns Bob, Carol, and Dana — regardless of how deep the chain goes. Without it, you'd need to know there are exactly 3 levels and chain 3 traversals manually. ## Basic Usage Add `.recursive()` between `.traverse()` and `.to()`: ```typescript const allManagers = await store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq("Alice")) .traverse("reportsTo", "e") .recursive() .to("Person", "manager") .select((ctx) => ({ employee: ctx.p.name, manager: ctx.manager.name, })) .execute(); // Returns every manager above Alice, at any depth ``` ## Options Reference ```typescript .recursive(options?) ``` | Option | Type | Default | Description | | ------------- | ----------------------------------------------------------------- | ----------- | ------------------------------------------------ | | `minHops` | `number` | `1` | Minimum traversal depth before including results | | `maxHops` | `number` | `10`* | Maximum traversal depth | | `cyclePolicy` | `"prevent" \| "allow"` | `"prevent"` | How to handle cycles | | `depth` | `boolean \| string` | — | Expose hop count in `select()` context | | `path` | `boolean \| string \| { format: "qualified"; alias?: string }` | — | Expose an ID or qualified path | *When `maxHops` is omitted, an implicit cap of 10 is applied. See [Depth Limits](#depth-limits). ## Controlling Depth ### maxHops Cap the traversal depth: ```typescript const nearbyManagers = await store .query() .from("Person", "p") .traverse("reportsTo", "e") .recursive({ maxHops: 3 }) .to("Person", "manager") .select((ctx) => ({ employee: ctx.p.name, manager: ctx.manager.name, })) .execute(); ``` ### minHops Skip nearby results. With `minHops: 2`, direct connections (1 hop) are excluded: ```typescript const distantConnections = await store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq("Alice")) .traverse("knows", "e") .recursive({ minHops: 2 }) .to("Person", "friend") .select((ctx) => ({ person: ctx.p.name, distantFriend: ctx.friend.name, })) .execute(); ``` ### Combining minHops and maxHops ```typescript // Friends-of-friends: 2–4 hops away .recursive({ minHops: 2, maxHops: 4 }) ``` `minHops` must be ≤ `maxHops` when both are specified. ## Tracking Depth and Path When `depth` or `path` are enabled, they become available as properties on the `select()` context. Pass a string to control the property name. Pass `true` to derive the default name from the target alias: `${targetAlias}_depth` or `${targetAlias}_path`. ### depth Expose the hop count as a number in each result row: ```typescript const orgChart = await store .query() .from("Person", "ceo") .whereNode("ceo", (p) => p.role.eq("CEO")) .traverse("manages", "e") .recursive({ depth: "level" }) .to("Person", "employee") .select((ctx) => ({ ceo: ctx.ceo.name, employee: ctx.employee.name, level: ctx.level, // 1 = direct report, 2 = skip-level, etc. })) .execute(); ``` The string `"level"` passed to `depth` becomes `ctx.level` in the select callback — TypeScript infers this automatically, so `ctx.level` is fully typed. ### path Expose the traversal path as an array of node IDs: ```typescript const pathsToRoot = await store .query() .from("Category", "cat") .whereNode("cat", (c) => c.name.eq("Electronics")) .traverse("parentCategory", "e") .recursive({ path: "trail" }) .to("Category", "ancestor") .select((ctx) => ({ category: ctx.cat.name, ancestor: ctx.ancestor.name, trail: ctx.trail, // Array of node IDs from start to ancestor })) .execute(); ``` ### Using both together ```typescript const networkAnalysis = await store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq("Alice")) .traverse("knows", "e") .recursive({ maxHops: 6, depth: "distance", path: "route", }) .to("Person", "connection") .select((ctx) => ({ person: ctx.p.name, connection: ctx.connection.name, distance: ctx.distance, // number route: ctx.route, // string[] of node IDs })) .execute(); ``` For a path that identifies both nodes and traversed edges, request the qualified format: ```typescript const routes = await store .query() .from("Person", "person") .traverse("knows", "connection") .recursive({ maxHops: 4, path: { alias: "route", format: "qualified" }, }) .to("Person", "friend") .select((ctx) => ctx.route) .execute(); ``` Each route alternates node and edge references, starting and ending with a node: ```typescript [ { type: "node", kind: "Person", id: "alice" }, { type: "edge", kind: "knows", id: "edge-1", direction: "out" }, { type: "node", kind: "Person", id: "bob" }, ]; ``` `direction` records how the edge was followed relative to its stored endpoints: `"out"` follows `from` to `to`, while `"in"` follows `to` to `from`. Qualified paths contain references only; they do not hydrate node or edge properties. Legacy `path: true` and `path: "alias"` continue to return node ID arrays. ## Chaining Fixed and Recursive Traversals Fixed-hop and recursive traversals compose from left to right. Each later stage expands from the completed identities produced by its `from` alias, then rejoins those results to the earlier rows. This preserves upstream multiplicity while avoiding repeated expansion of the same `(kind, id)` source within a stage. ```typescript const routes = await store .query() .from("Person", "root") .traverse("manages", "management") .recursive({ maxHops: 3, depth: "managementDepth" }) .to("Person", "manager") .traverse("worksAt", "employment", { from: "manager" }) .recursive({ maxHops: 2, path: "organizationPath" }) .to("Organization", "organization") .where((expr) => expr.organization.active.eq(true)) .orderBy("organization", "name") .limit(20) .select((ctx) => ({ manager: ctx.manager.name, organization: ctx.organization.name, managementDepth: ctx.managementDepth, organizationPath: ctx.organizationPath, })) .execute(); ``` The `minHops`, `maxHops`, `cyclePolicy`, `stopExpansion`, `path`, and `depth` settings apply to their own recursive stage. An `optionalTraverse()`, including the first stage, retains the earlier row when it finds no match; its target, path, and depth values are `undefined`. A completed `.where()`, final ordering, and final limit apply after every stage. The `from` option may also branch from any earlier materialized node alias. ### Mixing fixed hops with recursion A fixed hop can precede or follow recursion. Fixed-hop edges retain their ordinary property bindings; recursive edges are represented by path references. ```typescript const reports = await store.query() .from("Person", "root") .traverse("manages", "directManagement") .to("Person", "directReport") .traverse("manages", "management") .recursive({ minHops: 0, maxHops: 3, depth: "depth" }) .to("Person", "report") .traverse("worksAt", "employment") .to("Organization", "organization") .select((ctx) => ({ report: ctx.report.name, organization: ctx.organization.name, employment: ctx.employment, depth: ctx.depth, })) .execute(); ``` ### An optional first recursive stage Use `optionalTraverse()` before `.recursive()` to retain roots that have no eligible endpoint. With a positive `minHops`, a root with no matching path returns `undefined` for its target, depth, and path. With `minHops: 0`, an eligible root is a real zero-hop match: depth `0` and a one-node path. Endpoint eligibility includes `stopExpansion()` and its `emitStopNode` setting. Match constraints can remove every endpoint while retaining the optional row. A completed `.where()` comparison against the absent target removes that row unless its predicate explicitly allows absence. A subsequent required traversal from an absent target produces no match; a subsequent optional traversal preserves absence. You can still branch from a present earlier alias using the `from` option. ### Boolean shorthand Pass `true` instead of a string to derive output names from the target alias: ```typescript .recursive({ depth: true, path: true }) .to("Person", "target") // ctx.target_depth and ctx.target_path are available in select() ``` ## Cycle Detection Graphs often contain cycles: `A → B → C → A`. Without protection, a recursive traversal on this graph would loop forever. ### cyclePolicy: "prevent" (default) The default policy tracks visited nodes per path and stops when a node would be visited twice. This is safe for any graph topology: ```typescript // Safe even with circular relationships (A → B → C → A) const allReachable = await store .query() .from("Node", "start") .traverse("linkedTo", "e") .recursive() // cyclePolicy: "prevent" is the default .to("Node", "reachable") .select((ctx) => ctx.reachable.id) .execute(); ``` Under the hood, the compiled SQL maintains a path structure at each recursive step and checks whether the next node has already been visited. On PostgreSQL this uses `ARRAY` operations; on SQLite it uses string-delimited path tracking. ### cyclePolicy: "allow" Skips cycle checking entirely. The traversal relies solely on `maxHops` to terminate. Use this when: - You know your graph is acyclic (trees, DAGs) - You want maximum query performance and accept that nodes may appear multiple times - You're using a strict `maxHops` that prevents runaway recursion ```typescript // Tree structure — no cycles possible const ancestors = await store .query() .from("Category", "cat") .traverse("parentCategory", "e") .recursive({ maxHops: 20, cyclePolicy: "allow" }) .to("Category", "ancestor") .select((ctx) => ctx.ancestor.name) .execute(); ``` :::caution With `cyclePolicy: "allow"` on a cyclic graph, the traversal **will** revisit nodes until it hits `maxHops`. If `maxHops` is not set, the implicit cap of 10 prevents infinite recursion, but you may get many duplicate results. ::: ## Expansion and completed-result filters Predicates placed on the target node or edge apply **at every step** of the recursion — not just the final results. This lets you prune paths early: ```typescript // Only follow "active" edges and land on "active" nodes const activeNetwork = await store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq("Alice")) .traverse("knows", "e") .whereEdge("e", (e) => e.status.eq("active")) .recursive({ maxHops: 5 }) .to("Person", "connection") .whereNode("connection", (c) => c.active.eq(true)) .select((ctx) => ctx.connection.name) .execute(); ``` Source node predicates (on `"p"` above) apply only to the starting set. Edge and target node predicates are included in the recursive CTE, so unreachable branches are pruned at each level rather than filtered after the fact. Use `.where()` to filter completed matches without pruning intermediate nodes. In this example, inactive intermediate nodes can still lead to an active endpoint: ```typescript const activeEndpoints = await store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq("Alice")) .traverse("knows", "e") .recursive({ maxHops: 5 }) .to("Person", "connection") .where((fields) => expr.eq(fields.connection.active, expr.literal(true))) .select((ctx) => ctx.connection.name) .execute(); ``` `whereNode("connection", ...)` is an every-hop constraint. `.where(...)` is an endpoint/result constraint applied after recursive expansion and the minimum-depth check. ## Stop expansion at a boundary `stopExpansion()` prevents a matching endpoint from becoming the next recursive frontier. The stopping node is included by default: ```typescript const managersThroughDirectors = await store .query() .from("Person", "employee") .traverse("reportsTo", "edge") .recursive({ maxHops: 10 }) .to("Person", "manager") .stopExpansion("manager", (manager) => manager.role.eq("Director")) .select((ctx) => ctx.manager.name) .execute(); ``` This emits the matching director but does not follow that director's outgoing `reportsTo` edge. Pass `{ emitStopNode: false }` to omit the director as well: ```typescript .stopExpansion( "manager", (manager) => manager.role.eq("Director"), { emitStopNode: false }, ) ``` A stop predicate can use ordinary fields from the recursive target alias. Subqueries, aggregate predicates, ranked search predicates, and fields from other aliases are refused. Only the matching branch stops; other recursive branches continue. With `minHops: 0`, the source node is also the depth-zero target, so a matching source is emitted or omitted according to `emitStopNode` and its branch does not expand. SQL `NULL` does not stop a branch. ## Duplicate Results When a node is reachable via multiple paths, it appears once per path: ```typescript // Graph: A → B → D, A → C → D (D is reachable via two paths) const results = await store .query() .from("Node", "start") .whereNode("start", (n) => n.name.eq("A")) .traverse("linkedTo", "e") .recursive() .to("Node", "reachable") .select((ctx) => ctx.reachable.name) .execute(); // Returns: ["B", "D", "C", "D"] — D appears twice (once per path) ``` To get unique nodes, deduplicate in your application or use [set operations](/queries/combine). ## Depth Limits Two safety caps prevent runaway recursion: | Constant | Value | When it applies | | ------------------------------ | ----- | ---------------------------------- | | `MAX_RECURSIVE_DEPTH` | 10 | `maxHops` is omitted | | `MAX_EXPLICIT_RECURSIVE_DEPTH` | 1000 | Upper bound for explicit `maxHops` | Graphs with branching factor *B* produce O(*B*^depth) rows before cycle detection can prune them. The default of 10 covers typical neighborhood, shortest-path, and hierarchy queries without risking exponential blowup on dense graphs. Pass `maxHops` to `.recursive({ maxHops: N })` to opt in to deeper traversals when you know the graph structure. ```typescript import { MAX_EXPLICIT_RECURSIVE_DEPTH, MAX_RECURSIVE_DEPTH, } from "@nicia-ai/typegraph"; .recursive() // Implicitly capped at 10 .recursive({ maxHops: 50 }) // Honored (≤ 1000) .recursive({ maxHops: 2000 }) // Refused before SQL compilation ``` :::note[Breaking change in v0.14] The default depth was lowered from 100 to 10. If your traversals relied on the implicit 100-hop cap, add `.recursive({ maxHops: 100 })`. ::: ## Limitations - **Match predicates in a staged query can reference only the alias they constrain.** Cross-alias match predicates are refused before execution. Use completed-row `.where()` for comparisons between aliases; remember that a completed-row comparison can remove an unmatched optional row. - **Recursive queries cannot be aggregated in place yet.** `groupBy()`, aggregate projections, aggregate ordering, and `having()` are refused. Project node columns with `project()`, then use `asRelation()` to aggregate that completed relation. - **Recursive edge properties are not projected.** You can filter them with `whereEdge()`. Fixed-hop edge properties remain selectable when fixed hops and recursion appear in the same query. - **Recursive edges are not materialized in the result.** Qualified paths expose edge references, but selected recursive edge fields are refused. ## Real-World Examples ### Organizational Hierarchy Find all reports (direct and indirect) under a manager: ```typescript const allReports = await store .query() .from("Person", "manager") .whereNode("manager", (p) => p.name.eq("VP Engineering")) .traverse("manages", "e") .recursive({ depth: "level" }) .to("Person", "report") .select((ctx) => ({ manager: ctx.manager.name, report: ctx.report.name, level: ctx.level, department: ctx.report.department, })) .orderBy("level", "asc") .execute(); ``` ### Dependency Graph Find all transitive dependencies of a package: ```typescript const dependencies = await store .query() .from("Package", "pkg") .whereNode("pkg", (p) => p.name.eq("my-app")) .traverse("dependsOn", "e") .recursive({ path: "chain", depth: "depth" }) .to("Package", "dep") .select((ctx) => ({ package: ctx.pkg.name, dependency: ctx.dep.name, version: ctx.dep.version, depth: ctx.depth, chain: ctx.chain, })) .orderBy("depth", "asc") .execute(); ``` ### Social Network — Friends of Friends ```typescript const recommendations = await store .query() .from("Person", "me") .whereNode("me", (p) => p.id.eq(currentUserId)) .traverse("follows", "e") .recursive({ minHops: 2, maxHops: 3 }) .to("Person", "suggestion") .select((ctx) => ({ id: ctx.suggestion.id, name: ctx.suggestion.name, })) .limit(20) .execute(); ``` ### Category Breadcrumbs ```typescript const breadcrumbs = await store .query() .from("Category", "current") .whereNode("current", (c) => c.slug.eq("smartphones")) .traverse("parentCategory", "e") .recursive({ path: "pathIds", depth: "depth" }) .to("Category", "ancestor") .select((ctx) => ({ name: ctx.ancestor.name, slug: ctx.ancestor.slug, depth: ctx.depth, })) .orderBy("depth", "desc") .execute(); // Returns: [{ name: "Root", depth: 3 }, { name: "Electronics", depth: 2 }, { name: "Phones", depth: 1 }] ``` ### Access Control — Permission Inheritance Check if a user has access through a group hierarchy: ```typescript const inheritedPermissions = await store .query() .from("Group", "group") .whereNode("group", (g) => g.name.eq("Engineering")) .traverse("parentGroup", "e") .recursive({ depth: "level", maxHops: 10 }) .to("Group", "ancestor") .select((ctx) => ({ group: ctx.ancestor.name, level: ctx.level, })) .execute(); // Returns: [{ group: "Product", level: 1 }, { group: "Company", level: 2 }] // Alice inherits permissions from Engineering → Product → Company ``` ## Next Steps - [Traverse](/queries/traverse) — Single-hop and multi-hop traversals - [Filter](/queries/filter) — Filter nodes and edges with predicates - [Shape](/queries/shape) — Transform output with `select()` - [Combine](/queries/combine) — Merge results from multiple queries # Composing Relations > Combine, filter, aggregate, and batch explicit SQL results Call `asRelation()` on a database projection to work with its output columns. A relation supports SQL filtering, projection, aggregation, ordering, deduplication, top-N per partition, and set operations. Its callbacks see the projected columns, with their original value types and SQL nullability. ```typescript import { expr } from "@nicia-ai/typegraph"; const names = store.query().from("Person", "person") .project((fields) => ({ name: fields.person.name })) .asRelation(); const uniqueNames = await names.distinct() .orderBy((columns) => columns.name) .execute(); ``` Before `asRelation()`, projection ordering uses the graph alias context. After it, ordering and filtering use the output-column context. Graph traversal requires a graph-node binding; a projected object does not automatically become one. ## Set operations over visible columns `union()`, `unionAll()`, `intersect()`, and `except()` combine explicit SQL projections. Operands must have the same ordered column names, value types, and nullability, and compatible graph, execution-target, and temporal provenance. Create projections in the same field order. ```typescript const employees = store.query().from("Employee", "employee") .project((fields) => ({ name: fields.employee.name })).asRelation(); const contractors = store.query().from("Contractor", "contractor") .project((fields) => ({ name: fields.contractor.name })).asRelation(); const people = await employees.union(contractors) .orderBy((columns) => columns.name) .limit(20) .execute(); ``` `union()` removes duplicate projected rows, even when different graph entities produced those rows. `unionAll()` retains duplicates. Distinct set operations and `distinct()` require portable scalar columns; structured JSON equality differs between database engines. `unionAll()` can retain structured columns because it does not compare them for equality. Legacy `select()` callbacks are JavaScript result transformations and keep their existing set-operation behavior. Use `project()` for set operations whose equality is defined by the visible output. ## Filter and aggregate completed results Derived relations make filtering after aggregation and aggregation of aggregated results explicit: ```typescript const totals = store.query().from("Purchase", "purchase") .groupBy((fields) => [fields.purchase.customerId]) .aggregate((fields) => ({ customerId: fields.purchase.customerId, total: expr.sum(fields.purchase.amount), })) .asRelation(); const largeCustomers = await totals .where((columns) => expr.gt(columns.total, expr.literal(100))) .orderBy((columns) => columns.total, "desc", "last") .execute(); const grandTotal = await totals .aggregate((columns) => ({ total: expr.sum(columns.total) })) .first(); ``` Repeated relation `groupBy()` calls accumulate grouping expressions. When `aggregate()` or `project()` completes that grouping, filters, distinctness, ordering, limits, and offsets apply to the input rows first. Order or limit the returned relation to apply those operations to the grouped results instead. SQL NULL still decodes to `undefined`. Ordering accepts an explicit `"first"` or `"last"` null position; the defaults are NULLS LAST for ascending and NULLS FIRST for descending. `distinct()` compares the whole projection before the relation's limit and offset. It does not select an arbitrary edge or path to represent an entity. `count()` counts the current relation, including distinctness and its range; `exists()` checks whether it has a row. Neither runs `map()`. ## Ordered collections Use `expr.collect()` to return a list of scalar values or explicit flat records per group on a backend declaring [`orderedAggregates: true`](/backend-setup#backend-capabilities). Project the input columns into a relation first, then define collection ordering explicitly: ```typescript const purchases = store.query().from("Purchase", "purchase") .project((fields) => ({ id: fields.purchase.id, customerId: fields.purchase.customerId, amount: fields.purchase.amount, purchasedAt: fields.purchase.purchasedAt, })).asRelation(); const histories = await purchases .groupBy((columns) => [columns.customerId]) .aggregate((columns) => ({ customerId: columns.customerId, amounts: expr.collect(columns.amount, { orderBy: [ { expression: columns.purchasedAt, direction: "asc", nulls: "last" }, { expression: columns.id }, ], filter: expr.gte(columns.amount, expr.literal(10)), }), })) .orderBy((columns) => columns.customerId) .execute(); // One row per customer, with matching amounts in purchase order. ``` Project a record when each parent needs the fields from each matching child together. Record fields must be explicitly named scalar expressions; nested objects, arrays, and raw object expressions are not collection elements: ```typescript const histories = await purchases .groupBy((columns) => [columns.customerId]) .aggregate((columns) => ({ customerId: columns.customerId, purchases: expr.collect({ id: columns.id, amount: columns.amount, purchasedAt: columns.purchasedAt, }, { orderBy: [ { expression: columns.purchasedAt }, { expression: columns.id }, ], }), })) .execute(); // Each purchases value is a readonly array of readonly records. // purchasedAt decodes to Date; nullable fields decode to undefined. ``` Import `CollectOptions` to type reusable options or helper parameters without restating the nonempty ordering tuple. It names the options contract for `expr.collect()`. `distinct` and aggregate-local `limit` are not supported. `orderBy` must contain at least one scalar expression. Each item accepts `direction` (`"asc"` by default) and `nulls` (last for ascending, first for descending). Include a unique tie-breaker when other ordering values can tie. Collection ordering controls elements inside each list; the relation's outer `orderBy()` controls result rows. Source ordering is not an implicit collection order. The optional `filter` is a Boolean database expression in the same scope as the value and ordering expressions. SQL TRUE includes an element; false and SQL NULL exclude it. This follows SQL aggregate filter semantics, including for prepared parameters. Aggregate-local filtering matters with optional traversals. It can exclude a missing child while retaining the parent's group: ```typescript const projects = store.query().from("Project", "project") .optionalTraverse("hasTask", "assignment").to("Task", "task") .project((fields) => ({ project: fields.project.name, taskId: fields.task.id, taskTitle: fields.task.title, priority: fields.task.priority, })).asRelation(); const rows = await projects .groupBy((columns) => [columns.project]) .aggregate((columns) => ({ project: columns.project, tasks: expr.collect(columns.taskTitle, { orderBy: [{ expression: columns.priority }, { expression: columns.taskId }], filter: expr.isNotNull(columns.taskId), }), })) .orderBy((columns) => columns.project) .execute(); // [{ project: "Launch", tasks: ["Fix blocker", "Write announcement"] }, // { project: "Research", tasks: [] }] ``` Putting `expr.isNotNull(columns.taskId)` in the relation's outer `where()` instead removes the childless row before grouping, so `Research` has no result row. Use the collection filter when the parent must remain visible with an empty collection. The same rule applies to records. Replace `columns.taskTitle` in the optional-traversal example with `{ id: columns.taskId, title: columns.taskTitle }`, keeping its `orderBy` and `filter: expr.isNotNull(columns.taskId)` options. This returns `[]` for a parent without a child. Without that filter, an admitted optional-traversal row creates a record even when every projected field is SQL NULL; its fields decode to `undefined`. Scalar collection elements may be strings, numbers, Booleans, or dates. Record fields may use those same scalar types, and their Boolean, Date, and SQL NULL values decode to `boolean`, `Date`, and `undefined` respectively. The result is a readonly array with the projected element types and nullability preserved. Filtering does not change those types. An included SQL NULL operand decodes to `undefined` and remains in the collection; filtering is based only on the `filter` expression. This includes NULL values from missing optional targets when the filter admits them. An empty ungrouped collection aggregate returns one row containing `[]`; an empty grouped relation returns no rows. Source filters, distinctness, and ranges apply before collection aggregation. Duplicates remain unless you deduplicate the input projection explicitly. Collections support prepared and batched relation execution. They are materialized arrays, with no implicit truncation or response-byte limit. Arbitrary object/nested collection elements and aggregate-local limits are outside this API. Structured equality restrictions still apply: collection columns cannot be used as relation ordering keys, with `distinct()`, distinct set operations, grouping keys, or the existing scalar-only paging contract. Use compatible `unionAll()` to retain collection rows without equality. ## Top-N per parent Use `topPerPartition()` to select up to N rows independently for each parent in one SQL query. Partition keys identify the parent; the stage's ordering chooses its winning children. Both callbacks must return nonempty tuples of scalar expressions. Include the parent's kind as well as its ID when IDs can overlap across node kinds. ```typescript const recentPurchases = purchases.topPerPartition({ partitionBy: (columns) => [columns.customerId], orderBy: (columns) => [ { expression: columns.purchasedAt, direction: "desc", nulls: "last" }, { expression: columns.id }, ], limit: 3, }); const rows = await recentPurchases .orderBy((columns) => columns.customerId) .orderBy((columns) => columns.purchasedAt, "desc", "last") .orderBy((columns) => columns.id) .execute(); ``` `limit` must be a positive safe integer. Import `TopPerPartitionOptions` to type reusable options for a projected relation, and `TopPerPartitionOrder` for reusable ordering entries. The stage uses `ROW_NUMBER()`: ties do not expand the limit. Supply a stable final tie-breaker, usually the child's ID, to choose repeatable winners. Use kind and ID for multi-kind children whose IDs can overlap. The API cannot prove that your ordering is unique. Nullable partition keys group together, and ordering accepts the same explicit null positions and defaults as relation ordering. The ranking column is private and never appears in the result. Stage ordering chooses winners; it does not guarantee final result order. Add relation `orderBy()` after the stage to order returned rows. A `where()` before `topPerPartition()` chooses candidates; a `where()` afterward removes winners without selecting replacements. For example, filter purchases by a minimum amount before ranking to retrieve the latest three qualifying purchases, or after ranking to inspect which of the latest three qualify. Source `distinct()`, limits, and offsets apply before ranking. A source limit is global and can remove a parent's candidates entirely. Limits and offsets added after ranking apply globally to the winners. Complete any pending `groupBy()` with `aggregate()` or `project()` before ranking; post-execution JavaScript `map()` results cannot be ranked in SQL. Ranked rows can feed ordered record collections, keeping the per-parent bound in SQL: ```typescript const histories = await recentPurchases .groupBy((columns) => [columns.customerId]) .aggregate((columns) => ({ customerId: columns.customerId, purchases: expr.collect({ id: columns.id, amount: columns.amount, purchasedAt: columns.purchasedAt, }, { orderBy: [ { expression: columns.purchasedAt, direction: "desc", nulls: "last" }, { expression: columns.id }, ], }), })) .orderBy((columns) => columns.customerId) .execute(); ``` For optional traversals, partition by the parent's identity. A childless parent's placeholder row survives ranking. Keep `filter: expr.isNotNull(columns.taskId)` inside `expr.collect()` to turn that placeholder into `[]`, as in the optional-traversal example above. Filtering the placeholder out of the relation would remove the parent. Ranked relations preserve projected types and codecs and support prepared queries and `batchOnce()`. They require [`windowFunctions: true`](/backend-setup#backend-capabilities); unsupported backends throw `UnsupportedBackendCapabilityError` before execution. This bounds returned rows per partition, but the database may still scan and sort all candidates. It does not impose a response-byte budget or add an aggregate-local collection limit. ## Typed prepared composition Reuse parameter expressions in the query and pass their declaration to `prepare()` to infer the binding object's names and value types. The declaration must match every parameter used across all operands and derived stages. Undeclared, missing, extra, and incompatible parameters are refused. ```typescript const parameters = { minimum: expr.param("minimum", "number") }; const prepared = totals .where((columns) => expr.gt(columns.total, parameters.minimum)) .prepare(parameters); const rows = await prepared.execute({ minimum: 100 }); const bound = prepared.bind({ minimum: 200 }); ``` `prepare()` without a declaration preserves the compatibility binding type `Readonly>` and validates bindings at runtime. Parameter types are inferred from the explicit declaration, not recovered automatically from every earlier query callback. ## Entity identities and deterministic pages To deduplicate nodes, project only a node's `kind` and `id`, then call `distinctNodes({ kind: "kind", id: "id" })`. The relation verifies that both columns came from the same graph-node alias. The representative policy is identity-only: adding payload, edge, or path columns is refused, because different matches might disagree on those values. `page({ limit, offset? })` and `stream({ pageSize? })` require a provably unique ordering: use a whole-row `distinct()` relation with scalar columns and order by every output column exactly once. Computed sort expressions, missing tie-breakers, structured columns, and unproven uniqueness are refused. Page limits and stream page sizes must be positive integers. ```typescript const orderedNames = names.distinct().orderBy((columns) => columns.name); const secondPage = await orderedNames.page({ limit: 20, offset: 20 }); for await (const row of orderedNames.stream({ pageSize: 100 })) { console.log(row.name); } ``` Streaming fetches bounded pages; it is not a database cursor. Concurrent writes can change later pages unless the relation runs inside a transaction with an appropriate stable snapshot. Ordinary `limit()` and `offset()` remain available for relations without a proven unique order. ## Batch derived results A bound relation implements the same one-statement read contract as other `batchOnce()` inputs: ```typescript const [highValue, allTotals] = await store.batchOnce(() => [ prepared.bind({ minimum: 200 }), totals.orderBy((columns) => columns.customerId), ] as const); ``` The requests execute in one SQL statement and retain independent result arrays. Normal batch budgets, execution-target checks, and materialized-response tradeoffs apply. For retrieving multiple subgraphs, continue to use the recommended [`batchOnce(read => roots.map(root => read.subgraph(...)))` pattern](/performance/overview/). Derived relations currently refuse execution inside `withCheckedReads()`. Recorded views retain their recorded coordinate for direct relation execution and preparation; recorded relations remain unavailable in `batchOnce()`. These refusals occur before any member executes. # Shape > Transform output with select(), project(), map(), and aggregate() Shape operations transform how results are returned. Use `select()` for the compatibility result mapper, `project()` for an explicit SQL projection, `map()` for a final JavaScript transformation, and `aggregate()` for grouped results. ## select() The `select()` method defines what data to return: ```typescript const results = await store .query() .from("Person", "p") .select((ctx) => ({ name: ctx.p.name, email: ctx.p.email, })) .execute(); ``` ### Parameters ```typescript .select(selectFunction) ``` | Parameter | Type | Description | |-----------|------|-------------| | `selectFunction` | `(ctx) => T` | Function that receives a context and returns the output shape | The context provides typed access to all nodes and edges in the query via their aliases. For new queries that need SQL composition, use `project()` and the typed `expr` helpers described in [Database Expressions](/queries/expressions). A projected query can become an [`asRelation()`](/queries/relations/) for derived filtering, aggregation, distinctness, and set operations. Use `select()` when you specifically need the compatibility JavaScript mapper. ## Selection Patterns ### Full Node Return all properties as an object: ```typescript .select((ctx) => ctx.p) // Returns: { id, kind, name, email, ... } ``` ### Specific Fields Return only the fields you need: ```typescript .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name, email: ctx.p.email, })) ``` :::tip[Performance] Selecting specific fields triggers TypeGraph's **smart select optimization**. Instead of fetching the entire `props` blob, TypeGraph generates SQL that extracts only the requested fields. This can significantly improve performance, especially with well-designed [indexes](/performance/indexes). ::: ### Node Metadata Include system metadata fields: ```typescript .select((ctx) => ({ id: ctx.p.id, kind: ctx.p.kind, // "Person" version: ctx.p.version, // Optimistic concurrency version createdAt: ctx.p.createdAt, updatedAt: ctx.p.updatedAt, })) ``` ### Multiple Nodes Select from multiple nodes in a traversal: ```typescript const results = await store .query() .from("Person", "p") .traverse("worksAt", "e") .to("Company", "c") .select((ctx) => ({ person: ctx.p.name, company: ctx.c.name, role: ctx.e.role, // Edge property })) .execute(); ``` ### Nested Objects Structure output with nested objects: ```typescript .select((ctx) => ({ employee: { id: ctx.p.id, name: ctx.p.name, }, company: { id: ctx.c.id, name: ctx.c.name, }, employment: { role: ctx.e.role, startDate: ctx.e.startDate, }, })) ``` ### Renamed Fields Rename fields in the output: ```typescript .select((ctx) => ({ personName: ctx.p.name, // Renamed from 'name' companyName: ctx.c.name, // Renamed from 'name' jobTitle: ctx.e.role, // Renamed from 'role' })) ``` ## Type Inference TypeScript infers the result type from your selection: ```typescript // TypeScript infers: Array<{ name: string; email: string | undefined }> const results = await store .query() .from("Person", "p") .select((ctx) => ({ name: ctx.p.name, // string (required in schema) email: ctx.p.email, // string | undefined (optional in schema) })) .execute(); // Invalid property access caught at compile time: .select((ctx) => ({ invalid: ctx.p.nonexistent, // TypeScript error! })) ``` ## Optional Traversal Results When using `optionalTraverse()`, accessed nodes and edges may be `undefined`: ```typescript const results = await store .query() .from("Person", "p") .optionalTraverse("worksAt", "e") .to("Company", "c") .select((ctx) => ({ person: ctx.p.name, company: ctx.c?.name, // May be undefined role: ctx.e?.role, // May be undefined })) .execute(); ``` ## aggregate() Use `aggregate()` with aggregate functions for grouped queries: ```typescript import { count, sum, avg, field } from "@nicia-ai/typegraph"; const stats = await store .query() .from("Person", "p") .traverse("worksAt", "e") .to("Company", "c") .groupBy("c", "name") .aggregate({ companyName: field("c", "name"), employeeCount: count("p"), totalSalary: sum("e", "salary"), avgSalary: avg("e", "salary"), }) .execute(); ``` See [Aggregate](/queries/aggregate) for full aggregate documentation. ## Selecting Path Information With recursive traversals, include path and depth: ```typescript const results = await store .query() .from("Category", "cat") .traverse("parentCategory", "e") .recursive({ path: "pathIds", depth: "depth" }) .to("Category", "ancestor") .select((ctx) => ({ category: ctx.cat.name, ancestor: ctx.ancestor.name, path: ctx.pathIds, // Array of node IDs depth: ctx.depth, // Number of hops })) .execute(); ``` ## Temporal Metadata When using [temporal queries](/queries/temporal), access validity information: ```typescript const history = await store .query() .from("Article", "a") .temporal("includeEnded") .select((ctx) => ({ title: ctx.a.title, validFrom: ctx.a.validFrom, // When this version became valid validTo: ctx.a.validTo, // When superseded (undefined if current) version: ctx.a.version, // Version number })) .execute(); ``` ## Return Type `select()` returns an `ExecutableQuery` that provides: - `execute()` - Run the query and get results - `paginate()` - Cursor-based pagination - `stream()` - Stream results for large datasets - `first()` - Get the first result or undefined - `count()` - Count matching results - `exists()` - Check if any results exist - `toAst()` - Get the query AST - `compile()` - Compile to SQL ## Next Steps - [Expressions](/queries/expressions) - SQL projection and typed expression callbacks - [Relations](/queries/relations) - Compose explicit projected results - [Aggregate](/queries/aggregate) - Grouping and aggregate functions - [Order](/queries/order) - Ordering and limiting results - [Execute](/queries/execute) - Running queries and pagination # Source > Starting queries with from() Every query starts with `from()`, which specifies the node kind or explicit list of kinds to query and assigns an alias for referencing it throughout the query. ## Basic Usage ```typescript const results = await store .query() .from("Person", "p") // Start from Person nodes, alias as "p" .select((ctx) => ctx.p) .execute(); ``` ## Parameters ```typescript .from(kind, alias, options?) ``` | Parameter | Type | Description | |-----------|------|-------------| | `kind` | `string` | The node kind to query (must exist in your graph definition) | | `alias` | `string` | A unique identifier for referencing this node in the query | | `options.includeSubClasses` | `boolean` | Include nodes of subclass kinds (default: `false`) | ## Multiple kinds in one query Pass a nonempty list to scan several registered kinds in one SQL query: ```typescript // Person and Company both declare a string name property. const page = await store .query() .from(["Person", "Company"], "entity") .whereNode("entity", (entity) => entity.name.startsWith("A")) .orderBy("entity", "name") .select((ctx) => ctx.entity) .paginate({ first: 20 }); ``` The source scans exactly these kinds. Repeated kinds are normalized, an empty list is rejected, and unknown kinds throw `KindNotFoundError`. List sources do not accept an options bag or expand subclasses. For a reusable list, preserve its nonempty tuple type with `as const`. Predicates and database expressions expose compatible properties shared by every selected kind, plus system fields such as `id` and `kind`. A property present on only one kind, or declared with incompatible types across kinds, cannot be used as a shared query field. Full-node selections retain each kind's properties as a discriminated union; check the returned node's `kind` before reading kind-specific properties. Multi-kind sources use the existing graph scope, visibility, temporal, identity, traversal, aggregation, and prepared-query paths. They also compose with `batchOnce()` like single-kind sources. Use them for a single ordered stream, such as a directory of people and companies. For separate result lists with independent filters and limits, use [one-statement batches](/queries/execute#batch-execution) instead. Node identity across kinds is `(kind, id)`. Cursor pagination and streaming append missing `kind ASC` and `id ASC` keys to make page boundaries deterministic, including when two kinds contain the same ID. For queries that fan out through traversals, also order by the traversed row identities when one source node can produce multiple rows. The input list's order does not order results: specify `orderBy()`. Restart saved multi-kind cursors from older versions if they lack the `kind` identity column. This is one source scan, but it does not guarantee an index-only plan or eliminate a database sort; verify the execution plan for your selected kinds, filters, and ordering. ## Aliases The alias is used throughout the query to reference the node: ```typescript const results = await store .query() .from("Person", "person") .whereNode("person", (p) => p.status.eq("active")) // Reference in filter .orderBy("person", "name", "asc") // Reference in ordering .select((ctx) => ({ name: ctx.person.name, // Reference in selection email: ctx.person.email, })) .execute(); ``` Aliases must be unique within a query. TypeScript enforces this at compile time: ```typescript store .query() .from("Person", "p") .traverse("worksAt", "e") .to("Company", "p") // TypeScript error: alias "p" already in use ``` ## Subclass Expansion If your ontology defines subclass relationships, you can query a parent kind and include all subclasses: ```typescript // Graph definition with subclass relationships: // subClassOf(Podcast, Media) // subClassOf(Article, Media) // subClassOf(Video, Media) // Query only exact Media nodes (default behavior) const exactMedia = await store .query() .from("Media", "m") .select((ctx) => ctx.m) .execute(); // Query Media and all subclasses const allMedia = await store .query() .from("Media", "m", { includeSubClasses: true }) .select((ctx) => ({ kind: ctx.m.kind, // "Media" | "Podcast" | "Article" | "Video" title: ctx.m.title, })) .execute(); ``` When `includeSubClasses: true`: - Results include nodes of the specified kind AND all subclass kinds - The `kind` field in results reflects the actual node kind - All properties common to the parent kind are accessible ## Runtime-declared kinds `from()` requires `kind` to be a compile-time literal in your graph definition. For kinds added at runtime via [graph extensions](/graph-extensions), use `fromDynamic()`: ```typescript // "Paper" was added by store.evolve(extension), so it isn't in the // compile-time graph type. fromDynamic accepts arbitrary string kinds. const recent = await store .query() .fromDynamic("Paper", "p") .whereNode("p", (p) => p.field("year").number().gte(2020)) .select((ctx) => ctx.p) .execute(); ``` The kind is validated against the registry — typos throw `KindNotFoundError` instead of silently producing an empty query. The alias's predicate accessor is `DynamicNodeAccessor`, which exposes schema properties through a `.field(name)` discriminator — `.field("year").number().gte(2020)` for type-narrow predicates, `.field("year").eq(...)` directly for `BaseFieldAccessor` methods. See [Traverse ▸ The `.field()` discriminator](/queries/traverse#the-field-discriminator) for the full surface. `fromDynamic` mixes freely with typed `traverse` / `to` and the dynamic siblings. See [graph extensions ▸ Querying extension kinds](/graph-extensions#querying-extension-kinds) for the full story. Passing Store-issued runtime-kind evidence instead of a string narrows the alias to the graph-extension definition, so ordinary typed property access is available. String inputs retain the discriminator-based dynamic surface. ## Return Type `from()` returns a `QueryBuilder` that provides access to all query methods: - [Filter](/queries/filter) - `whereNode()`, `whereEdge()` - [Traverse](/queries/traverse) - `traverse()`, `optionalTraverse()` - [Shape](/queries/shape) - `select()`, `aggregate()` - [Order](/queries/order) - `orderBy()`, `limit()`, `offset()` - [Aggregate](/queries/aggregate) - `groupBy()`, `groupByNode()` - [Temporal](/queries/temporal) - `temporal()` - [Compose](/queries/compose) - `pipe()` ## Next Steps - [Filter](/queries/filter) - Reduce results with `whereNode()` - [Traverse](/queries/traverse) - Navigate to related nodes - [Shape](/queries/shape) - Define output with `select()` # Traverse > Navigate relationships with traverse() and optionalTraverse() Traversals let you navigate relationships in your graph. Instead of writing complex SQL joins, describe the path you want to follow. ## Single-Hop Traversal Follow one edge from a node to connected nodes: ```typescript const employments = await store .query() .from("Person", "p") .whereNode("p", (p) => p.id.eq("alice-123")) .traverse("worksAt", "e") // Follow worksAt edges .to("Company", "c") // Arrive at Company nodes .select((ctx) => ({ person: ctx.p.name, company: ctx.c.name, role: ctx.e.role, // Edge properties are accessible })) .execute(); ``` ## Parameters ### traverse() ```typescript .traverse(edgeKind, edgeAlias, options?) ``` | Parameter | Type | Description | |-----------|------|-------------| | `edgeKind` | `string` | The edge kind to traverse | | `edgeAlias` | `string` | Unique alias for referencing this edge | | `options.direction` | `"out" \| "in"` | Traversal direction (default: `"out"`) | | `options.expand` | `"none" \| "implying" \| "inverse" \| "all"` | Ontology edge expansion mode (default: `"inverse"`) | | `options.from` | `string` | Fan-out from a different node alias | ### optionalTraverse() ```typescript .optionalTraverse(edgeKind, edgeAlias, options?) ``` Uses the same options as `traverse()`, but returns optional edge/node values in the result context. ### to() ```typescript .to(nodeKind, nodeAlias, options?) ``` | Parameter | Type | Description | |-----------|------|-------------| | `nodeKind` | `string` | The target node kind | | `nodeAlias` | `string` | Unique alias for referencing this node | | `options.includeSubClasses` | `boolean` | Include subclass kinds (default: `false`) | ## Direction By default, traversals follow edges in their defined direction (from → to). Use `direction: "in"` to traverse backwards: ```typescript // Edge definition: worksAt goes from Person → Company // Forward: Find companies where Alice works .from("Person", "p") .traverse("worksAt", "e") // Person → Company .to("Company", "c") // Backward: Find people who work at Acme .from("Company", "c") .whereNode("c", (c) => c.name.eq("Acme")) .traverse("worksAt", "e", { direction: "in" }) // Company ← Person .to("Person", "p") ``` ## Edge Properties Edges can carry properties. Access them through the edge alias: ```typescript const employments = await store .query() .from("Person", "p") .traverse("worksAt", "e") .to("Company", "c") .select((ctx) => ({ person: ctx.p.name, company: ctx.c.name, role: ctx.e.role, // Edge property salary: ctx.e.salary, // Edge property startDate: ctx.e.startDate, // Edge property })) .execute(); ``` ### Edge Object Structure Each edge provides these fields: | Property | Type | Description | |----------|------|-------------| | `id` | `string` | Unique edge identifier | | `kind` | `string` | Edge type name | | `fromId` | `string` | ID of the source node | | `toId` | `string` | ID of the target node | | `meta.createdAt` | `string` | When the edge was created | | `meta.updatedAt` | `string` | When the edge was last updated | | `meta.deletedAt` | `string \| undefined` | Soft delete timestamp | | `meta.validFrom` | `string \| undefined` | Temporal validity start | | `meta.validTo` | `string \| undefined` | Temporal validity end | | *schema props* | varies | Properties defined in edge schema | ### Filtering on Edge Properties Use `whereEdge()` to filter based on edge values: ```typescript const highPaying = await store .query() .from("Person", "p") .traverse("worksAt", "e") .whereEdge("e", (e) => e.salary.gte(100000)) .to("Company", "c") .select((ctx) => ({ person: ctx.p.name, company: ctx.c.name, salary: ctx.e.salary, })) .execute(); ``` ## Multi-Hop Traversals Chain traversals to follow multiple relationships: ```typescript const projectTasks = await store .query() .from("Person", "person") .whereNode("person", (p) => p.name.eq("Alice")) .traverse("worksOn", "e1") .to("Project", "project") .traverse("hasTask", "e2") .to("Task", "task") .select((ctx) => ({ person: ctx.person.name, project: ctx.project.name, task: ctx.task.title, })) .execute(); ``` Each hop starts from the previous node set and arrives at new nodes. ### Mixed Directions Combine forward and backward traversals: ```typescript const teamStructure = await store .query() .from("Person", "p") .traverse("worksAt", "e1") // Forward: Person → Company .to("Company", "c") .traverse("manages", "e2", { direction: "in" }) // Backward: Person ← manages .to("Person", "manager") .select((ctx) => ({ employee: ctx.p.name, company: ctx.c.name, manager: ctx.manager.name, })) .execute(); ``` ## Optional Traversals Use `optionalTraverse()` for LEFT JOIN semantics—include results even when the traversal has no matches: ```typescript const peopleWithOptionalEmployer = await store .query() .from("Person", "p") .optionalTraverse("worksAt", "e") .to("Company", "c") .select((ctx) => ({ person: ctx.p.name, company: ctx.c?.name, // May be undefined if no employer })) .execute(); // Includes all people, even those without a worksAt edge ``` ### Mixing Required and Optional ```typescript const employeesWithOptionalManager = await store .query() .from("Person", "p") .traverse("worksAt", "e1") // Required: must work at a company .to("Company", "c") .optionalTraverse("reportsTo", "e2") // Optional: might not have manager .to("Person", "manager") .select((ctx) => ({ employee: ctx.p.name, company: ctx.c.name, manager: ctx.manager?.name, // undefined for top-level employees })) .execute(); ``` ### Optional Edge Access With optional traversals, the edge may be `undefined`: ```typescript .select((ctx) => ({ person: ctx.p.name, company: ctx.c?.name, // Node may be undefined role: ctx.e?.role, // Edge may be undefined salary: ctx.e?.salary, })) ``` ## Ontology-Aware Traversals If your ontology defines edge implications, expand queries to include implying edges: ```typescript // Ontology: implies(marriedTo, knows), implies(bestFriends, knows) const connections = await store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq("Alice")) .traverse("knows", "e", { expand: "implying" }) .to("Person", "other") .select((ctx) => ctx.other.name) .execute(); // Returns people connected via "knows", "marriedTo", or "bestFriends" ``` `expand: "implying"` is only reachable for endpoint-compatible implications: every node kind an implying edge (e.g. `marriedTo`) allows on a side must be assignable to a kind the implied edge (`knows`) allows on that same side. `implies()` relations that don't satisfy this are rejected with a `ConfigurationError` when the graph is built into a store, so an `expand: "implying"` traversal can never fold in rows whose kind couldn't actually satisfy the traversal's own endpoints. See [Ontology → Edge Relationships](/ontology#edge-relationships) for details. If your ontology defines inverse edge kinds, you can expand traversals to include inverse edges: ```typescript // Ontology: inverseOf(manages, managedBy) const relationships = await store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq("Alice")) .traverse("manages", "e", { expand: "inverse" }) .to("Person", "other") .select((ctx) => ({ name: ctx.other.name, viaEdgeKind: ctx.e.kind, })) .execute(); // Traverses both "manages" and "managedBy" ``` You can combine both options: ```typescript .traverse("knows", "e", { expand: "all" }) ``` :::note[Default expansion mode] The default expansion mode is `"inverse"`, meaning traversals automatically include inverse edge kinds from your ontology. To opt out for a single traversal, pass `expand: "none"`. To change the default for all traversals, set `queryDefaults.traversalExpansion` in `createStore` options. ::: ## Runtime-declared kinds For kinds and edges added at runtime via [graph extensions](/graph-extensions), use the string-keyed siblings `fromDynamic` (covered on [Source](/queries/source#runtime-declared-kinds)), `traverseDynamic`, `optionalTraverseDynamic`, and `toDynamic`: ```typescript const rows = await store .query() .fromDynamic("Paper", "p") .traverseDynamic("authoredBy", "a") .toDynamic("Author", "u") .whereNode("p", (p) => p.field("year").number().gte(2020)) .select((ctx) => ({ paper: ctx.p, author: ctx.u, edge: ctx.a })) .execute(); ``` Each method runtime-validates against the registry — typos throw `KindNotFoundError`, and a `toDynamic` target that isn't a valid endpoint for the current edge / direction throws `EndpointError`. ### The `.field()` discriminator Predicate accessors on dynamic-declared aliases expose schema properties through a `.field(name)` discriminator. `BaseFieldAccessor` methods (`eq`, `isNull`, `in`, `notIn`) work directly. Type-specific predicates sit behind one of: | Discriminator | Returns | | --- | --- | | `.string()` | `StringFieldAccessor` (`gte`, `contains`, `like`, …) | | `.number()` | `NumberFieldAccessor` (`gte`, `between`, …) | | `.date()` | `DateFieldAccessor` | | `.array()` | `ArrayFieldAccessor` | | `.object()` | `ObjectFieldAccessor<...>` | | `.embedding()` | `EmbeddingFieldAccessor` (`similarTo`) | Each discriminator validates against the registered Zod schema at query-build time and throws `TypeError` on mismatch: ```typescript .whereNode("p", (p) => p.field("year").number().gte(2020)) // ✓ .whereNode("p", (p) => p.field("year").string().eq("2020")) // throws TypeError .whereNode("p", (p) => p.field("yera").number().gte(2020)) // throws (unknown property) // BaseFieldAccessor methods don't need a discriminator. .whereNode("p", (p) => p.field("year").isNotNull()) // ✓ ``` The same `.field()` API is available on edge accessors: ```typescript .whereEdge("a", (e) => e.field("order").number().eq(1)) ``` ### Mixed typed and dynamic aliases Typed and dynamic aliases interleave in one query. Each alias's predicate accessor is resolved independently — typed aliases keep their narrow accessors, dynamic aliases get `.field()`: ```typescript const rows = await store .query() .from("Document", "d") // compile-time kind .traverseDynamic("taggedWith", "e") // runtime edge .toDynamic("Tag", "n") // runtime target .whereNode("d", (d) => d.title.eq("the doc")) // typed: direct .whereNode("n", (n) => n.field("label").string().eq("research")) // dynamic .select((ctx) => ({ doc: ctx.d, tag: ctx.n })) .execute(); ``` A typed `traverse("knownEdge", "e")` followed by `toDynamic(target, "n")` keeps `e` typed — `e.role.eq(...)` works without `.field()` because the edge schema is known at compile time. Only aliases declared via `fromDynamic` / `traverseDynamic` / `optionalTraverseDynamic` / `toDynamic` go through the discriminator. ### Optional dynamic traversal `optionalTraverseDynamic` is the LEFT-JOIN sibling. Source nodes without a matching edge still surface, with the edge and target aliases as `undefined`: ```typescript const papersWithOptionalAuthor = await store .query() .fromDynamic("Paper", "p") .optionalTraverseDynamic("authoredBy", "a") .toDynamic("Author", "u") .select((ctx) => ({ paperTitle: ctx.p.title, authorName: ctx.u?.name, // undefined for orphan papers order: ctx.a?.order, })) .execute(); ``` ## Real-World Examples ### Organizational Hierarchy ```typescript const teamMembers = await store .query() .from("Person", "manager") .whereNode("manager", (p) => p.name.eq("VP Engineering")) .traverse("manages", "e") .to("Person", "report") .select((ctx) => ({ manager: ctx.manager.name, report: ctx.report.name, department: ctx.report.department, })) .execute(); ``` ### Social Graph ```typescript const friends = await store .query() .from("Person", "me") .whereNode("me", (p) => p.id.eq(currentUserId)) .traverse("follows", "e") .to("Person", "friend") .select((ctx) => ({ id: ctx.friend.id, name: ctx.friend.name, followedAt: ctx.e.createdAt, })) .orderBy("e", "createdAt", "desc") .limit(50) .execute(); ``` ### E-Commerce ```typescript const orderDetails = await store .query() .from("Order", "o") .whereNode("o", (o) => o.id.eq(orderId)) .traverse("contains", "e") .to("Product", "p") .select((ctx) => ({ product: ctx.p.name, quantity: ctx.e.quantity, unitPrice: ctx.e.unitPrice, })) .execute(); ``` ## Next Steps - [Recursive](/queries/recursive) - Variable-length paths with `recursive()` - [Filter](/queries/filter) - Filter nodes and edges with predicates - [Shape](/queries/shape) - Transform output with `select()` # Common Patterns > Short, focused patterns for common graph problems Recipes are **short, focused patterns** that solve a specific problem in a few code blocks. For complete, end-to-end implementations, see [Examples](/examples/document-management). | Pattern | Use Case | | ------------------------------------------------------ | ---------------------------------------------- | | [RBAC](#role-based-access-control-rbac) | Permission checks through role hierarchies | | [Social Network](#social-network-followers--feeds) | Feeds, followers, friend recommendations | | [Content Versioning](#content-versioning-with-history) | Valid-time history and bitemporal audit trails | | [Tagging System](#tagging-system) | Flexible categorization with tag clouds | | [Tree Navigation](#tree-navigation) | Hierarchical menus, org charts, file systems | | [Weighted Relationships](#weighted-relationships) | Scoring, relevance, confidence levels | | [Soft Deletes](#soft-deletes-with-cascade) | Safe deletion with relationship cleanup | | [Unique Constraints](#enforcing-unique-constraints) | Preventing duplicates | ## Role-Based Access Control (RBAC) TypeGraph's traversal capabilities make it excellent for modeling permission systems, where access can be inherited through roles or groups. ### Schema Definition ```typescript import { z } from "zod"; import { defineNode, defineEdge, defineGraph } from "@nicia-ai/typegraph"; // 1. Define Nodes const User = defineNode("User", { schema: z.object({ username: z.string() }), }); const Role = defineNode("Role", { schema: z.object({ name: z.string() }), }); const Permission = defineNode("Permission", { schema: z.object({ action: z.string(), resource: z.string() }), }); const Resource = defineNode("Resource", { schema: z.object({ type: z.string(), externalId: z.string() }), }); // 2. Define Edges const hasRole = defineEdge("hasRole"); const hasPermission = defineEdge("hasPermission"); const appliesTo = defineEdge("appliesTo"); // 3. Define Graph (endpoints are specified here, not in defineEdge) const rbacGraph = defineGraph({ id: "rbac_system", nodes: { User: { type: User }, Role: { type: Role }, Permission: { type: Permission }, Resource: { type: Resource }, }, edges: { hasRole: { type: hasRole, from: [User], to: [Role] }, hasPermission: { type: hasPermission, from: [Role, User], to: [Permission] }, appliesTo: { type: appliesTo, from: [Permission], to: [Resource] }, }, }); ``` ### Checking Permissions To check if a user has a specific permission, we can query for a path from the User to the Permission, either directly or through a Role. ```typescript async function checkPermission(userId: string, action: string, resourceId: string) { const result = await store .query() .from("User", "u") .whereNode("u", (p) => p.id.eq(userId)) // Traverse optional roles .optionalTraverse("hasRole", "r_edge") .to("Role", "r") // From either User or Role, look for permissions .traverse("hasPermission", "p_edge") .to("Permission", "p") .whereNode("p", (p) => p.action.eq(action)) .execute(); return result.length > 0; } ``` ## Social Network (Followers & Feeds) Modeling social features requires efficient handling of relationships and recursive queries for recommendations. ### Schema Definition ```typescript const User = defineNode("User", { schema: z.object({ handle: z.string() }), }); const Post = defineNode("Post", { schema: z.object({ content: z.string(), timestamp: z.string() }), }); const follows = defineEdge("follows"); const authored = defineEdge("authored"); const socialGraph = defineGraph({ id: "social", nodes: { User: { type: User }, Post: { type: Post }, }, edges: { follows: { type: follows, from: [User], to: [User] }, authored: { type: authored, from: [User], to: [Post] }, }, }); ``` ### Generating a Feed Retrieve posts from users that the current user follows, ordered by time. ```typescript const feed = await store .query() .from("User", "me") .whereNode("me", (u) => u.id.eq(currentUserId)) .traverse("follows", "f") .to("User", "author") .traverse("authored", "p") .to("Post", "post") .select((ctx) => ({ author: ctx.author.handle, content: ctx.post.content, date: ctx.post.timestamp, })) .orderBy("post", "timestamp", "desc") .execute(); ``` ### Friend Recommendations Find "Friends of Friends" that the user doesn't follow yet. ```typescript const recommendations = await store .query() .from("User", "me") .whereNode("me", (u) => u.id.eq(currentUserId)) .traverse("follows", "f1") .to("User", "friend") .traverse("follows", "f2") .to("User", "fof") // Exclude people I already follow (simplified - in practice use EXCEPT or client filtering) .select((ctx) => ({ handle: ctx.fof.handle, })) .limit(10) .execute(); ``` ## Content Versioning with History TypeGraph has built-in support for temporal data. Every node and edge tracks `validFrom` and `validTo` timestamps, allowing you to travel through time without complex schema changes. For audit cases that need TypeGraph's captured state before a correction, enable recorded/system-time capture with `history: true` and use [`store.asOfRecorded(T)`](/queries/temporal#recorded-time-bitemporal). ### Enabling Temporal Mode Ensure your graph definition allows for history. By default, TypeGraph uses `temporalMode: "current"`, which only returns currently valid data. ```typescript const cmsGraph = defineGraph({ id: "cms", nodes: { /* ... */ }, edges: { /* ... */ }, defaults: { // This allows us to query past states temporalMode: "current", // Default, but can be overridden per query }, }); ``` ### Updating Content When you update a node, TypeGraph automatically: 1. Marks the old row as valid until `now()`. 2. Inserts a new row valid from `now()`. ```typescript // 1. Create initial version const article = await store.nodes.Article.create({ title: "Draft 1", content: "Work in progress...", }); // 2. Update it (automatically versions) await store.nodes.Article.update(article.id, { title: "Final Version", content: "Ready to publish!", }); ``` ### Querying Past States You can query the state of the graph as it existed at any point in time using `asOf`. ```typescript // Get the current version (Final Version) const current = await store .query() .from("Article", "a") .whereNode("a", (a) => a.id.eq(article.id)) .select((ctx) => ctx.a) .execute(); // Get the version from 5 minutes ago (Draft 1) const fiveMinutesAgo = new Date(Date.now() - 5 * 60 * 1000).toISOString(); const past = await store .query() .from("Article", "a") .temporal("asOf", fiveMinutesAgo) .whereNode("a", (a) => a.id.eq(article.id)) .select((ctx) => ctx.a) .execute(); ``` ### Audit Logs To see the full history of changes for a specific node, you can use `includeEnded`. ```typescript const history = await store .query() .from("Article", "a") .temporal("includeEnded") // Include historical rows .whereNode("a", (a) => a.id.eq(article.id)) .orderBy("a", "validFrom", "desc") .select((ctx) => ({ title: ctx.a.title, validFrom: ctx.a.validFrom, validTo: ctx.a.validTo, })) .execute(); ``` ### Replaying a Recorded Belief When `history: true` is enabled, capture a recorded anchor after a decision or report, then replay the graph later even if rows were corrected or deleted: ```typescript const checkpoint = await store.recordedNow(); if (checkpoint === undefined) throw new Error("expected recorded history"); const replay = store.asOfRecorded(checkpoint); const original = await replay.nodes.Article.getById(article.id); ``` See [Bitemporal Time Travel](/examples/bitemporal-time-travel) and [Agent Decision Replay](/examples/agent-decision-replay) for complete runnable examples. ## Tagging System A flexible tagging system where items can have multiple tags, and you can query by tag combinations. ### Schema ```typescript const Item = defineNode("Item", { schema: z.object({ title: z.string(), type: z.string() }), }); const Tag = defineNode("Tag", { schema: z.object({ name: z.string(), color: z.string().optional() }), }); const taggedWith = defineEdge("taggedWith"); const graph = defineGraph({ id: "tagging", nodes: { Item, Tag }, edges: { taggedWith: { type: taggedWith, from: [Item], to: [Tag] } }, }); ``` ### Find Items by Tag ```typescript const photoshopItems = await store .query() .from("Tag", "t") .whereNode("t", (t) => t.name.eq("photoshop")) .traverse("taggedWith", "e", { direction: "in" }) .to("Item", "i") .select((ctx) => ctx.i) .execute(); ``` ### Tag Cloud (Count Items per Tag) ```typescript import { count, field } from "@nicia-ai/typegraph"; const tagCounts = await store .query() .from("Item", "i") .traverse("taggedWith", "e") .to("Tag", "t") .groupBy("t", "name") .aggregate({ tag: field("t", "name"), count: count("i"), }) .execute(); // Sort by count descending const tagCloud = tagCounts.toSorted((a, b) => b.count - a.count); ``` ### Items with Multiple Tags (AND) ```typescript // Find items tagged with BOTH "javascript" AND "tutorial" const jsTag = await store .query() .from("Tag", "t") .whereNode("t", (t) => t.name.eq("javascript")) .select((ctx) => ctx.t) .first(); const tutorialTag = await store .query() .from("Tag", "t") .whereNode("t", (t) => t.name.eq("tutorial")) .select((ctx) => ctx.t) .first(); if (!jsTag || !tutorialTag) { return []; // Tags don't exist } const items = await store .query() .from("Item", "i") .traverse("taggedWith", "e1") .to("Tag", "t1") .whereNode("t1", (t) => t.id.eq(jsTag.id)) .traverse("taggedWith", "e2", { direction: "in" }) .to("Item", "i2") .traverse("taggedWith", "e3") .to("Tag", "t2") .whereNode("t2", (t) => t.id.eq(tutorialTag.id)) .select((ctx) => ctx.i) .execute(); ``` ## Tree Navigation Hierarchical structures like menus, org charts, or file systems. ### Schema ```typescript const Category = defineNode("Category", { schema: z.object({ name: z.string(), slug: z.string(), depth: z.number().default(0), }), }); const parentOf = defineEdge("parentOf"); const graph = defineGraph({ id: "categories", nodes: { Category }, edges: { parentOf: { type: parentOf, from: [Category], to: [Category] } }, }); ``` ### Get All Ancestors (Breadcrumb) ```typescript const breadcrumb = await store .query() .from("Category", "c") .whereNode("c", (c) => c.slug.eq("electronics/phones/iphone")) .traverse("parentOf", "e") .recursive({ path: "path" }) .to("Category", "ancestor") .select((ctx) => ({ name: ctx.ancestor.name, slug: ctx.ancestor.slug, })) .execute(); // Returns: [{ name: "Phones", slug: "..." }, { name: "Electronics", slug: "..." }, ...] ``` ### Get All Descendants ```typescript const allChildren = await store .query() .from("Category", "root") .whereNode("root", (c) => c.slug.eq("electronics")) .traverse("parentOf", "e", { direction: "in" }) .recursive({ depth: "level" }) .to("Category", "child") .select((ctx) => ({ name: ctx.child.name, level: ctx.level, })) .orderBy((ctx) => ctx.level, "asc") .execute(); ``` ### Build a Tree Structure ```typescript async function buildTree(rootSlug: string): Promise { const descendants = await store .query() .from("Category", "root") .whereNode("root", (c) => c.slug.eq(rootSlug)) .traverse("parentOf", "e", { direction: "in" }) .recursive({ maxHops: 10 }) .to("Category", "child") .select((ctx) => ({ id: ctx.child.id, name: ctx.child.name, parentId: ctx.e.fromId, })) .execute(); // Build tree in memory const nodeMap = new Map(); for (const d of descendants) { nodeMap.set(d.id, { ...d, children: [] }); } for (const d of descendants) { if (d.parentId && nodeMap.has(d.parentId)) { nodeMap.get(d.parentId)!.children.push(nodeMap.get(d.id)!); } } return nodeMap.get(rootSlug)!; } ``` ## Weighted Relationships Edges with scores for relevance, confidence, or priority. ### Schema ```typescript const Document = defineNode("Document", { schema: z.object({ title: z.string() }), }); const relatedTo = defineEdge("relatedTo", { schema: z.object({ score: z.number().min(0).max(1), type: z.enum(["similar", "cites", "extends"]), }), }); const graph = defineGraph({ id: "documents", nodes: { Document }, edges: { relatedTo: { type: relatedTo, from: [Document], to: [Document] } }, }); ``` ### Find Highly Related Documents ```typescript const related = await store .query() .from("Document", "d") .whereNode("d", (d) => d.id.eq(documentId)) .traverse("relatedTo", "e") .to("Document", "r") .whereEdge("e", (e) => e.score.gte(0.8)) .select((ctx) => ({ title: ctx.r.title, score: ctx.e.score, type: ctx.e.type, })) .orderBy((ctx) => ctx.score, "desc") .execute(); ``` ### Aggregate Relationship Scores ```typescript import { avg, count, field } from "@nicia-ai/typegraph"; const docStats = await store .query() .from("Document", "d") .traverse("relatedTo", "e") .to("Document", "r") .groupByNode("d") .aggregate({ docId: field("d", "id"), title: field("d", "title"), relationCount: count("e"), avgScore: avg("e", "score"), }) .execute(); ``` ## Soft Deletes with Cascade Delete nodes while preserving relationships for undo capability. ### Mark as Deleted ```typescript // TypeGraph uses soft deletes by default await store.nodes.Document.delete(documentId); // The node still exists but has deleted_at set // Queries automatically filter it out ``` ### Restore Deleted Nodes ```typescript // upsertById "un-deletes" soft-deleted nodes await store.nodes.Document.upsertById(documentId, { title: "Restored Document", content: "...", }); ``` ### Find Deleted Nodes ```typescript // Use temporal queries to see deleted nodes const deletedDocs = await store .query() .from("Document", "d") .temporal("includeEnded") .whereNode("d", (d) => d.deletedAt.isNotNull()) .select((ctx) => ({ id: ctx.d.id, title: ctx.d.title, deletedAt: ctx.d.deletedAt, })) .execute(); ``` ### Cascade Delete Pattern ```typescript async function cascadeDelete(documentId: string): Promise { await store.transaction(async (tx) => { // Find all related edges const edges = await tx .query() .from("Document", "d") .whereNode("d", (d) => d.id.eq(documentId)) .traverse("relatedTo", "e") .to("Document", "r") .select((ctx) => ({ edgeId: ctx.e.id })) .execute(); // Delete edges first for (const { edgeId } of edges) { await tx.edges.relatedTo.delete(edgeId); } // Then delete the node await tx.nodes.Document.delete(documentId); }); } ``` ### Cross-Store Transactions (Drizzle + TypeGraph) When your app writes to the **same database** through both a directly-owned Drizzle connection (relational rows) and a TypeGraph `Store`, use `store.withTransaction(sqlTx)` to enlist both layers in **one** transaction. The caller owns the transaction boundary; TypeGraph adopts that exact connection, so a failure rolls back both layers — no stray relational row, no graph node with a dangling foreign reference. :::caution[Do not use the root store inside a caller-owned transaction] Inside `db.transaction(async (sqlTx) => ...)`, do not call a managed write on the root `store` (for example, `store.nodes.Document.create(...)`). On a single-connection PostgreSQL handle, Drizzle starts that root-store write with `BEGIN` and finishes it with `COMMIT`; PostgreSQL treats the nested `BEGIN` as a warning and the `COMMIT` ends the caller's transaction. Use `store.withTransaction(sqlTx)` and make every graph write through the returned transaction context instead. For a history-enabled store, use `store.withRecordedTransaction(sqlTx, async (tx) => ...)`. ::: This is the explicit adapter surface: create the store with `createAdapterStore` or `createAdapterStoreWithSchema`. The portable `createStore` factories intentionally do not expose native handles or caller-owned transaction adoption. How you open the transaction depends on whether your driver is async or synchronous. `withTransaction` itself is driver-agnostic — it adopts whatever connection you hand it. :::caution[One connection — sequence your writes, don't overlap them] A transaction pins one connection, and both layers share it. TypeGraph serializes the statements *its own* collections issue, but your raw Drizzle writes are yours to order. Awaiting each write (as every example below does) is correct; a `Promise.all` that mixes a graph write with a relational write — or runs two relational writes at once — races two queries on the one connection, the overlap PostgreSQL removes in `pg@9`. The same holds for `store.withRecordedTransaction(...)`. ::: **Async drivers** — node-postgres, `neon-serverless` (Pool/WebSocket), libsql. Use Drizzle's `db.transaction(async …)`: ```typescript await db.transaction(async (sqlTx) => { const [connector] = await sqlTx .insert(connectors) .values({ name: "github" }) .returning({ id: connectors.id }); const txStore = store.withTransaction(sqlTx); await txStore.nodes.ArtifactSource.create({ connectorId: connector.id, label: "primary", }); }); // one COMMIT / ROLLBACK across both layers ``` **Synchronous `better-sqlite3`.** better-sqlite3 is a synchronous driver: Drizzle's `db.transaction()` rejects an `async` callback (`Transaction function cannot return a promise`), and the async continuation would then run *outside* the rolled-back transaction. Open the transaction with explicit `BEGIN`/`COMMIT`/`ROLLBACK` on the single connection instead, and pass that connection to `withTransaction`: ```typescript db.run(sql`BEGIN`); try { const [connector] = db .insert(connectors) .values({ name: "github" }) .returning({ id: connectors.id }) .all(); const txStore = store.withTransaction(db); await txStore.nodes.ArtifactSource.create({ connectorId: connector.id, label: "primary", }); db.run(sql`COMMIT`); } catch (error) { db.run(sql`ROLLBACK`); throw error; } ``` This is safe because better-sqlite3 is a single, single-threaded connection: the `await` between statements only yields the microtask queue, and nothing else touches the connection between `BEGIN` and `COMMIT`/`ROLLBACK`. Do not interleave other async work that writes to the same connection inside the `try`. **Cloudflare Durable Objects (`do-sqlite`).** A store backed by `drizzle(ctx.storage)` is auto-detected as `transactionMode: "do-sqlite"` and advertises `capabilities.execution.interactiveTransactions: true`. Drizzle's own `db.transaction()` here is `ctx.storage.transactionSync` and cannot span an `await`; TypeGraph instead delegates to the async storage runner `ctx.storage.transaction(async …)` (surfaced by Drizzle as `db.$client.transaction`), which rolls back SQL writes across `await`. There is no Drizzle transaction handle on Durable Objects — the storage transaction is ambient on the object — so the caller hands `withTransaction` the same `db`: ```typescript await ctx.storage.transaction(async () => { const txStore = store.withTransaction(db); await txStore.nodes.Document.update(documentId, props); await db.insert(documentVersions).values(versionRow); await db.insert(changeEvents).values(eventRow); }); // one storage-transaction COMMIT / ROLLBACK across both layers ``` `store.transaction(async (tx) => …)` works the same way (TypeGraph opens the storage transaction for you). Boot the parent store with `createAdapterStoreWithSchema` at object startup: bootstrap DDL and the durable materialization marker run *outside* any storage transaction (no DDL is ever emitted inside the business transaction), while the schema-version commit uses the `do-sqlite` runner. Cloudflare D1 is **not** `do-sqlite`: `D1Database.batch(...)` is transactional but not an interactive runner, so D1-backed stores stay `transactionMode: "none"` pending separate work. The adopted context exposes the `{ nodes, edges }` surface (same as `store.transaction`) and reuses the parent store's resolved schema — it runs no migration and emits no DDL inside your transaction. Boot the parent store once at startup with `createAdapterStoreWithSchema(graph, backend)` so fulltext-backed writes find their durable materialization marker; an uninitialized store throws `StoreNotInitializedError` instead of migrating mid-transaction. On SQLite, open caller-owned transactions that may perform constrained writes with `BEGIN IMMEDIATE`. TypeGraph probes the writer slot before any decision-driving constraint read; if a `DEFERRED` snapshot was invalidated by a concurrent commit, the operation refuses with `CONSTRAINT_TRANSACTION_NOT_WRITE_FENCED` and the caller must roll back and retry the whole transaction. `withTransaction` throws `ConfigurationError` on backends that cannot provide real rollback (`backend.capabilities.execution.interactiveTransactions === false`: `drizzle-orm/neon-http`, Cloudflare D1, SQLite `transactionMode: "none"`) — it never silently degrades to a non-atomic fallback, because the relational write would still commit. #### Adapter-owned graph transaction: `store.transaction` + `tx.sql` `withTransaction` is for when the **caller** owns the transaction boundary. When **TypeGraph** should own it on an `AdapterStore`, use `store.transaction(async (tx) => …)` and write your own relational tables through `tx.sql` — the adapter-native handle bound to that same transaction. Stores created from the SQLite and PostgreSQL adapter entrypoints infer this handle automatically: ```typescript await store.transaction(async (tx) => { await tx.nodes.Document.update(documentId, props); if (tx.sqlAvailability !== "available") { throw new Error(`Native transaction unavailable: ${tx.sqlAvailability}`); } const sqlTx = tx.sql; await sqlTx.insert(documentVersions).values(versionRow); await sqlTx.insert(changeEvents).values(eventRow); }); // one COMMIT / ROLLBACK across both layers ``` Custom `AdapterBackend` implementations determine the inferred native handle through their generic parameter. Managed `/sqlite/local` and `/postgres/pglite` entrypoints return `Store`, whose transaction context exposes TypeGraph collections, graph reads, and backend ports—but no adapter-native `tx.sql` handle. Use an adapter entrypoint when application tables must share the transaction. On **Postgres / libsql** this is mandatory for correctness — using the outer `db` would write on a *different* connection and silently escape the transaction. On **better-sqlite3** it is the single connection framed by TypeGraph's `BEGIN`/`COMMIT`/`ROLLBACK`; on **Durable Objects** (`do-sqlite`) it is the bound handle (the storage transaction is ambient). To decide at runtime whether `tx.sql` is usable, branch on **`tx.sqlAvailability`**, not on `tx.sql`'s truthiness: ```typescript await store.transaction(async (tx) => { switch (tx.sqlAvailability) { case "available": { const sqlTx = tx.sql; await sqlTx.insert(documentVersions).values(versionRow); break; } case "unavailable": // Raw Store only: no tx.sql, atomicity, or schema-version fence. await writeWithoutAtomicity(); break; case "history": case "revisionTracking": // Raw SQL is disabled here (it would bypass recorded-time capture or the // revision anchor). Use tx.nodes / tx.edges, or write your own tables // through the external handle you pass to store.withRecordedTransaction. break; } }); ``` When `tx.sqlAvailability` is `"unavailable"`, the callback is running without atomicity and there is no native transaction to join. The non-available union arms omit the `sql` property, so even reading `tx.sql` requires first narrowing the discriminant to `"available"`. Suppressed JavaScript or TypeScript access under `history: true` or `revisionTracking: true` still throws `ConfigurationError`; the error carries a branchable `details.code` — see [Recorded-capture guard codes](/errors/#recorded-capture-guard-codes). ## Enforcing Unique Constraints Prevent duplicate nodes or relationships. ### Schema-Level Uniqueness ```typescript const User = defineNode("User", { schema: z.object({ email: z.string().email(), username: z.string(), }), }); const graph = defineGraph({ id: "users", nodes: { User: { type: User, unique: [ { name: "user_email", fields: ["email"], scope: "kind", collation: "caseInsensitive" }, { name: "user_username", fields: ["username"], scope: "kind", collation: "caseSensitive" }, ], }, }, edges: {}, }); ``` ### Use `getOrCreateByConstraint` ```typescript async function createOrUpdateUserByEmail( email: string, username: string, ): Promise<{ user: Node; action: "created" | "found" | "updated" | "resurrected"; }> { return store.nodes.User.getOrCreateByConstraint( "user_email", { email, username }, { ifExists: "update" }, ); } ``` ### Use `getOrCreateByEndpoints` for Edge Deduplication ```typescript async function followUser(followerId: string, followeeId: string): Promise { const follower = await store.nodes.User.getById(followerId); const followee = await store.nodes.User.getById(followeeId); if (!follower || !followee) { throw new Error("User not found"); } await store.edges.follows.getOrCreateByEndpoints( follower, followee, {}, { ifExists: "return" }, ); } ``` ## Next Steps For complete, end-to-end implementations, see the [Examples](/examples/document-management) section: - [Document Management](/examples/document-management) - CMS with semantic search - [Product Catalog](/examples/product-catalog) - Categories, variants, inventory - [Workflow Engine](/examples/workflow-engine) - State machines with approvals - [Audit Trail](/examples/audit-trail) - Complete change tracking - [Multi-Tenant SaaS](/examples/multi-tenant) - Tenant isolation patterns # Evolving Schemas in Production > Step-by-step guide for safely evolving your graph schema across deployments Your graph schema will change as your application grows. This guide covers how to make those changes safely — from adding a field to renaming a node type. For API reference, see [Schema Migrations](/schema-management). For evolving the kind set itself **at runtime** (agent-induced kinds, plugin-supplied kinds, multi-tenant kind sets), see [Graph Extensions](/graph-extensions). ## How Schema Evolution Works When you call `createStoreWithSchema()`, TypeGraph: 1. Serializes your current graph definition 2. Compares it against the stored schema (by hash, then by diff) 3. **Safe changes** — auto-migrates and bumps the version 4. **Breaking changes** — throws `MigrationError` (or returns `status: "breaking"`) The key insight: TypeGraph manages **schema metadata**, not data migration. When you add an optional field, TypeGraph records that the schema now includes it. It does not alter existing rows — Zod defaults handle that at read time. ## Safe Changes These changes are backwards compatible and auto-migrate without intervention: - Adding new node types - Adding new edge types - Adding optional properties (with defaults) - Adding ontology relations - Changing per-kind annotations (UI hints, audit policy, etc.) - Changing graph-scoped annotations (display metadata, capabilities, etc.) ### Adding an Optional Property ```typescript // Version 1 const Person = defineNode("Person", { schema: z.object({ name: z.string(), }), }); // Version 2 — safe, auto-migrates const Person = defineNode("Person", { schema: z.object({ name: z.string(), email: z.string().optional(), }), }); ``` On startup, `createStoreWithSchema()` returns `status: "migrated"`. Existing Person nodes return `email: undefined` — no data transformation needed. ### Adding a Node Type with Edges ```typescript // Version 2 — add Company and worksAt in one deploy const Company = defineNode("Company", { schema: z.object({ name: z.string() }), }); const worksAt = defineEdge("worksAt", { schema: z.object({ role: z.string() }), }); const graph = defineGraph({ id: "my_app", nodes: { Person: { type: Person }, Company: { type: Company }, }, edges: { worksAt: { type: worksAt, from: [Person], to: [Company] }, }, }); ``` This is a single safe migration. New node and edge types don't affect existing data. ### Changing Annotations The `annotations` field on `defineNode` and `defineEdge` is part of the canonical schema, so any change bumps the schema version. Changes are classified as `safe` — no data migration needed, only the schema document is updated. ```typescript // Version 1 const Incident = defineNode("Incident", { schema: z.object({ title: z.string() }), annotations: { ui: { titleField: "title", icon: "alert-triangle" }, }, }); // Version 2 — swap the icon, add audit policy const Incident = defineNode("Incident", { schema: z.object({ title: z.string() }), annotations: { ui: { titleField: "title", icon: "circle-alert" }, audit: { pii: false, retentionDays: 365 }, }, }); ``` `getSchemaChanges()` reports each annotations-only change per kind: ```typescript import { getSchemaChanges } from "@nicia-ai/typegraph/schema"; const diff = await getSchemaChanges(backend, graph); for (const change of diff?.nodes ?? []) { if (change.details.includes("Annotations")) { console.log(`${change.kind}: annotations changed (${change.severity})`); // → "Incident: annotations changed (safe)" } } ``` The hash is computed with stable sorted-key order at every depth, so re-formatting the annotations object — or swapping sibling key order — does not bump the version. Only structural or value changes do. A few things worth knowing: - Graphs that never set `annotations` produce identical canonical-form hashes to graphs from before this field existed. Adoption requires no migration. - The canonical form omits empty / default annotations, so absent, explicit `undefined`, and explicit `{}` all hash identically — no migration is triggered just by writing `annotations: {}`. - Annotations values must be JSON-serializable (`bigint`, `function`, `Date`, and other class instances are rejected at definition time). See the [schemas-stores reference](/schemas-stores#per-kind-annotations) for the full annotations contract. ### Rolling out graph-scoped annotations Graph-scoped annotations use a top-level `SerializedSchema` field. During a mixed-version rollout, an older schema writer can otherwise recommit a document without a field it does not understand. Use this two-step deployment invariant: 1. Upgrade **every process that can write schema versions** to TypeGraph 0.54 or newer, without adding graph annotations yet. 2. After no older schema writer remains, enable `defineGraph({ annotations })` or `defineGraphExtension({ annotations })` and commit the safe schema change. Readers may be upgraded independently, but the writer floor must be complete before annotations are enabled. TypeGraph 0.54+ preserves unknown top-level schema fields across parse-and-recommit cycles, so later additive metadata slices follow the same rollout rule. ## Breaking Changes These require explicit handling: - Removing node or edge types - Removing properties - Adding required properties (no default) - Renaming types or properties TypeGraph will throw `MigrationError` by default. You have two options: fix the schema to be backwards compatible, or use the expand-contract pattern. ## The Expand-Contract Pattern For breaking changes, use a multi-deploy strategy. This is the same pattern used in relational database migrations — deploy in phases so there's never a moment where running code is incompatible with the schema. ### Renaming a Property Rename `name` to `fullName` on Person in three deploys: #### Deploy 1 — Expand: add the new property ```typescript const Person = defineNode("Person", { schema: z.object({ name: z.string(), fullName: z.string().optional(), // New property, optional for now }), }); ``` Safe migration. Then backfill existing data: ```typescript const [store] = await createStoreWithSchema(graph, backend); const people = await store.query(Person).execute(); for (const person of people) { if (!person.properties.fullName) { await store.nodes.Person.update(person.id, { fullName: person.properties.name, }); } } ``` #### Deploy 2 — Switch: use the new property everywhere Update all application code to read/write `fullName` instead of `name`. Both properties still exist, so this deploy is safe. #### Deploy 3 — Contract: remove the old property ```typescript const Person = defineNode("Person", { schema: z.object({ fullName: z.string(), }), }); ``` This is a breaking change (removing `name`). Use `migrateSchema()` to force it: ```typescript import { getSchemaChanges, migrateSchema } from "@nicia-ai/typegraph/schema"; const [store, result] = await createStoreWithSchema(graph, backend, { throwOnBreaking: false, }); if (result.status === "breaking") { // We've already backfilled — safe to force migrate const activeSchema = await backend.getActiveSchema(graph.id); await migrateSchema(backend, graph, activeSchema!.version); } ``` Two things `migrateSchema()` will not let you do by accident: - **Drop a kind that still holds rows.** The commit is refused with a `MigrationError` whose `details.reason` is `"kind-removal"`. Committing would make those rows unreachable, and the next `materializeRemovals()` would delete them — it re-derives removals by walking schema history, so the drop is not reversible by putting the kind back. Export or delete the rows first (see [Removing a Node Type](#removing-a-node-type)), or pass `{ discardDroppedKindRows: true }` if losing them is the intent. Dropping an *empty* kind needs no flag. - **Erase kinds added at runtime.** `migrateSchema()` folds the persisted graph extension into the graph you hand it, the same way `createStoreWithSchema()` does, so passing your compile-time graph never drops a kind that `evolve()` committed. To remove one of those deliberately, use `removeKinds()` — it queues the cleanup rows that make the removal reconcilable. ### Removing a Node Type #### Deploy 1 — Stop creating new instances Update application code to stop creating the deprecated node type. Existing data remains. #### Deploy 2 — Clean up references Delete edges that reference the deprecated node type, then delete the nodes themselves: ```typescript // Delete all edges connected to deprecated nodes const deprecated = await store.query(OldNode).execute(); for (const node of deprecated) { await store.nodes.OldNode.delete(node.id); } ``` #### Deploy 3 — Remove from schema Remove the node type from `defineGraph()` and force migrate. Deploy 2 is what makes this step legal: `migrateSchema()` refuses to drop a kind that still holds rows, so if any remain you will get a `MigrationError` with `details.reason === "kind-removal"` naming the kind and its row count rather than silent data loss. ### Changing a Property Type Change `age` from `z.string()` to `z.number()`: #### Deploy 1 — Add the new property ```typescript const Person = defineNode("Person", { schema: z.object({ age: z.string(), ageNumeric: z.number().optional(), }), }); ``` #### Deploy 2 — Backfill and switch ```typescript const people = await store.query(Person).execute(); for (const person of people) { if (person.properties.ageNumeric === undefined) { await store.nodes.Person.update(person.id, { ageNumeric: parseInt(person.properties.age, 10), }); } } ``` #### Deploy 3 — Contract Remove `age`, rename `ageNumeric` to `age` with the new type, and force migrate. ### Changing an Embedding Dimension Switching embedding models usually changes the vector dimension (e.g. `embedding(1536)` → `embedding(3072)`). The stored vectors are invalid under the new dimension — they must be recomputed, not converted — so this is handled out-of-band from the schema diff. Update the field's `embedding(N)` in the schema, then call `store.reembedVectorField()`. It drops and recreates the field's per-`(graphId, kind, field)` `tg_vec_*` storage at the new dimension and, when you pass an `embed` callback, pages the kind's nodes and re-embeds them: ```typescript const result = await store.reembedVectorField("Document", "embedding", { embed: async (nodes) => { const texts = nodes.map((node) => node.content); // schema fields are top-level const vectors = await batchEmbed(texts); // your new model return new Map(nodes.map((node, index) => [node.id, vectors[index]])); }, }); // result.recreated === true, result.reembedded === ``` Without an `embed` callback, the storage is recreated empty and you re-embed via normal `update()` writes. Until a field is re-embedded at the new dimension, a stray write at the **old** dimension throws `EmbeddingDimensionChangedError`. ## Pre-Deploy Schema Checks Use `getSchemaChanges()` in CI to catch breaking changes before they reach production. ### CI/CD Script ```typescript import { getSchemaChanges } from "@nicia-ai/typegraph/schema"; async function checkSchema(backend: GraphBackend, graph: GraphDef) { const diff = await getSchemaChanges(backend, graph); if (!diff) { console.log("No existing schema — first deploy"); return; } if (!diff.hasChanges) { console.log("Schema unchanged"); return; } console.log("Schema changes detected:"); console.log(diff.summary); for (const change of [...diff.nodes, ...diff.edges]) { const icon = change.severity === "safe" ? "[safe]" : change.severity === "warning" ? "[warn]" : "[BREAKING]"; console.log(` ${icon} ${change.details}`); } if (diff.hasBreakingChanges) { console.error("Breaking changes require migration before deploy."); process.exit(1); } } ``` ### Staging Validation Before deploying to production, run against a staging database that mirrors production schema state: ```typescript const [store, result] = await createStoreWithSchema(graph, stagingBackend); switch (result.status) { case "initialized": console.log("Staging DB was empty — initialized"); break; case "migrated": console.log( `Auto-migrated v${result.fromVersion} → v${result.toVersion}`, ); console.log("Changes:", result.diff.summary); break; case "breaking": console.error("Would break in production. Fix before deploying."); process.exit(1); break; } ``` ## Testing Schema Changes ### Unit Testing Migrations Test that your migration code handles existing data correctly: ```typescript import { createStoreWithSchema, defineGraph, defineNode } from "@nicia-ai/typegraph"; import { createTestBackend } from "./test-utils"; it("migrates name to fullName", async () => { const backend = createTestBackend(); // Set up v1 with data const graphV1 = defineGraph({ id: "test", nodes: { Person: { type: PersonV1 } }, edges: {}, }); const [storeV1] = await createStoreWithSchema(graphV1, backend); await storeV1.nodes.Person.create({ name: "Alice" }); // Migrate to v2 (expand phase) const graphV2 = defineGraph({ id: "test", nodes: { Person: { type: PersonV2WithBothFields } }, edges: {}, }); const [storeV2, result] = await createStoreWithSchema(graphV2, backend); expect(result.status).toBe("migrated"); // Run backfill const people = await storeV2.query(PersonV2WithBothFields).execute(); for (const person of people) { await storeV2.nodes.Person.update(person.id, { fullName: person.properties.name, }); } // Verify const updated = await storeV2.query(PersonV2WithBothFields).execute(); expect(updated[0].properties.fullName).toBe("Alice"); }); ``` ### Previewing Changes Without Applying Use `getSchemaChanges()` to see what would change without modifying the database: ```typescript import { getSchemaChanges } from "@nicia-ai/typegraph/schema"; const diff = await getSchemaChanges(backend, newGraph); if (diff?.hasChanges) { console.log("Pending changes:", diff.summary); console.log("Breaking:", diff.hasBreakingChanges); for (const change of diff.nodes) { console.log(` ${change.severity}: ${change.details}`); } } ``` ## Version History TypeGraph preserves all schema versions in the `typegraph_schema_versions` table. Only one version is active at a time. ```text typegraph_schema_versions ├── version 1 (initial) ← inactive ├── version 2 (added email) ← inactive ├── version 3 (added Company) ← active ``` Access version history through the backend: ```typescript // Get a specific version const v1 = await backend.getSchemaVersion("my_app", 1); console.log("V1 created at:", v1?.created_at); // Get the active version const active = await backend.getActiveSchema("my_app"); console.log("Current version:", active?.version); ``` ## Summary: Change Classification | Change | Classification | Auto-Migrated? | | ------------------------------ | -------------- | -------------- | | Add node type | Safe | Yes | | Add edge type | Safe | Yes | | Add optional property | Safe | Yes | | Add ontology relation | Safe | Yes | | Change kind annotations | Safe | Yes | | Add required property | Breaking | No | | Remove property | Breaking | No | | Remove node/edge type | Breaking | No | | Rename node/edge type | Breaking | No | | Change property type | Breaking | No | | Change onDelete behavior | Warning | Yes | | Change unique constraints | Warning | Yes | | Change edge cardinality | Warning | Yes | | Change edge endpoint kinds | Warning | Yes | | Remove allowed source-dependent endpoint pairs | Breaking | No | ## Rollback If a deployment goes wrong, you can switch back to a previous schema version. Version history is always preserved — `rollbackSchema()` simply changes which version is active. ```typescript import { rollbackSchema } from "@nicia-ai/typegraph/schema"; // Roll back to version 2 await rollbackSchema(backend, "my_app", 2); ``` This does not delete newer versions. You can migrate forward again later. ## Migration Hooks Use `onBeforeMigrate` and `onAfterMigrate` for observability — logging, metrics, and alerts during schema migrations: ```typescript const [store, result] = await createStoreWithSchema(graph, backend, { onBeforeMigrate: (context) => { console.log(`Migrating ${context.graphId} v${context.fromVersion} → v${context.toVersion}`); console.log("Changes:", context.diff.summary); }, onAfterMigrate: (context) => { console.log(`Migration complete: v${context.toVersion}`); metrics.increment("schema_migrations_total"); }, }); ``` For data transformations (backfill scripts), run them explicitly after store creation rather than inside hooks. This gives you control over retries and error handling: ```typescript const [store, result] = await createStoreWithSchema(graph, backend); if (result.status === "migrated" && result.toVersion === 3) { // Backfill fullName from name for the expand phase const people = await store.query(Person).execute(); for (const person of people) { if (!person.properties.fullName) { await store.nodes.Person.update(person.id, { fullName: person.properties.name, }); } } } ``` ## Reclaiming Removed Embedding Storage Embeddings live in per-`(graphId, kind, field)` tables (`tg_vec_*`), provisioned by the privileged migrator (`createStoreWithSchema`, or `evolve()` for a runtime-added field). When you remove an `embedding()` field from a **surviving** kind, the schema change commits fast but the field's now-orphaned vector table remains until you reconcile it. `store.materializeRemovals()` drops it — and clears its durable contribution marker so a later re-add re-provisions cleanly (this is the same pass that cleans up storage for fully removed kinds): ```typescript const result = await store.materializeRemovals(); for (const reclaimed of result.reclaimedVectorFields) { // → { kind: "Document", fieldPath: "embedding", status: "reclaimed" } console.log(`Dropped vector table for ${reclaimed.kind}.${reclaimed.fieldPath}`); } ``` The pass is idempotent and derived from immutable schema history, so re-running it lists the same removed fields and the underlying `DROP ... IF EXISTS` is a no-op on subsequent calls. ## Current Limitations - **No automatic data transformation.** TypeGraph tracks schema metadata changes but does not transform existing rows. Use backfill scripts (or `onAfterMigrate` hooks) for data migration. - **No rename detection.** Renaming a property looks like a removal + addition. Use the expand-contract pattern instead. - **Schema-level only.** Migrations operate on the graph definition, not on underlying database tables. TypeGraph's storage tables are schema-agnostic (nodes and edges are stored as JSON properties), so "schema migration" means updating the schema document that TypeGraph tracks, not running `ALTER TABLE`. # Schema Migrations > Schema versioning, migration, and lifecycle management For a practical guide on evolving schemas across deployments, see [Evolving Schemas in Production](/schema-evolution). ## When Do You Need Schema Management? As your application evolves, your graph schema changes: - **Adding features**: New node types, new properties, new relationships - **Refactoring**: Renaming types, changing property formats - **Deploying safely**: Ensuring schema changes don't break running applications Without schema management, you'd face: - No way to know if the database matches your code - Silent failures when property names change - Manual migration scripts for every deployment TypeGraph's schema management: 1. **Stores the schema in the database** alongside your data 2. **Detects changes** between your code and the stored schema 3. **Auto-migrates safe changes** (adding types, optional properties) 4. **Blocks breaking changes** until you handle them explicitly ## How It Works TypeGraph stores your graph schema in the database, enabling version tracking, safe migrations, and runtime introspection. When you create a store with `createStoreWithSchema()`, TypeGraph: 1. Creates the base tables if the database is fresh (auto-bootstrap) 2. Serializes your graph definition to JSON 3. Compares it with the stored schema (if any) 4. Returns the result so you can act on it ## Schema Lifecycle When you create a store, TypeGraph can automatically manage schema versions: ```typescript import { createStoreWithSchema } from "@nicia-ai/typegraph"; const [store, result] = await createStoreWithSchema(graph, backend); switch (result.status) { case "initialized": console.log(`Schema initialized at version ${result.version}`); break; case "unchanged": console.log(`Schema unchanged at version ${result.version}`); break; case "migrated": console.log(`Migrated from v${result.fromVersion} to v${result.toVersion}`); break; case "pending": console.log(`Safe changes pending at version ${result.version}`); break; case "breaking": console.log("Breaking changes detected:", result.actions); break; } ``` ## Basic vs Managed vs Verified Store TypeGraph provides three ways to create a store, each suited to a different deployment role: ### Basic Store (No Schema Management) Use `createStore()` when you manage schema versions yourself: ```typescript import { createStore } from "@nicia-ai/typegraph"; const store = createStore(graph, backend); // No schema versioning or write fence - you handle migrations manually ``` Because a basic Store has no committed schema-version metadata, its writes do not participate in the schema-version fence. Direct backend writes have the same raw semantics. Use this mode only when the application accepts responsibility for quiescing writers around schema changes. :::caution[Fulltext requires the managed store] `createStore()` is attach-only. If the graph has `searchable()` fields, use `createStoreWithSchema()` (below) at boot — it durably materializes the fulltext storage. Bare `createStore()` throws `StoreNotInitializedError` on the first fulltext operation. ::: ### Managed Store (Automatic Schema Management) Use `createStoreWithSchema()` for automatic version tracking: ```typescript import { createStoreWithSchema } from "@nicia-ai/typegraph"; const [store, result] = await createStoreWithSchema(graph, backend, { autoMigrate: true, // Auto-apply safe changes (default: true) throwOnBreaking: true, // Throw on breaking changes (default: true) onBeforeMigrate: (context) => { console.log(`Migrating ${context.graphId} from v${context.fromVersion} to v${context.toVersion}`); }, onAfterMigrate: (context) => { console.log(`Migration complete: v${context.toVersion}`); }, }); ``` ### Verified Store (Zero-DDL Attach With Verification Gate) Use `createVerifiedStore()` at runtime when the application runs under a least-privilege, DML-only database role and a separate privileged step has already advanced the schema. It is the runtime counterpart of `createStoreWithSchema()`: a synchronous-semantics attach that **issues no DDL** and fails fast if the database is not at the same schema version as the code graph. ```typescript import { createVerifiedStore } from "@nicia-ai/typegraph"; // Runtime — least-privilege, DML-only role. Zero DDL. const [store, result] = await createVerifiedStore(graph, backend); // result.status === "unchanged" on success. ``` It throws: - `BaseSchemaMigrationError` if deployment-wide base storage is missing, stale, or newer than the running library. Its details report `installedVersion`, `requiredVersion`, and `reason`. - `ConfigurationError` if no schema has been initialized (run the privileged migration step first). - `MigrationError` if the persisted schema is behind the code graph by **any** pending change (safe or breaking) — the least-privilege runtime cannot migrate. - `StoreNotInitializedError` if the schema is current but the runtime-contribution markers (e.g. fulltext) are missing/stale. The attach itself can succeed on a non-transactional or custom backend. On a backend whose `capabilities.execution.unitOfWork` is `"batch"` (Cloudflare D1, Neon HTTP), a fused write commonly succeeds — see [The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for which writes fuse — and a write that cannot fuse throws `ConfigurationError` with `details.code === "SCHEMA_WRITE_FENCE_UNSUPPORTED"`, or, for a proven need such as an interactive callback or a schema commit, a typed error naming `BATCH_WRITE_UNSUPPORTED` under `details.batchRefusal`. On any other backend that provides neither an interactive transaction nor the schema-write fence, every managed write throws `ConfigurationError` with `details.code === "SCHEMA_WRITE_FENCE_UNSUPPORTED"`. Reads remain available. If you only need the check without building a Store (e.g. a readiness probe), call `assertSchemaCurrent(backend, graph)` directly — it returns the same `SchemaValidationResult` or throws the same errors. :::note[Database privileges] Only `createStoreWithSchema()` runs DDL. `createStore()` is a synchronous zero-I/O attach; `createVerifiedStore()` is a SELECT-only attach (zero DDL — reads the base-schema marker, active graph schema, and contribution markers, nothing else). Graph-template registration and instantiation are also DML-only; instantiation copies the source graph's graph-local activation markers while deployment-scoped physical attestations remain shared by the database. A target can therefore be reopened by `createVerifiedStore()` from a later serverless isolate. To run the application under a least-privilege, DML-only role, do the privileged migration step once with `createStoreWithSchema(graph, adminBackend)` to adopt and stamp the current base schema before using the template APIs at runtime. See [Database roles & least privilege](/backend-setup#database-roles--least-privilege) for the canonical breakdown. ::: ### Which Stores are schema-managed? A Store is schema-managed when it carries committed schema metadata: `store.introspect().schemaVersion !== undefined`. The following paths create or preserve that state: - `createStoreWithSchema()` and `createAdapterStoreWithSchema()` - `createVerifiedStore()` and `createVerifiedAdapterStore()` - `createAdapterStore(..., { reconciled })` with a cached reconciled snapshot - Stores returned by `evolve()` and Stores rebound from an already-managed Store Managed writes acquire a transaction-scoped fence and revalidate that version before changing graph data. On the official SQLite and PostgreSQL backends this prevents a stale Store write from landing across a schema commit. A custom or non-transactional backend fails closed on the first managed write that cannot fuse the fence into its own statement — see [The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for which writes fuse and which refuse. `createStore()` and `createAdapterStore()` without `{ reconciled }` are raw, unversioned attaches. Their writes—and calls made directly through a backend—do not participate in the fence. `store.clear()` deletes the graph's schema rows and resets that Store to the same raw state; reopen it through a managed factory before resuming writes when the versioned guarantee is required. ### Store lifetime after a schema commit Managed Stores are immutable schema snapshots. A schema-changing operation such as `evolve()` returns the Store for the resulting schema; it does not update the instance on which it was called. Switch immediately to the returned Store for all subsequent work in the same request: ```typescript const evolved = await store.evolve(extension); await evolved.getNodeCollectionOrThrow("Paper").create({ title: "..." }); ``` For a long-lived local handle, pass a `StoreRef` and use either the return value or the updated `ref.current` after the call: ```typescript import type { StoreRef } from "@nicia-ai/typegraph"; const ref: StoreRef = { current: store }; const evolved = await ref.current.evolve(extension, { ref }); // `ref.current === evolved`; do not resume through the pre-evolve Store. await ref.current.getNodeCollectionOrThrow("Paper").create({ title: "..." }); ``` Capturing `ref.current` once at request entry is safe only for requests that do not change the schema. The ref also cannot observe commits made by another process or isolate. Before reusing a cross-request cache, compare the cached Store's `introspect().schemaVersion` (or its reconciled snapshot version) with `getCommittedSchemaVersion()`, then run `createVerifiedStore()` or `createVerifiedAdapterStore()` when the version changes. The [per-request connection recipe](/integration#per-request-connections-cache-the-verified-store) shows the complete single-flight cache pattern. ## Schema Validation Results The validation result indicates what happened during store initialization: | Status | Meaning | | ------------- | -------------------------------------------------- | | `initialized` | First run - schema version 1 was created | | `unchanged` | Schema matches stored version - no changes | | `migrated` | Safe changes auto-applied, new version created | | `pending` | Safe changes detected but `autoMigrate` is `false` | | `breaking` | Breaking changes detected, action required | The `initialized` and `migrated` results also include `committedRow: SchemaVersionRow`, the schema row that was just written. Most applications only need the version fields shown above, but integrations that build schema metadata can use `committedRow` without issuing another `getActiveSchema` read. ## Safe vs Breaking Changes ### Safe Changes (Auto-Migrated) These changes are backwards compatible and can be auto-migrated: - Adding new node types - Adding new edge types - Adding optional properties with defaults - Adding new ontology relations ### Breaking Changes (Require Manual Action) These changes require manual migration: - Removing node or edge types - Renaming node or edge types - Changing property types - Removing properties - Changing cardinality constraints to be more restrictive - Removing allowed endpoint pairs from a source-dependent edge ### Endpoint Pair Changes [Source-dependent targets](/core-concepts#source-dependent-targets) are part of the serialized schema. The `targetKindsBySource` field preserves the allowed pairs alongside the source and target kind lists, so export/import and schema round trips retain the restriction. For compile-time declarations, reordering map entries or target arrays does not change the schema hash. Persisted runtime extension documents also contribute to the hash and retain their array order. Narrowing a target map is breaking even when the overall source and target kind sets remain unchanged. For example, changing an edge from allowing every `Employee`/`Student` to `Department`/`Course` combination to allowing only `Employee → Department` and `Student → Course` removes two pairs. Existing rows using those pairs need migration before adopting the narrower schema. Adding allowed pairs, or changing the representation without removing any pairs, is nonbreaking. Runtime extension changes have an additional empty-kind check when tightening endpoints; see [extension edges](/graph-extensions#edges). ## Handling Breaking Changes When breaking changes are detected: ```typescript const [store, result] = await createStoreWithSchema(graph, backend, { throwOnBreaking: false, // Don't throw, inspect instead }); if (result.status === "breaking") { console.log("Breaking changes detected:"); console.log("Summary:", result.diff.summary); console.log("Required actions:"); for (const action of result.actions) { console.log(` - ${action}`); } // Option 1: Fix your schema to be backwards compatible // Option 2: Force migration (data loss possible!) // import { migrateSchema } from "@nicia-ai/typegraph/schema"; // await migrateSchema(backend, graph, currentVersion); } ``` ### Pre-flighting before you commit Both checks below are **SELECT-only** — no DDL, no writes — so a least-privilege runtime can decide what to do *before* it hits the privileged migration wall: ```typescript import { classifySchemaChanges } from "@nicia-ai/typegraph/schema"; // Cheapest: does this need the privileged path at all? // (true when the schema is behind, and when nothing is committed yet) if (await store.requiresMigration()) { // Route to the privileged bootstrap instead of failing mid-request. } // Or get the three-way decision: const diff = await store.schemaChanges(); const classification = diff === undefined ? "uninitialized" : classifySchemaChanges(diff); // "identical" | "additive" | "incompatible" ``` ### Classifying a failure If a commit does fail, branch on the structured outcome rather than the message text, which is free to be reworded in any release: ```typescript import { MigrationError } from "@nicia-ai/typegraph"; try { await commitSomething(); } catch (error) { if (error instanceof MigrationError) { switch (error.details.reason) { case "schema-behind": { // The runtime can't migrate. `diff` says whether it's safe to proceed. const additive = error.details.diff?.hasBreakingChanges === false; break; } case "breaking-change": { break; } case "kind-removal": { // The commit would drop a kind that still holds rows. Narrowing on // `reason` makes `droppedKinds` non-optional — the details type is a // discriminated union, so each reason carries exactly its own payload. const { nodes, edges } = error.details.droppedKinds; console.error("still populated:", [...nodes, ...edges]); break; } // "no-active-version" | "version-not-found" } } } ``` `details.reason` is a stable discriminant (the `MIGRATION_FAILURE_REASONS` union), and `details.diff` carries the same structured diff — with per-change `severity` — that `getSchemaChanges` returns, so you never need a second query to decide. ## Schema Introspection ### What Does This Database Already Have? `getActiveSchema` returns the committed schema document — the same JSON stored in `typegraph_schema_versions.schema_doc`, parsed into a `SerializedSchema`. Read it instead of querying that table by hand: ```typescript import { getActiveSchema, isSchemaInitialized, type SerializedSchema } from "@nicia-ai/typegraph"; // Check whether this graph has been committed at all const initialized = await isSchemaInitialized(backend, "my_graph"); const schema: SerializedSchema | undefined = await getActiveSchema(backend, "my_graph"); if (schema) { console.log("Version:", schema.version); console.log("Nodes:", Object.keys(schema.nodes)); // ["Person", "Company"] console.log("Edges:", Object.keys(schema.edges)); // ["worksAt"] } ``` These are exported from both the package root and the `@nicia-ai/typegraph/schema` subpath. Reach for `getCommittedSchemaVersion` instead when you only need the version number — for example, to invalidate a cached schema across isolates. ### Previewing Pending Changes ```typescript import { getSchemaChanges } from "@nicia-ai/typegraph/schema"; const diff = await getSchemaChanges(backend, graph); if (diff?.hasChanges) { console.log("Pending changes:", diff.summary); console.log("Is backwards compatible:", !diff.hasBreakingChanges); } ``` ## Manual Migration For full control over migrations: ```typescript import { initializeSchema, migrateSchema, rollbackSchema, ensureSchema } from "@nicia-ai/typegraph/schema"; // Initialize schema (first run only) const row = await initializeSchema(backend, graph); console.log("Created version:", row.version); // Migrate to new version. Folds the persisted graph extension into `graph` // first, and refuses (MigrationError, reason "kind-removal") if the commit // would drop a kind that still holds rows. const newVersion = await migrateSchema(backend, graph, currentVersion); console.log("Migrated to version:", newVersion); // Rollback to a previous version await rollbackSchema(backend, "my_graph", 1); console.log("Rolled back to version 1"); // Or use ensureSchema for automatic handling const result = await ensureSchema(backend, graph, { autoMigrate: true, throwOnBreaking: true, }); ``` ## Migrating Legacy Embedding Storage Embeddings now live in per-`(graphId, kind, field)` typed tables (`tg_vec___`), provisioned by `createStoreWithSchema` (the privileged migrator) at boot. This replaces the single shared `typegraph_node_embeddings` table. New deployments need no action — the per-field tables are materialized by `createStoreWithSchema`, which the legacy migration below also relies on having run. Deployments that already hold rows in the legacy table run a one-time, idempotent cutover with `migrateLegacyEmbeddings()`, exported from the package root: ```typescript import { migrateLegacyEmbeddings } from "@nicia-ai/typegraph"; // `backend` is the post-cutover backend, wired with its VectorStrategy. const result = await migrateLegacyEmbeddings({ backend }); console.log("Rows migrated:", result.migrated); console.log("Per field:", result.perField); console.log("Skipped (dimension mismatch):", result.skippedDimensionMismatch); console.log("Legacy table existed:", result.legacyTablePresent); ``` The run re-inserts every legacy embedding into per-field storage and is a clean no-op on a fresh install or a re-run (`legacyTablePresent: false`). A non-empty `skippedDimensionMismatch` flags `(kind, field)` slots that held mixed dimensions and need a deliberate re-embed at a single dimension — see [`reembedVectorField`](/schema-evolution#changing-an-embedding-dimension). The vector and hybrid query API (`.similarTo()`, `store.search.vector`, `store.search.hybrid`) is storage-transparent and unchanged by this cutover. ## Migrating Preview Recorded Time The initial recorded-time preview stored timestamps directly in `recorded_from`, `recorded_to`, and the graph clock. Versioned anchors now keep the durable string API while recorded relations compare numeric revisions. **Stop writers and run the one-time migration before enabling `history: true` with the new library version.** `createStoreWithSchema` and `createVerifiedStore` validate the recorded table shapes during an async open and reject an unmigrated preview schema before returning a store: ```typescript import { deleteLegacyRecordedAnchorMap, migrateLegacyRecordedTime, migrateRecordedAnchor, } from "@nicia-ai/typegraph"; const result = await migrateLegacyRecordedTime({ backend }); console.log(result.graphs, result.anchors); // Translate anchors stored in an application-owned checkpoint table. const upgraded = await migrateRecordedAnchor({ backend, graphId: "event-materializer", anchor: oldTimestampOnlyAnchor, }); await checkpoints.replaceAnchor(oldTimestampOnlyAnchor, upgraded); // Do this only after every external checkpoint for the graph is upgraded. await deleteLegacyRecordedAnchorMap({ backend, graphId: "event-materializer", dropWhenEmpty: true, }); ``` The bundled SQLite and PostgreSQL backends provide the recorded-relation DDL needed by this rewrite. A custom backend that created the preview schema must implement `backend.recordedTableDdl(tableNames)` before running `migrateLegacyRecordedTime`; otherwise the migration throws `UnsupportedBackendCapabilityError` with `details.capability: "recordedTableDdl"`. The callback is invoked for the temporary and final name sets so the backend, rather than TypeGraph's portable entrypoint, remains the owner of dialect-specific table and index DDL. When the engine names primary-key constraints, each callback result must name the constraint for both name sets or for neither. A one-sided declaration throws `ConfigurationError` with `details.code: "RECORDED_DDL_CONSTRAINT_NAME_MISMATCH"` before the replacement tables are published. See [`recordedTableDdl` in the backend contract](/backend-setup#recorded-table-migration-ddl-recordedtableddl) when adapting this migration to a custom backend. The migration dense-ranks distinct legacy commit timestamps independently per graph, preserving their exact total order. It rewrites the recorded relations and clock atomically and retains a durable old-anchor mapping so downstream stores can migrate separately. Re-running it after the cutover is a no-op. `migrateRecordedAnchor` also accepts an already-versioned `r1` anchor, making a mixed old/new checkpoint pass idempotent. The synchronous `createStore` factory is an attach-only, zero-I/O path, so it cannot inspect table shapes during construction. If used with `history: true`, an unmigrated schema still fails loudly on the first recorded operation. Prefer one of the async factories above at application startup when early schema verification matters. The old allocator may have pushed a hot graph's physical timestamp ahead of real wall time. Migration preserves that value because lowering it would put the clock behind recorded relation boundaries. New commits advance the logical revision normally, while the physical component remains pinned until wall time catches up. During that window, diagonal reads use the inherited future valid time; recorded-only ordering and replay remain exact. The mapping is graph-scoped: the same timestamp can correspond to different revisions in different graphs. Keep writers stopped for the schema rewrite, and delete mapping rows only after every external checkpoint for that graph has been translated. `dropWhenEmpty: true` atomically drops the mapping table when the deleted graph was the final one. Without that option, the empty table is retained intentionally and can be dropped by your normal migration tooling. ## Repairing Inverted Validity Windows Older library versions could store a row whose validity window runs backwards (`valid_from > valid_to`). Such a row is readable at **no** coordinate at all: `asOf(t)` needs `valid_from <= t < valid_to`, and backwards bounds admit no `t`. The write paths no longer produce one — a write that stamps a lower bound the caller did not state now stores no bound rather than an inverting one, see [Open-left rows](/queries/temporal#open-left-rows-validfrom-is-undefined) — but **upgrading rewrites nothing**. Rows already stored that way keep their window and stay invisible until an operator repairs them, which is deliberate: an upgrade that silently made previously-invisible rows appear in historical queries would be the worse surprise. `repairInvertedValidityWindows` is that explicit action. It has two modes: `report` counts and writes nothing, `apply` normalizes the rows it counted to `valid_from = NULL` ("ended at T, start unknown"). ```typescript import { repairInvertedValidityWindows } from "@nicia-ai/typegraph"; // Diagnose. `report` reads through `execute`, a required backend member, so it // runs against ANY backend — including a history-capturing one and one with no // statement-execution support. const report = await repairInvertedValidityWindows({ backend: anyBackend, relations: "live-and-recorded", mode: "report", }); // report.counts.recordedNodes === undefined means NOT SCANNED, never "clean". // report.atomic === false means the counts came from per-relation snapshots. // Repair, with writers stopped. On a history-enabled store pass the RAW backend // you constructed it from: the repair mints no revision by design, and the // capture wrapper refuses raw statements. await repairInvertedValidityWindows({ backend: rawBackend, relations: "live-and-recorded", mode: "apply", }); ``` If `tableNames` is supplied, it patches `backend.tableNames`; unstated relation names keep the backend's configured values. A partial override never sends the other relations back to TypeGraph's built-in defaults. `relations` is **required**, and `"live-and-recorded"` is the recommended scope. Repairing only the live axis leaves the recorded twin carrying the inverted window, which re-materializes the invisible row at any `asOfRecorded` coordinate — the same defect one axis over. `"live"` is right in exactly two cases: the store captures no history and the `recorded_*` tables do not exist (scanning them is then an error, not a no-op), or you are deliberately keeping the recorded axis as an audit record of the pre-repair state and accept that historical `asOfRecorded` reads keep returning the invisible shape. What an operator must know before running it: 1. **Run `apply` with writers stopped**, the same guidance `migrateLegacyRecordedTime()` carries. A concurrent window-bearing update may fence its write on the validity lower bound it read, so a repair landing in between can make the peer's first `UPDATE` match no row. Store node and edge updates re-read and re-judge against the repaired bound; interchange records a per-row target-changed error instead of claiming the row was written. `report` needs no quiescing: it scans in a read-only transaction (`BEGIN` rather than SQLite's writer-reserving `BEGIN IMMEDIATE`, and `BEGIN … READ ONLY` on PostgreSQL), so it cannot write itself. 2. **Repaired rows become visible** at `asOf` coordinates before their end. That is the point, and it is a read-visibility change to historical queries. 3. **Outstanding `base@V` merge tokens are invalidated** for repaired rows — `valid_from` is part of the base content fingerprint, so a merge whose base token predates the repair fails its precondition afterwards. Quiesce merges, repair, then re-baseline branches. 4. **The repair mints no revision and bumps no `version`**, and does not move `updated_at`. It normalizes a storage convention for rows that were never observable at any coordinate; it is not a logical write. That is why `apply` is run against the raw backend, and why bypassing recorded-time capture here is intended rather than a workaround. 5. **`apply` refuses when a scanned relation stores non-canonical bounds** (SQLite only — PostgreSQL stores `timestamptz`, so a scanned relation always reports `nonCanonical: 0`). SQLite compares the bounds as text, so a non-canonical value cannot be classified without a timestamp semantics this repair does not own. The refusal is total: the whole call is rejected before any row is updated, so `apply` never repairs the rows it understood and skips the rest. `report` still counts them, in `nonCanonical` — normalize those bounds, or narrow the call with `graphId`, and re-run. 6. **On a backend without transactions the call still runs**, per relation, and says so with `report.atomic === false`: the counts may span snapshots, and a crash mid-`apply` can leave the live axis repaired and the recorded axis not. Re-run — each statement is idempotent and convergent, and a later `report` proves it converged. 7. **Repair before exporting a legacy graph.** An exported inverted row is refused per row on re-import, so an unrepaired graph does not round-trip. The statement touches only rows the library mis-stored, so it is empty on a healthy graph and needs no batching. If a report returns a count large enough to worry about, narrow the call with `graphId` and run it per graph. ## Schema Serialization Schemas are stored as JSON documents with computed hashes for fast comparison: ```typescript import { serializeSchema, computeSchemaHash } from "@nicia-ai/typegraph/schema"; // Serialize a graph definition const serialized = serializeSchema(graph, 1); // Compute hash for comparison const hash = computeSchemaHash(serialized); ``` The serialized schema includes: - Graph ID and version - All node types with their Zod schemas (as JSON Schema) - All edge types with endpoints and constraints - Complete ontology relations - Uniqueness constraints and delete behaviors ## Version History TypeGraph maintains a history of all schema versions: ```text typegraph_schema_versions ├── version 1 (initial) ├── version 2 (added User node) ├── version 3 (added email property) ← active └── ... ``` Only one version is marked as "active" at a time. Previous versions are preserved for auditing and potential rollback. ## Best Practices ### 1. Use Managed Stores in Production ```typescript // Production: Use schema management const [store, result] = await createStoreWithSchema(graph, backend); // Development: Basic store is fine for rapid iteration const store = createStore(graph, backend); ``` ### 2. Check Migration Status on Startup ```typescript async function initializeApp() { const [store, result] = await createStoreWithSchema(graph, backend); if (result.status === "breaking") { console.error("Database schema incompatible with application!"); console.error("Run migrations before deploying this version."); process.exit(1); } if (result.status === "migrated") { console.log(`Schema auto-migrated to v${result.toVersion}`); } return store; } ``` ### 3. Preview Changes Before Deployment ```typescript import { getSchemaChanges } from "@nicia-ai/typegraph/schema"; // In your CI/CD pipeline or migration script const diff = await getSchemaChanges(backend, graph); if (diff?.hasChanges) { console.log("Schema changes detected:"); console.log(diff.summary); if (!diff.isBackwardsCompatible) { console.error("Breaking changes require manual migration!"); process.exit(1); } } ``` ### 4. Add Properties with Defaults When adding new properties, always provide defaults to ensure backwards compatibility: ```typescript // Good: Optional with default const User = defineNode("User", { schema: z.object({ name: z.string(), // New property with default - safe migration status: z.enum(["active", "inactive"]).default("active"), }), }); // Bad: Required without default - breaking change const User = defineNode("User", { schema: z.object({ name: z.string(), status: z.enum(["active", "inactive"]), // No default! }), }); ``` # Schemas & Stores > Schema definition functions and store API reference This reference documents the schema definition functions and store API for TypeGraph. ## Schema Definition ### `defineGraph(config)` annotations `defineGraph` accepts consumer-owned JSON metadata for the schema as a whole: ```typescript const graph = defineGraph({ id: "support", annotations: { displayName: "Support knowledge graph", description: "Queryable product and incident knowledge", capabilities: { search: true, temporal: true }, }, nodes: { Incident: { type: Incident } }, edges: {}, }); ``` The value is available as `graph.annotations`, is persisted in `SerializedSchema.annotations`, and is returned by `store.introspect().annotations`. It follows the same JSON-only validation and canonical hashing rules as per-kind annotations. An absent or empty object is omitted from canonical form and retains legacy hashes. ### `defineNode(name, options)` Creates a node type definition. ```typescript import { defineNode } from "@nicia-ai/typegraph"; function defineNode>( name: K, options: { schema: S; description?: string; annotations?: Readonly>; }, ): NodeType; ``` **Parameters:** | Parameter | Type | Description | |-----------|------|-------------| | `name` | `string` | Unique name for this node type | | `options.schema` | `z.ZodObject` | Zod object schema for node properties | | `options.description` | `string` | Optional description | | `options.annotations` | `KindAnnotations` | Optional consumer-owned per-kind annotations. See [Per-kind annotations](#per-kind-annotations). | **Example:** ```typescript const Person = defineNode("Person", { schema: z.object({ name: z.string(), email: z.string().email().optional(), }), description: "A person in the system", }); ``` **With annotations:** ```typescript const Incident = defineNode("Incident", { schema: z.object({ title: z.string(), summary: z.string(), occurredAt: z.string().datetime(), }), annotations: { ui: { titleField: "title", temporalField: "occurredAt", icon: "alert-triangle", }, audit: { pii: false, retentionDays: 365, }, }, }); ``` ### `defineEdge(name, options?)` Creates an edge type definition. ```typescript import { defineEdge } from "@nicia-ai/typegraph"; function defineEdge>( name: K, options?: { schema?: S; description?: string; annotations?: Readonly>; from?: NodeType[]; to?: NodeType[]; }, ): EdgeType; ``` **Parameters:** | Parameter | Type | Description | |-----------|------|-------------| | `name` | `string` | Unique name for this edge type | | `options.schema` | `z.ZodObject` | Optional Zod object schema (defaults to empty object) | | `options.description` | `string` | Optional description | | `options.annotations` | `KindAnnotations` | Optional consumer-owned per-kind annotations. See [Per-kind annotations](#per-kind-annotations). | | `options.from` | `NodeType[]` | Optional domain constraint (valid source node types) | | `options.to` | `NodeType[]` | Optional range constraint (valid target node types) | **Example:** ```typescript const worksAt = defineEdge("worksAt", { schema: z.object({ role: z.string(), startDate: z.string().optional(), }), }); const knows = defineEdge("knows"); // No schema needed ``` **With annotations:** ```typescript const reportedBy = defineEdge("reportedBy", { schema: z.object({ channel: z.string() }), from: [Incident], to: [Person], annotations: { ui: { showInTimeline: true, badge: "report" }, }, }); ``` **With Domain/Range Constraints:** When `from` and `to` are specified, the edge carries its endpoint constraints intrinsically: ```typescript const worksAt = defineEdge("worksAt", { schema: z.object({ role: z.string(), startDate: z.string().optional(), }), from: [Person], // Domain: only Person can be the source to: [Company], // Range: only Company can be the target }); ``` **Unconstrained Edges:** Edges without `from`/`to` are unconstrained — they can connect any node type to any node type: ```typescript const sameAs = defineEdge("sameAs"); const related = defineEdge("related", { schema: z.object({ reason: z.string() }), }); ``` **Direct use in defineGraph:** Any edge type can be used directly in `defineGraph` without an `EdgeRegistration` wrapper: ```typescript const graph = defineGraph({ id: "my_graph", nodes: { Person: { type: Person }, Company: { type: Company } }, edges: { worksAt, // Constrained — uses built-in from/to sameAs, // Unconstrained — connects any node to any node }, }); ``` See [Core Concepts](/core-concepts#domain-and-range-constraints) for detailed documentation on domain/range constraints. ### Per-kind annotations Both `defineNode` and `defineEdge` accept an optional `annotations` field — a plain JSON object for consumer-owned, structured per-kind data that doesn't belong in the Zod schema. Common uses: - **Generic UI rendering.** Which property is the title for list views? Which is the canonical date for sorting? Which icon represents the kind? - **Audit and compliance hints.** Mark a kind as PII, set retention windows, attach data-classification labels. - **Tooling annotations.** Group kinds in catalogs, mark provenance ("originated from agent run X"), attach feature-flag gates. ```typescript const Incident = defineNode("Incident", { schema: z.object({ title: z.string(), occurredAt: z.string().datetime(), }), annotations: { ui: { titleField: "title", temporalField: "occurredAt", icon: "alert-triangle" }, audit: { pii: false, retentionDays: 365 }, }, }); ``` Reading annotations back from a kind: ```typescript const titleField = (Incident.annotations?.ui as { titleField?: string })?.titleField; ``` Or from a stored schema: ```typescript import { getSchemaChanges, getActiveSchema } from "@nicia-ai/typegraph/schema"; const stored = await getActiveSchema(backend, "my_graph"); const incidentMeta = stored?.nodes.Incident?.annotations; ``` **Key guarantees and constraints:** - **TypeGraph never reads, validates, or interprets keys inside `annotations`.** Consumers own the entire namespace — no reserved prefixes, no `x-typegraph` extension convention. Future library-owned per-kind state, if needed, will use a separate sibling field rather than carving out keys here. - **Annotations participate in schema hashing and migration diffs.** Changing `annotations` bumps the schema version like any other structural change, and the diff is reported as a `safe`-severity change per kind. See [Schema Evolution](/schema-evolution#changing-annotations). - **Values must be JSON-serializable.** Strings, numbers, booleans, `null`, arrays, and plain objects only. `bigint`, `function`, `symbol`, `undefined`, `Date`, `Map`, `Set`, and other class instances are rejected at definition time with a `ConfigurationError` so they can never silently break hashing or storage round-trips. - **Default is `undefined`, not `{}`.** Graphs that never set `annotations` produce identical canonical-form hashes to graphs from before this field existed — adoption requires no migration. An explicit empty object (`{}`) is a structural opt-in and bumps the hash. - **Annotations are not a typed contract.** TypeScript types them as `Readonly>`. Wrap reads in your own typed accessors at consumer boundaries if you need stronger guarantees. ### `embedding(dimensions, options?)` Creates a Zod schema for vector embeddings with dimension validation. Carries optional vector-index configuration that the auto-derivation pass at `defineGraph()` time reads to produce `VectorIndexDeclaration` entries — see [Graph Extensions → Vector indexes](/graph-extensions#vector-indexes) for the full materialization flow. ```typescript import { embedding } from "@nicia-ai/typegraph"; function embedding( dimensions: D, options?: EmbeddingIndexOptions, ): EmbeddingSchema; type EmbeddingIndexOptions = Readonly<{ /** Distance metric. Default `"cosine"`. */ metric?: "cosine" | "l2" | "inner_product"; /** Vector index implementation. Default `"hnsw"`. */ indexType?: "hnsw" | "ivfflat" | "none"; /** HNSW: max connections per layer. Default `16`. */ m?: number; /** HNSW: build-time search depth. Default `64`. */ efConstruction?: number; /** IVFFlat: number of inverted-list partitions. */ lists?: number; }>; ``` **Parameters:** | Parameter | Type | Description | |-----------|------|-------------| | `dimensions` | `number` | Number of dimensions (e.g., 384, 512, 768, 1536, 3072) | | `options` | `EmbeddingIndexOptions?` | Optional index configuration. Defaults match pgvector recommendations. Pass `{ indexType: "none" }` to opt out of automatic materialization while keeping the embedding column. | **Example:** ```typescript // Defaults: cosine similarity, HNSW index, m=16, ef_construction=64. const Document = defineNode("Document", { schema: z.object({ title: z.string(), content: z.string(), embedding: embedding(1536), // OpenAI ada-002 }), }); // Override at the brand site — this is the load-bearing place to // signal index intent because the metric usually reflects model // output (cosine-normalized vs. raw inner-product). const Image = defineNode("Image", { schema: z.object({ embedding: embedding(512, { metric: "l2", m: 32, efConstruction: 100 }), }), }); // Opt out of automatic materialization while keeping the embedding column. const Manual = defineNode("Manual", { schema: z.object({ embedding: embedding(384, { indexType: "none" }), }), }); // Optional embeddings work as before — the brand survives `.optional()` / // `.nullable()` wrappers and auto-derivation walks through them. const Article = defineNode("Article", { schema: z.object({ content: z.string(), embedding: embedding(1536).optional(), }), }); ``` See [Semantic Search](/semantic-search) for query usage and [Graph Extensions](/graph-extensions#vector-indexes) for how the auto-derived index flows through `materializeIndexes()`. ### `externalRef(table)` Creates a Zod schema for referencing external data sources. Use this for hybrid overlay patterns where TypeGraph stores relationships while your existing tables remain the source of truth. ```typescript import { externalRef } from "@nicia-ai/typegraph"; function externalRef(table: T): ExternalRefSchema; ``` **Parameters:** | Parameter | Type | Description | |-----------|------|-------------| | `table` | `string` | Identifier for the external table (e.g., "users", "documents") | **Example:** ```typescript const Document = defineNode("Document", { schema: z.object({ source: externalRef("documents"), embedding: embedding(1536).optional(), }), }); // Create with explicit table reference await store.nodes.Document.create({ source: { table: "documents", id: "doc_123" }, }); // Query the external reference const results = await store .query() .from("Document", "d") .select((ctx) => ctx.d.source) .execute(); // results[0].source = { table: "documents", id: "doc_123" } ``` ### `createExternalRef(table)` Factory helper to create external reference values without repeating the table name. ```typescript import { createExternalRef } from "@nicia-ai/typegraph"; function createExternalRef( table: T ): (id: string) => ExternalRefValue; ``` **Example:** ```typescript const docRef = createExternalRef("documents"); await store.nodes.Document.create({ source: docRef("doc_123"), // { table: "documents", id: "doc_123" } }); ``` ### `defineGraph(config)` Creates a graph definition combining nodes, edges, and ontology. ```typescript import { defineGraph } from "@nicia-ai/typegraph"; function defineGraph(config: { id: string; nodes: Record; edges: Record; ontology?: OntologyRelation[]; indexes?: IndexDeclaration[]; defaults?: { onNodeDelete?: DeleteBehavior; temporalMode?: TemporalMode; }; }): G; ``` **Parameters:** | Parameter | Type | Description | |-----------|------|-------------| | `id` | `string` | Unique identifier for this graph | | `nodes` | `Record` | Node type registrations | | `edges` | `Record` | Edge registrations or edge types directly | | `ontology` | `OntologyRelation[]` | Optional semantic relationships | | `indexes` | `IndexDeclaration[]` | Optional explicit index declarations from `defineNodeIndex` / `defineEdgeIndex`. Vector indexes are auto-derived from `embedding()` brands; explicit declarations win on `(kind, fieldPath)` collisions. | | `defaults` | `{ onNodeDelete?, temporalMode? }` | Optional graph-wide defaults. `onNodeDelete` defaults to `"restrict"`; `temporalMode` defaults to `"current"`. | Edge entries can be: - **`EdgeRegistration`** — explicit `{ type, from, to }` with optional `cardinality` - **`EdgeType` with `from`/`to`** — uses built-in constraints - **`EdgeType` without `from`/`to`** — unconstrained, connects any node to any node **Example:** ```typescript const graph = defineGraph({ id: "my_graph", nodes: { Person: { type: Person }, Company: { type: Company, onDelete: "cascade" }, }, edges: { worksAt: { type: worksAt, from: [Person], to: [Company], cardinality: "many", }, sameAs, // Unconstrained — any→any }, ontology: [disjointWith(Person, Company)], }); ``` ### Graph identity and the kind namespace `id` is the isolation boundary, and it is the most important operational fact about a TypeGraph deployment: **kinds are scoped to the `graph_id`** — not to the module, file, or proposal that declared them. Everything committed under one `id` shares a single schema document and a single set of physical collections: - Two graphs with the **same `id`** share one kind namespace. Committing a graph that declares `Invoice` makes `Invoice` part of *that graph's* schema for every process that opens it — including ones whose compile-time graph never mentioned it. Re-committing a different graph under the same `id` is a **schema change**, diffed against what is already committed; it is not a separate scope. - Two graphs with **different `id`s** are fully independent: separate kind namespaces, separate data, separate schema versions. They can hold divergent schemas in the same database. The practical consequence: a namespace *is* a `graph_id`. If you want two units (tenants, test suites, per-customer graphs) to declare kinds without colliding, give them different `id`s — do not rely on separate declaration sites, module boundaries, or naming conventions to isolate them. A shared test database where every suite commits under one `id` will see those suites fight over one schema. See [Multi-Tenant Architecture](/integration#multi-tenant-architecture) for running many `graph_id`s in a single database. ### Pre-flighting a schema change Committing a schema is a privileged, migration-gated operation — in a least-privilege deployment the runtime role cannot run DDL at all. Both of these are **SELECT-only**: they report what a commit *would* do without attempting it. ```typescript import { classifySchemaChanges } from "@nicia-ai/typegraph/schema"; // Cheapest check: does this need the privileged path at all? if (await store.requiresMigration()) { // Route to the privileged bootstrap instead of failing mid-request. } // Or decide additive-vs-incompatible before committing: const diff = await store.schemaChanges(); const classification = diff === undefined ? "uninitialized" : classifySchemaChanges(diff); // "identical" | "additive" | "incompatible" ``` `requiresMigration()` is `true` when the committed schema is behind the graph **and** when nothing has been committed yet, since both need the privileged path. If a commit does fail, branch on the structured outcome rather than the message text (which is free to be reworded in any release): `MigrationError.details.reason` is a stable discriminant (`MIGRATION_FAILURE_REASONS`), and `details.diff` carries the same structured diff — with per-change `severity` — that `getSchemaChanges` returns. See [Handling Breaking Changes](/schema-management#handling-breaking-changes). ## Store Creation ### `createStore(graph, backend, options?)` Creates the portable store contract for a graph definition. It contains the complete TypeGraph API and graph-owned transactions, but deliberately omits adapter-native handles, caller-owned transaction adoption, and mutable backend internals. This synchronous factory performs no database I/O, including schema shape checks. Use an async factory below when startup must verify storage. Because it carries no committed schema-version metadata, its writes are raw and are not fenced against concurrent schema changes. ```typescript import { createStore } from "@nicia-ai/typegraph"; function createStore( graph: G, backend: GraphBackend, options?: StoreOptions ): Store; ``` **Options:** | Option | Type | Description | |--------|------|-------------| | `hooks` | `StoreHooks` | Observability hooks for monitoring operations | | `history` | `boolean` | Enable built-in recorded / system-time capture: every committed TypeGraph node/edge write is captured into the recorded-time relations read by [`store.asOfRecorded(T)`](/queries/temporal#recorded-time-bitemporal) (default: `false`) | | `recordedRead` | `ExternalRecordedReadSource` | Bind an already-populated recorded relation for `store.asOfRecorded(T)` reads without enabling TypeGraph-managed capture. Must be created with `recordedRelation({ schema })` using a `createSqlSchema(...)` schema; the store validates those factory descriptors at runtime. Use `history: true` when TypeGraph should capture writes and advance `store.recordedNow()`. | | `schema` | `SqlSchema` | Custom table name configuration created with `createSqlSchema(...)` | | `queryDefaults.traversalExpansion` | `TraversalExpansion` | Default ontology expansion mode for traversals (default: `"inverse"`) | | `autoRefreshStatistics` | `false \| number` | Row threshold at which a single autocommit `bulkCreate`/`bulkInsert` triggers an automatic planner-statistics refresh (default: `1000`); `false` disables. See [Refreshing planner statistics](/backend-setup#refreshing-planner-statistics-after-bulk-loads). | | `coalesceUnchangedUpserts` | `boolean` | Skip the write for an `upsertById` or endpoint get-or-create update whose validated props and requested window already equal the existing live row; bulk forms behave identically (default: `false`). Node `getOrCreateByConstraint` updates are not coalesced; use `upsertById` for replay projectors that must avoid unchanged node history churn. For at-least-once / replay materializers: a byte-identical re-delivery performs no write, no history row, and no revision advance. See [`upsertById`](#upsertbyidid-props-options), [`getOrCreateByEndpoints`](#getorcreatebyendpointsfrom-to-props-options), and [Materializing external event logs](/materializing-event-logs). | **Example:** ```typescript const store = createStore(graph, backend); ``` When an application owns an adapter connection and must coordinate native SQL with TypeGraph, use `createAdapterStore(graph, adapterBackend)` instead. It returns `AdapterStore`, which adds precisely typed `tx.sql`, `withTransaction`, `withRecordedTransaction`, and the adapter backend surface. A plain `GraphBackend` cannot be passed to this factory. `createAdapterStore` is likewise raw unless passed a cached `{ reconciled }` snapshot. Writes issued directly through a backend are always outside the Store schema fence. Override the default traversal expansion: ```typescript const store = createStore(graph, backend, { queryDefaults: { traversalExpansion: "none" }, }); ``` ### `createStoreWithSchema(graph, backend, options?)` Creates a store and ensures the database schema is initialized or migrated. This is the recommended factory for production use, and it is **required** for any graph with `searchable()` fields: it durably materializes the fulltext storage. Bare `createStore()` does not, and the first fulltext operation against an uninitialized database throws `StoreNotInitializedError`. The privileged open also adopts versioned, deployment-wide physical base storage even when the per-graph TypeGraph schema document is unchanged. Base-schema version 1 covers the durable graph-template relation and nullable edge match-identity columns and arbiter. Fresh and published schemas also carry the nullable-pair CHECK; SQLite adoption does not require rebuilding an externally managed table solely to add that defensive constraint. The marker advances only after all adoption steps succeed. A database provisioned by an older release must either be opened once with `createStoreWithSchema()` under a DDL-capable role or receive the published additive migration before a DML-only runtime attaches. With `history: true`, the async open also verifies the recorded node, edge, and clock column shapes before returning. Databases created by the timestamp-only recorded-time preview must run `migrateLegacyRecordedTime({ backend })` first; an unmigrated schema throws a typed `ConfigurationError` with `details.code === "RECORDED_SCHEMA_INCOMPATIBLE"` at open rather than on the first write. For an existing history-enabled database that predates identity enablement, open it once through `createStoreWithSchema(...)` after adding `defineGraph(...).identity`. Bundled backends provision the identity relations before the migration preflight. A later missing relation reports `details.code === "IDENTITY_STORAGE_MISSING"` instead of opening over silently empty state. Restore a missing assertion ledger from backup. If only the derived closure is missing, recreate that relation with TypeGraph's standard DDL, then run `rebuildIdentityClosure()` before serving traffic. Any `history: true` open of an identity-enabled graph verifies the recorded identity relation exists and reports `RECORDED_IDENTITY_SCHEMA_MISSING` if it does not; bundled backends provision it, so this is rare there and more likely on a custom backend that skipped it. Restore that recorded ledger from backup; recreating it empty would silently discard identity history. Provision an empty relation through the backend's privileged setup path only when this is confirmed first-time identity enablement with no identity history to preserve. Malformed columns in an existing relation continue to report `RECORDED_SCHEMA_INCOMPATIBLE`. ```typescript import { createStoreWithSchema } from "@nicia-ai/typegraph"; function createStoreWithSchema( graph: G, backend: GraphBackend, options?: StoreOptions & SchemaManagerOptions ): Promise<[Store, SchemaValidationResult]>; ``` **Returns:** A tuple of `[store, validationResult]` The validation result indicates what happened: - `status: "initialized"` - Schema created for the first time - `status: "unchanged"` - Schema matches, no changes needed - `status: "migrated"` - Safe changes auto-applied (additive only) - `status: "pending"` - Safe changes detected but `autoMigrate` is `false` - `status: "breaking"` - Breaking changes detected, action required For `initialized` and `migrated`, the result also includes `committedRow: SchemaVersionRow`, which is the row TypeGraph just committed. Most callers can ignore it; it is useful when building schema metadata without performing another active-schema lookup. **Example:** ```typescript const [store, result] = await createStoreWithSchema(graph, backend); if (result.status === "initialized") { console.log("Schema initialized at version", result.version); } else if (result.status === "migrated") { console.log(`Migrated from v${result.fromVersion} to v${result.toVersion}`); } else if (result.status === "pending") { console.log(`Safe changes pending at version ${result.version}`); } ``` **Throws:** `MigrationError` if breaking changes are detected and `throwOnBreaking` is `true` (the default). Use `createAdapterStoreWithSchema` for the same provisioning behavior with an `AdapterStore` result. This explicit factory is required for native transaction adoption or `tx.sql`; schema provisioning alone does not expose adapter capabilities on the portable `Store`. #### Schema-managed write fence A Store is schema-managed when `store.introspect().schemaVersion !== undefined`. This includes Stores opened by `createStoreWithSchema`, `createAdapterStoreWithSchema`, `createVerifiedStore`, or `createVerifiedAdapterStore`; `createAdapterStore(..., { reconciled })`; Stores returned by `evolve()`; and Stores rebound from an already-managed Store. The official transactional backends fence and revalidate every managed write against schema commits. Rechecking matters when adapter-native SQL rolls back to a savepoint, because PostgreSQL releases row locks acquired after that savepoint. A custom or non-transactional backend without fence support fails closed when a managed write is attempted rather than racing. Raw Stores and direct backend writes remain available when the application deliberately owns schema/write coordination. Because a managed Store is pinned to the schema version it opened, use the Store returned by `evolve()` for every subsequent operation in that request. A Store captured before the commit is not updated, and its next managed write is rejected by this fence. See [Store lifetime after a schema commit](/schema-management#store-lifetime-after-a-schema-commit) for the `StoreRef` and cross-process cache patterns. #### Managed-write latency and batching A single managed write is intentionally a small read/write protocol, not one blind `INSERT` or `UPDATE`. Its statement count includes the schema-version fence, the per-graph fence for a declared check-then-write constraint, identity coordination when a caller-supplied node id can fold across kinds, the operation's existence/endpoint/constraint probes, and the row plus its claim and sidecar writes. The exact count depends on the graph declaration and backend capabilities. Outside an explicit Store transaction, the schema fence is rechecked for every managed write. Inside one `store.transaction(...)` callback, TypeGraph acquires the fence on the pinned transaction target and leases that held fence to later writes in the callback. Adapters must not introduce a broader cache: PostgreSQL releases row locks acquired after a savepoint when the transaction rolls back to that savepoint, and a backend-instance cache could therefore outlive the database protection it claims. The per-graph write lock follows the same transaction-scoped ownership rule. Endpoint, duplicate-id, claim, and projection work is folded only through the backend's semantic command port, never into a blind write. A command either applies every requested dimension atomically or returns `unsupported` before issuing SQL, after which the Store re-enters the portable validation path. That fallback is available only when the backend can preserve the requested atomicity; a root backend that cannot keep a row and its projection sidecars together refuses the write with a typed transaction-required error. Node creates still distinguish live duplicates from tombstones, and edge writes still preserve endpoint and claim ordering. `getOrCreateByEndpoints` uses an outside read only for its no-write found fast path; every create, resurrection, or update makes its decision on the fenced transaction target. An adopted PostgreSQL transaction can therefore return an existing match at any isolation level. If the operation reaches the create leg, TypeGraph checks the effective isolation captured by the graph-lock statement and refuses repeatable read before issuing the convergent write. Dynamic `matchOn` remains transaction-scoped because a call-level field list has no database uniqueness object. For latency-sensitive paths, declare a graph-local `matchIdentity` on the edge registration. TypeGraph stores its canonical endpoint/property key on every edge row and both bundled dialects enforce it with a unique database arbiter. PostgreSQL and SQLite then lower an eligible root, single-item `getOrCreateByEndpoints` create/found decision, endpoint validation, and schema fence to one conflict-arbitrated statement. Eligibility requires a bundled root backend (including D1/neon-http), a schema-managed generated-id node or `cardinality: "many"` edge, and no claims, sidecars, history, or revision work. Derived wrappers, adopted transactions, custom backends, and constrained edges remain on the interactive path. On a Neon WebSocket connection this removes the dispatcher read plus `BEGIN`, graph-lock, and `COMMIT` exchanges: the common eligible miss falls from roughly five sequential requests to one. Constrained cardinalities and history/revision stores retain a transaction and their required sidecar/fence work, so they are not eligible for the one-request root path; declared identities are still arbitrated by the durable key inside that transaction. Undeclared dynamic matches retain the fenced portable path and fail closed when a backend cannot provide it. For ingestion, bundled PostgreSQL roots using a recognized session-capable driver, Neon HTTP, Cloudflare D1, and libSQL roots have a narrower native path for schema-managed `nodes.bulkInsert()` and `nodes.bulkCreate()`: generated, caller-supplied, or mixed IDs can run as one schema-fenced atomic program when there is no Operational Identity, history, or revision work. The program composes fulltext/vector projections with its supported uniqueness and disjointness claim envelope. Session-capable PostgreSQL pins the program to one transaction; Neon HTTP, D1, and libSQL submit one transport batch. `bulkCreate()` restores rows to input order. Multiple claims, hierarchy-wide uniqueness scopes, generated/caller ID mixtures, and claim-plus-projection work share the program; compatibility probes preserve legacy claim-axis rows. Identity, history/revision capture, or a member beyond the executor's claim-input budget keep the existing transaction or fallback path. This is an internal implementation detail, not a general Store batch surface. The same bundled roots execute eligible node and edge `bulkDelete()` calls as closed schema-fenced mutation programs. Edge collection identity and node `restrict` behavior are rechecked by the write statement itself. Restricted node programs release owner-side uniqueness and disjointness claims in the same exchange. Nodes with search, identity, or capture sidecars, or with cascade or disconnect behavior, retain the interactive path, as do derived/custom backends and caller-owned transactions. On a portable backend, edge bulk deletion still resolves all IDs in one batched read and applies them through a set-based delete port when available, replacing the former per-ID loop. Keep these guarantees distinct: - An **interactive transaction** is the public `store.transaction(...)` API; it pins a session and groups the callback's Store writes. Internally, `runOptionallyInTransaction` reports this as `{ mode: "interactive-transaction" }`, or `{ mode: "sequential" }` when it cannot open one. - A **static internal adapter batch** is a backend implementation detail, such as a D1 batch or a bind-budgeted multi-row insert. It is not a Store API and does not make arbitrary Store calls atomic. - A **certified atomic SQL program** is a closed, ordered sequence submitted to a backend transport that has passed the framework conformance runner. The runner checks ordered result slots, exact parameter forwarding, empty-batch no-op behavior, and rollback of primary and sidecar writes after a later statement fails. This transport capability is separate from the semantic proof that makes a particular mutation eligible. - An **authoritative one-statement command** is the semantic `commands` port; its statement returns the created/found decision it owns. Durable `matchIdentity` convergence can qualify for this root path because the schema-declared key has a database arbiter. Operational Identity, claim/cardinality checks, undeclared dynamic `matchOn`, and history/revision sidecars remain interactive-transaction contracts. A durable edge match identity only authorizes its own canonical edge create/found arbitration; it does not make those other responsibilities transactionless. For networked deployments, amortize the safe costs at the call boundary: - Use `bulkCreate`, `bulkInsert`, `bulkUpsertById`, and the bulk get-or-create methods for batches. They batch endpoint and node validation and use one transaction and set-based writes where the backend supports them. - Group several related writes in `store.transaction(async (tx) => ...)` to amortize transaction framing and per-transaction graph-lock acquisition. Always use the `tx` collections inside the callback. Calling the root Store there can open a second transaction or use the wrong connection. - Keep schema-managed writes on a transactional backend. A non-transactional adapter cannot provide the schema or constraint fences and fails closed for those writes. There is no supported option to disable these checks for a single write. If an application has already established stronger invariants, it may use a direct backend write under its own transaction and coordination policy, but that is outside the Store's typed validation, claim, sidecar, history, and identity contracts. ## Store Projection ### `StoreProjection` A type-level utility that projects a store's collection surface onto a subset of node and edge keys. Use this to type reusable helpers that work with any store containing a shared subgraph. ```typescript import type { StoreProjection } from "@nicia-ai/typegraph"; type CoreStore = StoreProjection< typeof myGraph, "Document" | "Chunk", "hasChunk" >; async function ingestChunk(store: CoreStore, document: Node, text: string) { const chunk = await store.nodes.Chunk.create({ text }); await store.edges.hasChunk.create(document, chunk); return chunk; } ``` Both `Store` and `TransactionContext` are structurally assignable to a `StoreProjection` whose keys are a subset of `G`. Node constraint names are erased so the projection works across graphs that register the same node types with different unique constraints. See [Shared Subgraph Helpers](./multiple-graphs#shared-subgraph-helpers) for a full example with multiple graphs. ## Store API The store provides typed node and edge collections via `store.nodes.*` and `store.edges.*`. On creation, every method that accepts `validFrom` (`create`, `createFromRecord`, `upsertById`, `upsertByIdFromRecord`, `bulkCreate`, `bulkInsert`, `bulkUpsertById`, and their edge equivalents) distinguishes three inputs: | Input | Stored lower bound | | --- | --- | | Omitted | The operation's creation timestamp, except for the born-ended case below | | Canonical UTC timestamp | That timestamp | | `null` | No lower bound: an explicitly **open-left** window | ```typescript const person = await store.nodes.Person.create( { name: "Alice" }, { validFrom: null }, ); // person.meta.validFrom === undefined ``` An open-left row is visible at every valid-time coordinate before its `validTo`, or at every coordinate when there is no end. Returned metadata uses `undefined` for an absent bound; pass `validFrom: row.meta.validFrom ?? null` to preserve that state when creating a copy. Passing `undefined` requests the creation default. `validTo` remains optional and open-ended until set. On a live upsert, omission preserves the stored lower bound. An explicit `null` restates an already open-left row and is refused if the live row has a timestamp, just as a different timestamp is refused. `onImmutableLowerBound: "preserve"` continues to make the stated bound creation/resurrection-only input. The one exception is a row **born already ended**: a write that CREATES or RESETS a row's window while stating a `validTo` at or before its own instant and no `validFrom` stores **no lower bound** ("ended at T, start unknown") rather than a start after its own end, and `meta.validFrom` reads back as `undefined`. Such a row is visible at every `asOf` coordinate before its end and at none after it. A `validTo` in the future is unaffected — it still stamps the creation timestamp, so the row is invisible at instants before it existed. One stated window reaches one stored shape. Every **node** write that resets the window takes the exception: `create`, `bulkCreate` and `bulkInsert` on a fresh id or on one naming a tombstone, and `upsertById` / `bulkUpsertById` resurrecting a tombstoned node — all of them store no lower bound for a lone historical `validTo`, and `meta.validFrom` reads back as `undefined` in every case. An **edge** is different, and deliberately: an edge create never lands on a tombstone (a taken id raises `Edge already exists`), and a tombstoned edge is reachable again only through `bulkUpsertById` or `getOrCreateByEndpoints`, both of which RETAIN the bound that row already carries and judge the stated `validTo` against it — so a `validTo` before that bound is refused. Pass `validFrom` alongside a historical `validTo` when an edge upsert may land on a tombstone. Writes that accept a validity-end mutation have three explicit states: omit both fields to preserve the stored end, pass `{ validTo }` to set or move it, or pass `{ clearValidTo: true }` to remove it and reopen the window. `validTo` and `clearValidTo` are mutually exclusive. Clearing is supported by node `update`, `upsertById`, `upsertByIdFromRecord`, and `bulkUpsertById`, plus edge `update`, `bulkUpsertById`, and both endpoint get-or-create forms. A window may not have negative width. Stating both endpoints out of order, or updating a row with a `validTo` that precedes its stored `validFrom`, raises a `ValidationError` whose issue carries the code `INVERTED_VALIDITY_WINDOW` — such a row stopped being true before it started, so no `asOf` coordinate could ever observe it. Two related shapes are legal: a ZERO-width window (`validTo === validFrom`), which is what a same-instant retraction produces; and a create carrying only a historical `validTo`, which means "born already ended" and is read back at any `asOf` coordinate before that end, or through the `includeEnded` temporal mode. ### Node Collections Each node type has a collection with these methods: #### Naming Guidelines Method names follow what identifier is used to match an existing record: | If you have... | Read-only | Get-or-create | |----------------|-----------|---------------| | ID | `getById` | `upsertById` | | Unique constraint name + props | `findByConstraint` | `getOrCreateByConstraint` | | Declared index name + records (candidates) | `bulkFindByIndex` | — | | Edge endpoints (`from`, `to`) + optional `matchOn` | `findByEndpoints` | `getOrCreateByEndpoints` | #### `create(props, options?)` Creates a new node. ```typescript store.nodes.Person.create( props: { name: string; email?: string }, options?: { id?: string; validFrom?: string | null; validTo?: string } ): Promise>; ``` #### `getById(id)` Retrieves a node by ID. ```typescript store.nodes.Person.getById(id: NodeId): Promise | undefined>; ``` When a persisted id crosses an untyped boundary, brand it before passing it to read/update/delete APIs: ```typescript const id = asNodeId(row.personId); const person = await store.nodes.Person.getById(id); ``` `create({ id })` and `upsertById` still accept plain strings because those are write surfaces that mint or claim ids. #### `getByIds(ids)` Retrieves multiple nodes by ID, returning results in input order with `undefined` for missing IDs. Costs one statement per bind-limit chunk where the backend exposes a batch read; where it does not, it falls back to one lookup per distinct id, issued concurrently. ```typescript store.nodes.Person.getByIds( ids: readonly NodeId[], options?: QueryOptions ): Promise | undefined)[]>; ``` When the backend supports batch lookups (`getNodes`), this executes `SELECT ... WHERE id IN (...)` once per bind-limit chunk — a single statement for id counts under the limit. Otherwise it falls back to one lookup per distinct id, issued concurrently rather than sequentially. ```typescript const [alice, bob, unknown] = await store.nodes.Person.getByIds([ aliceId, bobId, "nonexistent", ]); // alice: Node // bob: Node // unknown: undefined ``` #### `update(id, props, options?)` Updates node properties. ```typescript store.nodes.Person.update( id: NodeId, props: Partial<{ name: string; email?: string }>, options?: { validTo?: string } | { clearValidTo: true }, ): Promise>; ``` #### `compareAndSet(id, params)` Applies a patch only while the current live row still has the supplied exact property values. The id, expected values, and update execute as one set-based mutation, so this closes the race left by a separate `getById()` followed by `update()`. ```typescript const reopened = await store.nodes.ChangeSet.compareAndSet(changeSetId, { expected: { tenantId, projectId, status: "adopted", }, patch: { status: "proposed" }, }); const claimedUnassigned = await store.nodes.ChangeSet.compareAndSet(changeSetId, { expected: { assigneeId: compareAndSetAbsent }, patch: { assigneeId }, }); if (!reopened) { // Missing row or stale/mismatched precondition; nothing was written. } ``` Expected values are JSON scalars (`string`, `number`, `boolean`, or `null`). Use the exported `compareAndSetAbsent` marker to require that an optional property is not stored. `undefined` is refused rather than treated as absence, because it can disappear while an object is assembled or serialized; arrays and objects are also refused so every backend uses the same exact scalar comparison. Expected predicates are re-checked directly on the target row by the outer `UPDATE`, so a concurrent writer cannot satisfy the guard and then change the row before the patch lands. The complete after-image still goes through ordinary schema validation, uniqueness, history/version bookkeeping, fulltext, and vector maintenance. This makes `compareAndSet()` suitable for narrowly authorized exceptional recovery without broadening a schema's normal transition map. It requires the same transactional set-update backend support as `updateWhere()`. #### `updateWhere(params)` Updates a set of current, live nodes in one transactional operation and returns the number of rows changed. A selector is mandatory: provide `candidates`, `where`, one or more independent `exists` relationship predicates, or the explicit `all: true` acknowledgement. `candidates` accepts a query created by the same Store (or transaction) that selects one concrete node kind; its root node ids are intersected with any other selectors in the same atomic write. Candidate queries must contain concrete predicate values and select rows directly. TypeGraph refuses `param()` references because `updateWhere()` has no binding argument, and refuses `groupBy()` / `having()` because replacing an aggregate projection with root node ids would change the query's grouping semantics. Queries created by `withCheckedReads()` are also refused: embedding one in a mutation would bypass the checked read's schema-version fence. ```typescript const result = await store.nodes.Person.updateWhere({ patch: { active: false }, where: (person) => person.lastSeen.lt(cutoff), exists: [ { edgeKind: "worksAt", direction: "out", relatedKind: "Company", whereRelated: (company) => company.field("status").string().eq("closed"), }, ], }); // { affectedCount: number } const candidates = store .query() .from("Person", "person") .whereNode("person", (person) => person.age.gte(18)) .select((context) => context.person.id); await store.nodes.Person.updateWhere({ candidates, patch: { active: true }, }); ``` Each `exists` entry is evaluated independently and ANDed with the other selectors. The patch is shallow: values replace top-level properties, explicit `null` is preserved, and `undefined` removes an optional property. TypeGraph validates every complete after-image and updates uniqueness, fulltext, vector, history, and revision state in the same transaction. If any row or sidecar is invalid, the whole update rolls back. Backends without transactional set-write and batched sidecar support reject the operation before writing. `updateWhere()` is current-state only and is intentionally unavailable on a `StoreView`. Use `all: true` for an intentional whole-kind update: ```typescript await store.nodes.Person.updateWhere({ patch: { needsBackfill: false }, all: true, }); ``` #### `delete(id)` Soft-deletes a node. ```typescript store.nodes.Person.delete(id: NodeId): Promise; ``` #### `hardDelete(id)` Permanently deletes a node. This is irreversible and should be used carefully. ```typescript store.nodes.Person.hardDelete(id: NodeId): Promise; ``` #### `find(filter?, temporal?)` Finds nodes of this kind with optional filtering and pagination. The temporal coordinate is a separate second argument (`temporalMode` / `asOf`), so the filter object never mixes filtering with temporal scope. ```typescript store.nodes.Person.find( filter?: { where?: (accessor) => Predicate; limit?: number; offset?: number; }, temporal?: { temporalMode?: TemporalMode; asOf?: string }, ): Promise[]>; ``` The optional `where` predicate uses the same accessor API as `whereNode()` in the query builder: ```typescript const activeUsers = await store.nodes.Person.find({ where: (p) => p.status.eq("active"), limit: 50, }); // Pass the temporal coordinate as the second argument. const asOfLastYear = await store.nodes.Person.find( { where: (p) => p.status.eq("active") }, { temporalMode: "asOf", asOf: "2024-01-01T00:00:00.000Z" }, ); ``` #### `count(temporal?)` Counts nodes of this kind (excluding soft-deleted nodes). Accepts the same optional temporal coordinate as `find`. ```typescript store.nodes.Person.count(temporal?: { temporalMode?: TemporalMode; asOf?: string; }): Promise; ``` #### `createFromRecord(data, options?)` Creates a node from untyped data, relying on runtime Zod validation. Use this for dynamic dispatch (changesets, migrations, imports) where the data shape is determined at runtime, not compile time. The return type is fully typed — only the input gate is relaxed. ```typescript store.nodes.Person.createFromRecord( data: Record, options?: { id?: string; validFrom?: string | null; validTo?: string } ): Promise>; ``` ```typescript // Data arrives from an external source at runtime const importedRow: Record = JSON.parse(line); const person = await store.nodes.Person.createFromRecord(importedRow); // person is fully typed as Node ``` #### `upsertById(id, props, options?)` Creates or updates a node by ID. ```typescript store.nodes.Person.upsertById( id: string, props: { name: string; email?: string }, options?: { validFrom?: string | null; validTo?: string; clearValidTo?: true; onImmutableLowerBound?: "refuse" | "preserve"; } ): Promise>; ``` **Behavior:** - Creates a new node if no node with the ID exists - Updates the existing node if one exists - Un-deletes soft-deleted nodes (clears `deletedAt`) **Coalescing unchanged upserts.** When the store is created with [`coalesceUnchangedUpserts: true`](#createstoregraph-backend-options), an upsert whose validated props are value-identical to the existing **live** row performs **no write at all** — no update, no recorded history row, no revision-anchor advance, and no `update` operation hooks — and resolves with the existing node (its original `validFrom` / `updatedAt` / `version`). Enable it for at-least-once / replay materializers, where a byte-identical re-delivery would otherwise rewrite every row and grow recorded history by one per delivery. A write still happens (never coalesced) when the row is soft-deleted (an upsert resurrects it), when an explicit `validFrom` / `validTo` MOVES the window the row already holds, or when any prop differs after Zod normalization. Re-stating the window a row already holds — including clearing an already-open end — coalesces like any other unchanged value after backend capability validation, and a `validFrom` naming a bound a live row does not hold is refused rather than written (see [Immutable validity lower bounds](/errors/#immutable_validity_lower_bound)) — with this option on or off, because coalescing must not decide whether an unappliable or malformed bound is reported. The default is off, because some consumers want an audit row per re-delivery as proof the event was reprocessed. In a receipt, a coalesced upsert still counts as one write intent (`writes.total`) but captures nothing (`recorded` stays `undefined`) — the same shape as a no-op delete. **Create-only event time.** The default `onImmutableLowerBound: "refuse"` treats `validFrom` as an assertion on every branch. Event projectors that carry the source start on every revision can use `"preserve"`: create and resurrection still validate and store `validFrom`, while a live-row update preserves the bound already stored and applies the new props and `validTo`. This is an explicit branch policy, not a silent drop, and avoids catching `IMMUTABLE_VALIDITY_LOWER_BOUND` as normal control flow. #### `upsertByIdFromRecord(id, data, options?)` Upserts a node from untyped data, relying on runtime Zod validation. Same behavior as `upsertById` but accepts `Record` instead of the typed schema input. ```typescript store.nodes.Person.upsertByIdFromRecord( id: string, data: Record, options?: { validFrom?: string | null; validTo?: string; clearValidTo?: true; onImmutableLowerBound?: "refuse" | "preserve"; } ): Promise>; ``` ```typescript // Pre-seeded ID with dynamic data from a changeset const run = await store.nodes.Run.upsertByIdFromRecord( prepared.runId, { status: "running", ...dynamicConfig }, ); ``` #### `bulkCreate(items)` Creates multiple nodes efficiently. Uses a single multi-row INSERT when the backend supports it. ```typescript store.nodes.Person.bulkCreate( items: readonly { props: { name: string; email?: string }; id?: string; validFrom?: string | null; validTo?: string; }[] ): Promise[]>; ``` Use `bulkInsert` when you don't need the created nodes back: ```typescript await store.nodes.Person.bulkInsert(batch); ``` #### `bulkInsert(items)` Inserts multiple nodes without returning results. This is the dedicated fast path for bulk ingestion — wrapped in a transaction when the backend supports it. ```typescript store.nodes.Person.bulkInsert( items: readonly { props: { name: string; email?: string }; id?: string; validFrom?: string | null; validTo?: string; }[] ): Promise; ``` #### `bulkUpsertById(items)` Creates or updates multiple nodes by ID. ```typescript store.nodes.Person.bulkUpsertById( items: readonly { id: string; props: { name: string; email?: string }; validFrom?: string | null; validTo?: string; clearValidTo?: true; onImmutableLowerBound?: "refuse" | "preserve"; }[] ): Promise[]>; ``` With [`coalesceUnchangedUpserts: true`](#createstoregraph-backend-options) the dirty-check is applied per item: value-identical items are skipped from the write batch but still appear in the returned array (the existing node, in input order). See [`upsertById`](#upsertbyidid-props-options). #### `bulkReplaceById(items)` Creates or completely replaces multiple node documents by distinct IDs: ```typescript store.nodes.Person.bulkReplaceById( items: readonly { id: string; props: { name: string; email?: string }; }[] ): Promise[]>; ``` Unlike `bulkUpsertById`, replacement does not merge with stored properties. An omitted optional field is removed. Missing IDs are created, live rows keep their stored validity window, and tombstones are resurrected with a freshly stamped validity window. The method deliberately accepts no temporal mutation options: its complete postimage is knowable before dispatch, which lets eligible serverless backends execute the call without a preimage read. Duplicate IDs are refused rather than interpreted as an ordered script. #### `bulkDelete(ids)` Soft-deletes multiple nodes. ```typescript store.nodes.Person.bulkDelete( ids: readonly NodeId[] ): Promise; ``` On eligible bundled roots, a restricted node kind runs this call as one schema-fenced atomic exchange, including owner-side uniqueness and disjointness claim cleanup. Kinds that owe search, identity, or capture cleanup, or cascade/disconnect work, use the transactional path instead. #### `getOrCreateByConstraint(constraintName, props, options?)` Looks up an existing node by a named uniqueness constraint. Returns the match if found, or creates a new node if not. ```typescript store.nodes.Person.getOrCreateByConstraint( constraintName: string, props: { name: string; email?: string }, options?: { ifExists?: "return" | "update" } // Default: "return" ): Promise<{ node: Node; action: "created" | "found" | "updated" | "resurrected"; }>; ``` #### `bulkGetOrCreateByConstraint(constraintName, items, options?)` Batch version of `getOrCreateByConstraint`. Returns results in input order. ```typescript store.nodes.Person.bulkGetOrCreateByConstraint( constraintName: string, items: readonly { props: { name: string; email?: string }; }[], options?: { ifExists?: "return" | "update" } ): Promise< { node: Node; action: "created" | "found" | "updated" | "resurrected"; }[] >; ``` #### `findByConstraint(constraintName, props)` Looks up a node by a named uniqueness constraint without creating. Returns the matching node or `undefined`. Soft-deleted nodes are excluded. ```typescript store.nodes.Person.findByConstraint( constraintName: string, props: { name: string; email?: string } ): Promise | undefined>; ``` ```typescript const alice = await store.nodes.Person.findByConstraint("email", { email: "alice@example.com", name: "Alice", }); if (alice) { console.log(alice.id, alice.name); } ``` Throws `NodeConstraintNotFoundError` if the constraint name is not defined on the node type. #### `bulkFindByConstraint(constraintName, items)` Batch version of `findByConstraint`. Returns results in input order, with `undefined` for non-matches. Deduplicates within-batch lookups automatically. ```typescript store.nodes.Person.bulkFindByConstraint( constraintName: string, items: readonly { props: { name: string; email?: string } }[] ): Promise<(Node | undefined)[]>; ``` ```typescript const results = await store.nodes.Person.bulkFindByConstraint("email", [ { props: { email: "alice@example.com", name: "Alice" } }, { props: { email: "nobody@example.com", name: "Nobody" } }, { props: { email: "bob@example.com", name: "Bob" } }, ]); // results[0]: Node (Alice) // results[1]: undefined // results[2]: Node (Bob) ``` #### `bulkFindByIndex(indexName, items, options?)` Batched candidate retrieval against a **declared node index** (from [`defineNodeIndex`](/performance/indexes)). For each input record, returns the live nodes that share its declared index key. Unlike `bulkFindByConstraint`, the index may be **non-unique**, so each input yields a (possibly empty) array rather than a single optional node — this is candidate discovery (import reconciliation, dedup candidates, joining records by a composite key), not a uniqueness guarantee. For unique lookups prefer `bulkFindByConstraint`. ```typescript store.nodes.Person.bulkFindByIndex( indexName: string, items: readonly { props: Partial<{ name: string; email?: string }> }[], options?: { limitPerInput?: number } ): Promise[][]>; ``` ```typescript // Index: defineNodeIndex(Person, { name: "by_tenant", fields: ["tenantId"] }) const candidates = await store.nodes.Person.bulkFindByIndex("by_tenant", [ { props: { tenantId: "t1" } }, { props: { tenantId: "t2" } }, ]); // candidates[0]: Node[] (everyone in t1) // candidates[1]: Node[] (everyone in t2) ``` Semantics: one bucket per input in input order (empty input → `[]`); live, non-soft-deleted nodes only; buckets ordered by node id; only `index.fields` are used (not `coveringFields` or `keySystemColumns`), with the index's partial `where` applied to stored rows. A missing/`undefined` indexed field matches stored `NULL`. - `options.limitPerInput` caps each bucket (positive integer); unbounded by default. On backends without SQL window functions (`capabilities.windowFunctions: false`) the cap is applied in memory rather than via `ROW_NUMBER()` — same result. - Throws `NodeIndexNotFoundError` for an unknown index, `ConfigurationError` for an index declared without `fields` (only `coveringFields` and/or `keySystemColumns` — nothing to probe by) or for a date-typed key field (which can't compare identically across SQLite and PostgreSQL), and `ValidationError` for a non-positive `limitPerInput` or a non-scalar probe value. See [Index-backed lookup](/performance/indexes#batched-index-lookup-bulkfindbyindex) for details. ### Edge Collections Each edge type has a type-safe collection. The `from` and `to` parameters are constrained to only accept node types declared in the edge registration. #### `create(from, to, props)` Creates an edge. TypeScript enforces valid endpoint types. ```typescript // Given: worksAt: { type: worksAt, from: [Person], to: [Company] } store.edges.worksAt.create( from: NodeRef, to: NodeRef, props: { role: string } ): Promise>; // Preferred: Pass node objects directly await store.edges.worksAt.create(alice, acme, { role: "Engineer" }); // Compile error - Company is not a valid 'from' type await store.edges.worksAt.create(acme, alice, { role: "Engineer" }); ``` #### Node References Both forms are **exactly equivalent**—TypeGraph extracts `kind` and `id` from either: ```typescript // Full node object (preferred - cleaner syntax) await store.edges.worksAt.create(alice, acme, { role: "Engineer" }); // Explicit reference (useful when you only have IDs) await store.edges.worksAt.create( { kind: "Person", id: aliceId }, { kind: "Company", id: acmeId }, { role: "Engineer" } ); ``` Use the explicit `{ kind, id }` form when you have IDs but not the full node objects (e.g., from a previous query or external input). #### `getById(id)` Retrieves an edge by ID. ```typescript store.edges.worksAt.getById(id: EdgeId): Promise | undefined>; ``` When a persisted id crosses an untyped boundary, brand it before passing it to read/update/delete APIs: ```typescript const id = asEdgeId(row.edgeId); const edge = await store.edges.worksAt.getById(id); ``` Edge write APIs that mint ids still accept plain strings. #### `getByIds(ids)` Retrieves multiple edges by ID, returning results in input order with `undefined` for missing IDs. Costs one statement per bind-limit chunk where the backend exposes a batch read (`getEdges`); where it does not, it falls back to one lookup per distinct id, issued concurrently. ```typescript store.edges.worksAt.getByIds( ids: readonly EdgeId[], options?: QueryOptions ): Promise | undefined)[]>; ``` ```typescript const [edge1, edge2] = await store.edges.worksAt.getByIds([id1, id2]); ``` #### `update(id, props, options?)` Updates edge properties. ```typescript store.edges.worksAt.update( id: EdgeId, props: Partial<{ role: string }>, options?: { validTo?: string } | { clearValidTo: true } ): Promise>; ``` #### `findFrom(from, options?)` Finds edges from a node. Honors the same temporal model as `getById` / `find`: with no `options`, the graph's default `temporalMode` applies (so under the default `"current"` mode, edges outside their `validFrom` / `validTo` window are excluded). Pass `temporalMode` / `asOf` to read the endpoint's edges at another coordinate — e.g. `{ temporalMode: "includeEnded" }` for every non-deleted edge. ```typescript store.edges.worksAt.findFrom( from: NodeRef, options?: { temporalMode?: TemporalMode; asOf?: string } ): Promise[]>; ``` #### `findTo(to, options?)` Finds edges to a node. Temporal semantics mirror `findFrom`. ```typescript store.edges.worksAt.findTo( to: NodeRef, options?: { temporalMode?: TemporalMode; asOf?: string } ): Promise[]>; ``` #### `bulkFindFrom(froms, options?)` / `bulkFindTo(tos, options?)` Finds the edges of a **set** of endpoints in one read. This is `findFrom` / `findTo` with the endpoint predicate widened from `from_id = ?` to `from_id IN (...)` — the same index prefix seek, the same temporal model, the same per-endpoint ordering — so a page of N nodes costs one statement per endpoint kind and bind-budget chunk instead of N singleton statements. Results are grouped per input: index `i` of the returned array holds the edges of `froms[i]`, an endpoint with no edges gets an empty array, and repeated inputs each get their own copy. Pass `limitPerInput` to bound each endpoint's fan-out; it keeps the leading edges of that endpoint's `findFrom` order. Large inputs are transparently split across statements to respect the backend's bound-parameter budget. Requires a backend that implements the `findEdgesByEndpointSet` operation — both bundled Drizzle backends do. On a custom backend without it, these methods throw a `ConfigurationError` instead of falling back to one `findFrom` per input: a caller reaching for a bulk read is asking for a set-oriented read, so quietly issuing N singleton statements would be the cost surprise the method exists to remove. Loop over `findFrom` / `findTo` yourself if that trade is fine. ```typescript store.edges.worksAt.bulkFindFrom( froms: readonly NodeRef[], options?: { temporalMode?: TemporalMode; asOf?: string; limitPerInput?: number } ): Promise[][]>; ``` ```typescript const people = await store.nodes.Person.find({ limit: 50 }); const jobsPerPerson = await store.edges.worksAt.bulkFindFrom(people); // jobsPerPerson[i] holds the worksAt edges of people[i] ``` #### `store.bulkFindEdgesFrom(params, options?)` Reads multiple edge kinds from heterogeneous source kinds through one set-oriented backend operation. The result preserves source-group and ID order, includes an empty `edges` array for a source with no matches, and preserves repeated source references as separate result entries. The bundled SQLite and PostgreSQL backends issue one statement per bind-budget chunk, independent of the number of licensed `(source kind, edge kind)` combinations. `limitPerInput` bounds each source's fan-out. Temporal options have the same meaning as edge-collection reads, and `StoreView.bulkFindEdgesFrom` supplies its pinned coordinate automatically. ```typescript const results = await store.bulkFindEdgesFrom( { sources: [ { kind: "Company", ids: companyIds }, { kind: "Person", ids: personIds }, ], edgeKinds: ["employs", "owns", "dependsOn"], }, { limitPerInput: 20 }, ); for (const { source, edges } of results) { console.log(source.kind, source.id, edges.length); } ``` The operation validates every dynamic kind against the Store's graph. It requires a backend that implements `findEdgesByHeterogeneousEndpointSet`; a custom backend without that capability gets a `ConfigurationError` instead of an implicit loop of singleton reads. #### `store.bulkFindEdgesTo(params, options?)` The reverse-direction mirror accepts `targets` and `edgeKinds`, returning `{ target, edges }` buckets in input order. It shares the forward method's chunking, empty/repeated bucket behavior, `limitPerInput`, capability checks, and temporal options. `StoreView.bulkFindEdgesTo` uses the view's pinned coordinate. ```typescript const inbound = await store.bulkFindEdgesTo({ targets: [{ kind: "Document", ids: documentIds }], edgeKinds: ["owns", "references"], }); ``` For a single edge kind, `store.edges.references.bulkFindTo(targets)` already provides a set-oriented reverse read. Neither API needs `store.batch()`. #### `batchFindFrom(from, options?)` / `batchFindTo(to, options?)` / `batchFindByEndpoints(from, to, options?)` Deferred variants of `findFrom`, `findTo`, and `findByEndpoints` for use with [`store.batch()`](#batch-query-execution). These return a `BatchableQuery` instead of executing immediately. `batchFindFrom` / `batchFindTo` accept the same temporal `options` as `findFrom` / `findTo`. Batching them does not merge the reads — each still costs its own statement, and they share a connection only when the backend supports transactions. It is not a snapshot: PostgreSQL's default read-committed isolation lets a later read observe a commit the earlier ones did not. To read edges for many sources in one statement, traverse from them in a single query. ```typescript store.edges.worksAt.batchFindFrom( from: NodeRef, options?: { temporalMode?: TemporalMode; asOf?: string } ): BatchableQuery>; store.edges.worksAt.batchFindTo( to: NodeRef, options?: { temporalMode?: TemporalMode; asOf?: string } ): BatchableQuery>; store.edges.worksAt.batchFindByEndpoints( from: NodeRef, to: NodeRef, options?: { matchOn?: readonly string[]; props?: Partial<{ role: string }> } ): BatchableQuery>; ``` ```typescript // Execute multiple edge lookups in sequence — one statement each const [skills, employer] = await store.batch( store.edges.hasSkill.batchFindFrom(alice), store.edges.worksAt.batchFindFrom(alice), ); ``` `batchFindByEndpoints` returns a 0-or-1 element array (matching the at-most-one semantics of `findByEndpoints`). #### `find(filter?, temporal?)` Finds edges with endpoint filtering. The temporal coordinate is a separate second argument, mirroring `store.nodes..find`. ```typescript store.edges.worksAt.find( filter?: { from?: NodeRef; to?: NodeRef; limit?: number; offset?: number; }, temporal?: { temporalMode?: TemporalMode; asOf?: string }, ): Promise[]>; ``` For edge property filters, use the query builder with `whereEdge(...)`. #### `count(filter?, temporal?)` Counts edges matching filters. ```typescript store.edges.worksAt.count( filter?: { from?: NodeRef; to?: NodeRef; }, temporal?: { temporalMode?: TemporalMode; asOf?: string }, ): Promise; ``` #### `delete(id)` Soft-deletes an edge. ```typescript store.edges.worksAt.delete(id: EdgeId): Promise; ``` #### `hardDelete(id)` Permanently deletes an edge. This is irreversible and should be used carefully. ```typescript store.edges.worksAt.hardDelete(id: EdgeId): Promise; ``` #### `bulkCreate(items)` Creates multiple edges efficiently. Uses a single multi-row INSERT when the backend supports it. ```typescript store.edges.worksAt.bulkCreate( items: readonly { from: NodeRef; to: NodeRef; props?: { role: string }; id?: string; validFrom?: string | null; validTo?: string; }[] ): Promise[]>; ``` Use `bulkInsert` for high-volume edge ingestion when you do not need returned payloads: ```typescript await store.edges.worksAt.bulkInsert(edgeBatch); ``` #### `bulkInsert(items)` Inserts multiple edges without returning results. This is the dedicated fast path for bulk ingestion — wrapped in a transaction when the backend supports it. ```typescript store.edges.worksAt.bulkInsert( items: readonly { from: NodeRef; to: NodeRef; props?: { role: string }; id?: string; validFrom?: string | null; validTo?: string; }[] ): Promise; ``` #### `bulkDelete(ids)` Soft-deletes multiple edges. ```typescript store.edges.worksAt.bulkDelete( ids: readonly EdgeId[] ): Promise; ``` On eligible bundled roots, the whole call is one schema-fenced atomic exchange. An ID belonging to another edge kind refuses the call and rolls back every chunk. Portable backends batch both the authoritative lookup and soft delete when their optional batch ports are available. #### `bulkUpsertById(items)` Creates or updates multiple edges by ID. ```typescript store.edges.worksAt.bulkUpsertById( items: readonly { id: EdgeId; from: NodeRef; to: NodeRef; props?: { role: string }; validFrom?: string | null; validTo?: string; clearValidTo?: true; }[] ): Promise[]>; ``` #### `getOrCreateByEndpoints(from, to, props, options?)` Looks up an existing edge by endpoints (and optionally by property fields via `matchOn`). Returns the match if found, or creates a new edge if not. Declare the durable identity in the graph registration. An empty `fields` array means directed endpoints only; otherwise the named top-level persisted properties join the endpoint key. ```typescript edges: { worksAt: { type: worksAt, from: [Person], to: [Company], matchIdentity: { name: "employment", fields: ["role"] }, }, } ``` When this declaration exists, omitting `matchOn` uses its fields. A supplied `matchOn` must name exactly the same field set or the call is refused; it can never silently select a different identity. The identity fields are immutable through ordinary updates, soft deletion retains the key for deterministic resurrection, and hard deletion releases it. Direct creates and import paths materialize the same key, so they cannot bypass endpoint convergence. The complete indexed identity tuple is limited to 2,000 UTF-8 bytes on every backend. Larger identities refuse with `EDGE_MATCH_IDENTITY_KEY_TOO_LARGE` before writing, rather than succeeding on SQLite and later exceeding PostgreSQL's btree tuple limit. Durable identities must use compact JSON-scalar fields: strings, finite numbers, booleans, literals, enums, and nullable/optional/readonly unions of those types. Schema defaults, prefaults, and catch values are accepted when their wrapped output type stays inside that grammar. Transforms, pipes, and codecs are refused because their runtime result cannot be proven portable. `z.date()`, objects, arrays, maps, and sets are refused for the same reason. Long scalar payloads refuse per row during import, so one malformed edge does not roll back unrelated rows. Adding, removing, or changing the declaration is a breaking schema change. The first release refuses that migration while the edge kind holds rows; export and hard-delete the rows, migrate, then import them to materialize the new key. It never activates a declaration over legacy rows with `NULL` keys. ```typescript store.edges.worksAt.getOrCreateByEndpoints( from: NodeRef, to: NodeRef, props: { role: string }, options?: { matchOn?: readonly ("role")[]; // Default: [] ifExists?: "return" | "update"; // Default: "return" validFrom?: string | null; validTo?: string; clearValidTo?: true; onImmutableLowerBound?: "refuse" | "preserve"; // Default: "refuse" } ): Promise<{ edge: Edge; action: "created" | "found" | "updated" | "resurrected"; }>; ``` `validFrom` applies when the operation creates or resurrects the edge. On an `"updated"` live match, `onImmutableLowerBound: "preserve"` treats it as create/resurrection-only input: the stored start remains unchanged while props and `validTo` are applied. The default `"refuse"` policy instead refuses a `validFrom` naming a different instant; restating the bound the edge already holds is accepted. See [Immutable validity lower bounds](/errors/#immutable_validity_lower_bound). On a resurrection, naming `validFrom` asserts the COMPLETE window: an accompanying `validTo` is applied, and an omitted one reopens the revived row rather than keeping the tombstoned incarnation's end. `validTo` applies when the edge is created, updated, or resurrected, and may not precede the row's effective start — see [Inverted validity windows](/errors/#inverted_validity_window). When `ifExists` is omitted or `"return"`, a live match produces the `"found"` action and neither temporal option changes the edge. The bundled one-statement arbiter implements a contended/found result through the database conflict target. PostgreSQL and SQLite consequently perform a no-op physical update on this path: it can acquire a row lock and produce write amplification even though the logical edge is unchanged. This is the trade-off that keeps a cache-safe create/found verdict to one database request; do not treat `getOrCreateByEndpoints` as a read primitive on a hot identity. Inside a caller-owned PostgreSQL `repeatable_read` or `serializable` transaction, contention can instead abort that whole transaction with SQLSTATE `40001`. Retry the complete caller transaction; TypeGraph cannot safely replay only a nested callback whose surrounding relational work it does not own. At any isolation level, two transactions converging on durable identities in opposite orders can deadlock on their incumbent row locks and PostgreSQL can abort one with SQLSTATE `40P01`. Apply the same whole-transaction retry policy. `clearValidTo` is the exception: a live match can apply it only under `ifExists: "update"`. Supplying it with the default/`"return"` mode refuses with `ConfigurationError` code `CLEAR_VALID_TO_REQUIRES_UPDATE` instead of silently returning an ended edge. A create or tombstone resurrection can still apply the clear request. With `coalesceUnchangedUpserts: true`, an `ifExists: "update"` match whose validated props and requested validity bounds already equal the live edge is a no-op. It returns the existing edge with action `"found"`; action `"updated"` therefore always means an UPDATE ran. The same rule applies per item to `bulkGetOrCreateByEndpoints`. Confirming that no-op requires the endpoint match-key convergence fence; a top-level backend without transactions refuses it with `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` rather than eliding the write from an unfenced read. #### `bulkGetOrCreateByEndpoints(items, options?)` Batch version of `getOrCreateByEndpoints`. Returns results in input order. The bundled backends read all candidate endpoint pairs with set-oriented statements rather than one lookup per item. These are exact directed-pair joins, so a high-fan-out source does not materialize all of its unrelated outgoing edges for client-side filtering. For a schema-declared durable `matchIdentity` with `cardinality: "many"`, the declaration's match fields, default `ifExists: "return"`, and no temporal mutation, bundled roots use one closed native atomic exchange. That program owns endpoint validation, durable identity arbitration and ordered live results, so it does not perform an outside probe or open a separate interactive transaction. Its authoritative upsert establishes an all-live `"found"` result in the same exchange, but can take incumbent-row locks and produce write amplification. An all-live default-`"return"` batch outside that envelope is read-only and returns from one set-oriented root read. Other outcomes retain the set-oriented read plus transactional write path. Tombstoned winners roll the native attempt back and refuse transactionless convergence with the typed `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` (`edgeMatchKeyConvergence`) error; use a transaction-capable backend when resurrection must preserve schema-aware partial updates. ```typescript store.edges.worksAt.bulkGetOrCreateByEndpoints( items: readonly { from: NodeRef; to: NodeRef; props: { role: string }; validFrom?: string | null; validTo?: string; clearValidTo?: true; onImmutableLowerBound?: "refuse" | "preserve"; }[], options?: { matchOn?: readonly ("role")[]; ifExists?: "return" | "update"; } ): Promise< { edge: Edge; action: "created" | "found" | "updated" | "resurrected"; }[] >; ``` Temporal fields and `onImmutableLowerBound` belong to each item so identities with different endpoints or `matchOn` values can carry independent validity windows and lower-bound policies in one batch. Items with the same endpoint-plus-`matchOn` identity are duplicates: the first item supplies the write values and later items return that edge with the `"found"` action. To represent multiple periods between the same endpoints, add a stable period or source-event field to the edge schema and include it in `matchOn`. Their create, update, and resurrection semantics otherwise match the single operation. #### `findByEndpoints(from, to, options?, temporal?)` Looks up an edge by its endpoints without creating. Returns the matching edge or `undefined`. Honors the same temporal model as `findFrom` / `findTo`: with no `temporal` argument the graph's default `temporalMode` applies (so under the default `"current"` mode, edges outside their validity window are excluded). Pass `temporalMode` / `asOf` to look up the edge as of another coordinate. When `matchOn` is omitted, returns the first matching edge between the two endpoints. When `matchOn` is provided, filters by the specified property fields. ```typescript store.edges.knows.findByEndpoints( from: NodeRef, to: NodeRef, options?: { matchOn?: readonly ("relationship" | "since")[]; props?: Partial<{ relationship: string; since: string }>; }, temporal?: { temporalMode?: TemporalMode; asOf?: string }, ): Promise | undefined>; ``` ```typescript // Find any edge between Alice and Bob const edge = await store.edges.knows.findByEndpoints(alice, bob); // Find the specific "colleague" edge between Alice and Bob const colleague = await store.edges.knows.findByEndpoints(alice, bob, { matchOn: ["relationship"], props: { relationship: "colleague" }, }); ``` ### Transactions #### `store.transaction(fn)` Executes a callback within an atomic transaction. All operations succeed together or are rolled back together. The transaction context (`tx`) provides the same `nodes.*` and `edges.*` collection API as the store, plus transaction-bound `query()`, `neighbors()`, `countNeighbors()`, `subgraph()`, and `batchOnce()` reads. Those reads see writes made earlier in the callback. ```typescript await store.transaction(async (tx) => { const person = await tx.nodes.Person.create({ name: "Alice" }); const company = await tx.nodes.Company.create({ name: "Acme" }); await tx.edges.worksAt.create(person, company, { role: "Engineer" }); }); ``` #### Return values The callback's return value is forwarded to the caller: ```typescript const personId = await store.transaction(async (tx) => { const person = await tx.nodes.Person.create({ name: "Alice" }); return person.id; }); // personId is available here ``` #### Transaction receipts Use `store.transactionWithReceipt()` when a caller needs a write summary without wrapping the transaction context itself. It runs the callback exactly like `store.transaction()` and returns the result together with a receipt: ```typescript const outcome = await store.transactionWithReceipt(async (tx) => { const alice = await tx.nodes.Person.create({ name: "Alice" }); const bob = await tx.nodes.Person.create({ name: "Bob" }); await tx.edges.knows.getOrCreateByEndpoints(alice, bob, { since: "2026", }); return alice.id; }); outcome.result; // Alice's id outcome.receipt.writes; // { nodes: { Person: 2 }, edges: { knows: 1 }, total: 3 } outcome.receipt.recorded; // RecordedInstant | undefined ``` Receipt counts are completed write intents at the collection surface, not rows affected: - Every successful completion of a write method on `tx.nodes.*` / `tx.edges.*` counts. The authoritative method list is `NodeWrites` / `EdgeWrites`. - Bulk methods count by input length; an empty bulk call (`bulkCreate([])`) counts 0. - Single-row methods count 1 on resolve — including `delete` of an absent id and `getOrCreate*` that found an existing row. Consumers that need "did anything actually change" semantics apply their own per-operation policy. - A method that rejects counts 0 — even when the backend applied part of a bulk input before failing. On SQLite a failed statement does not abort the surrounding transaction, so a caller that catches the rejection and commits can persist rows the receipt never counted. Do not read the receipt as rows-affected in that scenario. - A node `delete` under `cascade` / `disconnect` removes connected edges through the backend, not the edge-collection surface; those removals do not appear in `edges`. - Rows-affected fidelity is intentionally out of scope for this first version; a future extension could ask backends to return row counts. When the store was created with `{ history: true }` and the transaction flushed captured writes or an explicit revision request, `receipt.recorded` is the recorded commit instant allocated for this store's graph by this transaction. It is `undefined` when history capture is off, the transaction is read-only, or no captured writes or revision request were flushed. Writes that bypass the transaction collection surface — direct backend writes, raw SQL, and import helpers — are not counted. `store.withRecordedTransaction()` — the adopted-commit path for history stores — returns the same `TransactionOutcome`, so the exactly-once cursor pattern gets a receipt too (see [Recorded time](/queries/temporal/#raw-sql-under-history-capture)); only `withTransaction`, whose commit belongs entirely to the caller with no flush point, produces no receipt. On a Store backed by a non-transactional driver, `transactionWithReceipt()` refuses before invoking the callback, so it cannot produce a receipt. Ordinary Store writes remain available when the application deliberately owns non-atomic coordination. History transaction contexts also expose `tx.requestRecordedRevision()`. Call it when the transaction must produce a replay anchor even if it changes no nodes, edges, or identity assertions. Requests are idempotent: repeated calls, or a request combined with ordinary graph writes, allocate one revision at the terminal capture flush. The instant is not available inside the callback; read it from the outer `receipt.recorded` after the callback completes. A thrown callback rolls the request back with the transaction. Engine-native history refuses this method because revision allocation belongs to the database engine, and an `accessMode: "read_only"` transaction refuses it because allocation is a write. Plain `store.transaction()` can request a revision, but returns no receipt; callers that must persist the exact anchor produced by their own transaction use `transactionWithReceipt()` or `withRecordedTransaction()`. ```typescript const outcome = await store.transactionWithReceipt(async (tx) => { tx.requestRecordedRevision(); }); const checkpoint = outcome.receipt.recorded; ``` ##### Scoped receipts: `tx.measure()` The context handed to `transactionWithReceipt` and `withRecordedTransaction` also exposes `tx.measure(fn)`. It runs `fn` with a **scoped context** — a second view over the same transaction — and returns a `TransactionOutcome` whose receipt counts exactly the writes made **through that scoped context** (`scoped.nodes` / `scoped.edges`). This lets a framework attribute writes to user code it invoked (for example, an event-log materializer measuring `project(scoped, change)` to detect a change that wrote nothing) while its own bookkeeping — written through the outer `tx` — stays out of that count: ```typescript await store.transactionWithReceipt(async (tx) => { const projected = await tx.measure((scoped) => project(scoped, change)); if (projected.receipt.writes.total === 0 && change.operation !== "delete") { throw new DroppedChangeError(change); // the projector dropped the change } await tx.nodes.Cursor.upsertById("s1", { offset: change.offset }); // outer tx — not in `projected` }); ``` Attribution is by **which context you write through**, not by timing. A write through the scoped context counts in both the scope and the outer receipt (it happened in the transaction); a write through the outer `tx` during the scope counts only in the outer receipt. This makes overlapping and concurrent measures safe by construction — two scopes racing under `Promise.all`, each writing through its own scoped context, never cross-count. Nesting composes: `scoped.measure(...)` opens a child scope that chains up through its ancestors. A scoped receipt's `recorded` is **always `undefined`** — the recorded instant is a per-transaction flush concern, unknowable mid-transaction. Plain `store.transaction()` contexts have no `measure` (no receipt is being produced). The scoped history context still exposes `requestRecordedRevision()`; its request belongs to the outer transaction, so only the outer receipt contains the instant. #### Rollback and error propagation If the callback throws, the transaction is rolled back and the error re-throws to the caller. No partial writes are persisted. ```typescript try { await store.transaction(async (tx) => { await tx.nodes.Person.create({ name: "Alice" }); throw new Error("something went wrong"); // Alice is NOT persisted — the entire transaction is rolled back }); } catch (error) { // error.message === "something went wrong" } ``` #### Retrying on conflict PostgreSQL can abort a transaction with a serialization failure or deadlock when two writers race — its own protocol response is to re-run the whole transaction from the top. By default `store.transaction()` and `store.transactionWithReceipt()` do this once: a conflict the backend reports surfaces as `TransactionConflictError` (code `TRANSACTION_CONFLICT`, with the driver error as `cause`) instead of the raw error, and `details.attempts` is `1`. Pass `retry: { attempts }` to have TypeGraph re-run the callback itself, up to `attempts` times total, whenever a conflict is detected: ```typescript import { TransactionConflictError } from "@nicia-ai/typegraph"; try { await store.transaction( async (tx) => { // Read inside the callback: a replayed attempt must see fresh balances. const from = await tx.nodes.Account.getById(fromId); const to = await tx.nodes.Account.getById(toId); if (from === undefined || to === undefined) throw new Error("missing account"); await tx.nodes.Account.compareAndSet(fromId, { expected: { balance: from.balance }, patch: { balance: from.balance - 10 }, }); await tx.nodes.Account.compareAndSet(toId, { expected: { balance: to.balance }, patch: { balance: to.balance + 10 }, }); }, { retry: { attempts: 3 } }, ); } catch (error) { if (error instanceof TransactionConflictError) { // every attempt conflicted; error.details.attempts === 3 } } ``` A retried callback re-runs unconditionally on conflict, so it must satisfy the **replay contract**: - **Await all of its own work** before returning or throwing. A retried callback that leaves a fire-and-forget effect in flight from a failed try could duplicate that effect, or let the caller observe it after the try that started it was rolled back. - **Read and write only values created fresh on each call.** A failed attempt's transaction rolled back, so anything it left in a variable outside the callback — a counter, a buffer, an accumulated list — describes state no committed database agrees with. Reading it on the next attempt lets a rolled back try leak into the one that commits. - **Perform no effect outside the transaction itself.** The whole callback re-runs on conflict, so a network call, a write to a different store, or any other side effect the transaction does not own runs again too. - **Tolerate being invoked up to `attempts` times**, not exactly once. Every hook fired by the callback's operations (`onOperationStart`, `onBulkOperationStart`, `onQueryStart`, and a failing operation's own `onError`) carries the 1-based attempt number that produced it, so a listener can tell a replay apart from a new operation. A rolled-back attempt's completed operations report neither `onOperationEnd` nor `onError` of their own — only the attempt that actually commits (or the last one, once `attempts` is exhausted) is reported, and `transactionWithReceipt`'s receipt reflects only that same committed attempt's writes. `retry` changes nothing about which backends `store.transaction()` accepts in the first place: a backend without interactive transactions (see [Backend support](#backend-support) below) already refuses the call before the callback ever runs, `retry` present or not. #### Nesting Transactions do **not** nest. The transaction context intentionally omits the `transaction()` method, so attempting to start a transaction inside another transaction is a compile-time error. If you need to compose transactional operations, pass the `tx` context through your call chain. #### Backend support Not all backends support atomic transactions. Cloudflare D1 and `drizzle-orm/neon-http` cannot hold a multi-statement session and report `capabilities.execution.interactiveTransactions: false`. A schema-managed Store fails closed before writing on these backends because it cannot hold the schema fence. On a raw Store, `store.transaction(fn)` refuses because no interactive transaction is available. Eligible operations backed by a certified atomic SQL program remain separate from this interactive capability. If you require atomicity or version fencing, branch on the capability: ```typescript if (store.capabilities.execution.interactiveTransactions) { await store.transaction(async (tx) => { /* atomic */ }); } else { // Use individual operations or an eligible certified atomic operation. } ``` See [Limitations](/limitations#backends-without-atomic-transactions) for the full list of affected backends and edge-runtime alternatives. ### Clear #### `store.clear()` Hard-deletes the current graph's data: nodes, edges, uniqueness entries, embeddings, and schema versions. Contribution materialization markers are preserved by default so a clear does not invalidate their attestations. Pass `preserveContributionMaterializations: false` to remove those graph-local markers during a full cutover purge. Resets collection caches so the store is immediately reusable. ```typescript store.clear(options?: { preserveContributionMaterializations?: boolean }): Promise; ``` Wrapped in a transaction when the backend supports it. Does not affect other graphs sharing the same backend. To verify what remains afterward without reading TypeGraph's physical tables, use [`inspectGraphStorage(store)`](/multiple-graphs#inspectgraphstoragestore). ```typescript // Wipe all data and start fresh await store.clear(); // Also remove this graph's contribution markers during a full cutover purge. await store.clear({ preserveContributionMaterializations: false }); // Store is immediately reusable, now with raw/unversioned semantics. const person = await store.nodes.Person.create({ name: "Alice" }); ``` Because `clear()` deletes the committed schema rows, it also resets a formerly managed Store to `introspect().schemaVersion === undefined`. Subsequent writes are raw and unfenced. Reopen the graph through a managed factory before writing when the schema-version guarantee is required. ### Batch Query Execution #### `store.batchOnce(buildReads, options?)` Executes zero or more independent reads. A nonempty input is exactly one SQL statement; an empty input returns `[]` without issuing SQL. Each read is embedded as a CTE, and one JSON envelope carries the independently typed result sets back in input order. This is the batch surface for latency-bound page assembly: dozens of independent reads still form one statement and one database round trip. Each query's explicit `.orderBy()` is preserved even when its sort fields are not part of the public projection. The callback may return a heterogeneous tuple, a singleton, or a readonly runtime array. This makes dynamic multi-root subgraph retrieval direct. Prefer this form over awaiting `store.subgraph()` in a loop when one request needs several independent bounded neighborhoods, especially against a remote database: ```typescript const subgraphs = await store.batchOnce((read) => roots.map((root) => read.subgraph(root.id, { edges: ["knows"], maxDepth: 2 }), ), ); ``` Repeated roots remain separate results. Missing roots produce their ordinary empty subgraph result. The portable planning limit is 500 reads, matching SQLite's compound-select ceiling. The complete statement must also fit the backend's declared bind-parameter limit; otherwise `batchOnce()` refuses before execution. It never chunks. Result rows for every member are materialized in JSON envelopes, so this API is intended for bounded reads and is not streaming. TypeGraph does not guess response size or impose a response-byte cap; use explicit limits, projections, and subgraph bounds to control it. Each member keeps its own root and options. A tuple can therefore combine unrelated neighborhood shapes in the same call: ```typescript const [social, work] = await store.batchOnce((read) => [ read.subgraph(person.id, { edges: ["knows"], maxDepth: 2, edgeWindows: { knows: { limit: 25 } }, }), read.subgraph(person.id, { edges: ["worksAt"], maxDepth: 1, project: { nodes: { Company: ["name", "industry"] }, edges: { worksAt: ["role"] }, }, }), ]); ``` One statement means one round trip, not one shared traversal by default. Overlapping subgraphs are planned and hydrated independently. For compatible, overlapping, payload-heavy subgraphs, pass `{ shareSubgraphs: true }` as the second `batchOnce()` argument to traverse the roots together and hydrate each shared entity once: ```typescript const details = await store.batchOnce( (read) => roots.map((root) => read.subgraph(root.id, options)), { shareSubgraphs: true }, ); ``` Every request still receives independent result objects. Sharing adds membership and reconstruction work, and measurements show it can increase encoded response size when roots do not overlap or the projection is small. Keep the independent default unless the request shape benefits in a benchmark. ```typescript const [people, neighbors] = await store.batchOnce((read) => [ store.query().from("Person", "p").select((ctx) => ctx.p), read.neighbors(person, { edges: ["knows"], orderBy: { by: "node", field: "name", direction: "asc" }, limit: 5, }), ]); ``` The same API is available on a transaction context. Both the fluent query and the callback's read builder execute through the already-open transaction, so they see writes made earlier in the callback while the combined batch remains exactly one statement: ```typescript await store.transaction(async (tx) => { const alice = await tx.nodes.Person.create({ name: "Alice" }); const [people, neighborCount] = await tx.batchOnce((read) => [ tx.query().from("Person", "p").select((ctx) => ctx.p), read.countNeighbors(alice, { edges: ["knows"] }), ]); }); ``` `batchOnce()` has no sequential fallback. Every read must belong to the same graph and execution target as the Store or transaction running the batch; cross-Store and cross-transaction rebinding is refused before SQL. Its callback returns fluent relational queries, set operations, and batch-scoped composable reads built through the callback's `read.neighbors()`, `read.countNeighbors()`, and `read.subgraph()` methods. Prepared queries and edge collection `batchFind*` values are excluded because they cannot be embedded without changing their execution contract. Use `batch()` when those queued collection reads or sequential execution are the goal. #### `store.batch(...queries)` Runs several independent queries in sequence and returns a typed tuple of results preserving input order — N query executions, never one round trip. Accepts two or more queries (from `.select()`, set operations, or edge collection `batchFind*` methods), each keeping its own projection, filtering, sorting, and pagination. **Cost.** At least one statement per query. Whole-node, whole-edge, and spread selections detected during planning use a full fetch from the start. A selector branch that depends on actual row values can still trigger a second statement: selective-field mapping falls back after its statement has executed, then re-runs as a full fetch. That fallback clears the fast path for the query instance, so reusing it avoids repeating the extra statement. With `backend.capabilities.execution.interactiveTransactions` the queries share one transaction; how that reaches the wire is the adapter's business. A SQL backend frames them with `begin`/`commit`, putting a networked one at N+2 round trips **at best**, while Durable Objects use an ambient storage transaction with no framing statements. Without transactions there is no framing. Connection reuse is a separate question from transaction support: the no-transaction path passes the same backend object, so an adapter may reuse one client there too (see [Limitations](/limitations)). The portable guarantee is only that at most one query is in flight at a time. **Not a snapshot by default.** PostgreSQL defaults to read-committed isolation, so a later query in the batch can observe a commit the earlier ones did not. For one stable snapshot, call `tx.query()` or the transaction's set-oriented reads inside `store.transaction(fn, { isolationLevel: "repeatable_read" })`. `tx.batchOnce()` is already one statement and preserves that physical shape inside the transaction. Transactions require a backend with interactive transaction support; a history-enabled store on PostgreSQL additionally requires `accessMode: "read_only"` for a read-only transaction. **Will not fix an N+1.** Serializing N queries does not reduce their number. The alternatives are set-oriented or chunked rather than fixed-cost: `.traverse()` compiles a whole chain to one statement; [`store.subgraph()`](#subgraph-extraction) costs 2 statements on SQLite and 3 on PostgreSQL however large the result; `getByIds()` issues one statement per bind-limit chunk, falling back to one per distinct id where the backend exposes no batch read; `bulkFindByIndex()` costs one probe plus that same chunked hydration. **Versus `Promise.all`.** Workload- and adapter-dependent in both directions. `Promise.all` overlaps its queries against a pool with idle capacity, but it does not necessarily hold N connections, and against a single client or a saturated pool it queues. `batch()` keeps at most one query in flight, so it pays the sum of their latencies — but it can still come out ahead where connection acquisition dominates. Measure rather than assume. ```typescript store.batch( q1: BatchableQuery, q2: BatchableQuery, ...qn: BatchableQuery, ): Promise; ``` **Example:** ```typescript const [people, companies] = await store.batch( store .query() .from("Person", "p") .whereNode("p", (p) => p.status.eq("active")) .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })), store .query() .from("Company", "c") .select((ctx) => ({ id: ctx.c.id, name: ctx.c.name })) .orderBy("c", "name", "asc") .limit(5), ); // people: readonly { id: string; name: string }[] // companies: readonly { id: string; name: string }[] ``` **With traversals and mixed projections:** ```typescript const [skills, artifacts, recentGoals] = await store.batch( store .query() .from("Agent", "a") .whereNode("a", (a) => a.id.eq(agentId)) .traverse("has_skill", "e") .to("Skill", "s") .select((ctx) => ({ id: ctx.s.id, name: ctx.s.name })), store .query() .from("Agent", "a") .whereNode("a", (a) => a.id.eq(agentId)) .traverse("references", "ref") .to("Artifact", "art") .select((ctx) => ({ id: ctx.art.id, title: ctx.art.title, pin: ctx.ref.activeVersionId, })), store .query() .from("Agent", "a") .whereNode("a", (a) => a.id.eq(agentId)) .traverse("has_goal", "e") .to("Goal", "g") .select((ctx) => ({ id: ctx.g.id, name: ctx.g.name })) .orderBy("g", "name", "asc") .limit(10), ); ``` **Set operations work too:** ```typescript const [combined, separate] = await store.batch( store .query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("admin")) .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })) .union( store .query() .from("Person", "p") .whereNode("p", (p) => p.role.eq("owner")) .select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })), ), store .query() .from("Company", "c") .select((ctx) => ({ id: ctx.c.id, name: ctx.c.name })), ); ``` **Edge collection lookups:** ```typescript // Edge batchFind* methods return BatchableQuery — mix freely with fluent queries const [skills, employer, colleague] = await store.batch( store.edges.hasSkill.batchFindFrom(alice), store.edges.worksAt.batchFindFrom(alice), store.edges.knows.batchFindByEndpoints(alice, bob), ); ``` #### When to use batch() vs alternatives | Pattern | Use | |---------|-----| | Independent embeddable reads that must use one statement | `store.batchOnce()` / `tx.batchOnce()` | | Mixed fluent and queued collection queries | `store.batch()` | | Load entity with all relationships (uniform) | `store.subgraph()` | | Fixing an N+1 / reducing round trips | `.traverse()` or `store.batchOnce()` (one statement), `store.subgraph()` (2–3), `getByIds()` (chunked) | | Single query | `.execute()` directly | | Writes interleaved with fluent or set-oriented reads | `store.transaction()` | | Same-shape queries merged into one result | `.union()` / `.intersect()` / `.except()` | :::note `batch()` is read-only. For bulk writes, use `bulkCreate`, `bulkInsert`, or wrap operations in a `store.transaction()`. ::: ### Neighbor Reads `store.neighbors(source, options)` joins visible relationships to their adjacent nodes in one statement. Ordering and limiting happen before hydration, so append-only relationship histories can fetch only their newest target. ```typescript const [latest] = await store.neighbors(document, { edges: ["hasVersion"], direction: "out", orderBy: { field: "createdAt", direction: "desc" }, limit: 1, }); console.log(latest?.edge, latest?.node); ``` Use `store.countNeighbors(source, options)` for the matching aggregate-only read. It counts visible relationships through the selected edge kinds without hydrating edge or node objects. Both methods accept `direction: "out" | "in" | "both"` and the Store's temporal read options. `neighbors()` can order by edge `id`, `createdAt`, `updatedAt`, `validFrom`, or `validTo`, with a positive integer `limit`. Null metadata sorts last on both dialects, and ties are resolved by edge ID so a bounded read is deterministic. Set `orderBy.by` to `"node"` to order by a schema-declared adjacent-node property instead. Omit it or use `"edge"` for edge metadata. Inside `batchOnce()`, use the callback's `read.neighbors()` and `read.countNeighbors()` methods to compose these shapes with other independent reads. ### Schema-Checked Read Scopes `store.withCheckedReads(expectedVersion, fn)` binds one expected active schema version to every fluent query created through the callback scope. Calling `.execute()` on those queries automatically uses the same statement-level schema check as `.executeChecked(expectedVersion)`. ```typescript const page = await store.withCheckedReads(schemaVersion, async (reads) => { const people = await reads.query().from("Person", "p").select((ctx) => ctx.p).execute(); const teams = await reads.query().from("Team", "t").select((ctx) => ctx.t).execute(); return { people, teams }; }); ``` A mismatch throws `SchemaChangedError` from the read block, giving the caller one boundary at which to reload the store/schema cache and retry the whole block. The scope currently accepts ordinary `.execute()` fluent selects only. Prepared queries, aggregates, set operations, pagination, streaming, and `batchOnce()` refuse with `ConfigurationError` rather than silently running without the schema check. ### Subgraph Extraction #### `store.subgraph(rootId, options)` Extracts a typed subgraph by performing a BFS traversal from a root node, following the specified edge kinds. Returns an indexed result with adjacency maps for immediate traversal. `maxDepth` defaults to 10 and accepts integers from 0 through 1000. Zero returns only the root (unless excluded). Values outside that range are rejected; they are never silently clamped. `direction` accepts `"out"` or `"both"`, and `cyclePolicy` accepts `"prevent"` or `"allow"`. Use `edgeWindows` to choose direction and cap an append-only edge kind per source at every traversal hop. The ranking is applied inside the recursive traversal and again during edge hydration, so omitted targets are not loaded and do not remain as orphan nodes. ```typescript const detail = await store.subgraph(document.id, { edges: ["hasVersion", "hasSection"], maxDepth: 3, edgeWindows: { hasVersion: { direction: "in", limit: 1, orderBy: { field: "createdAt", direction: "desc" }, }, }, }); ``` Each window accepts `direction: "out" | "in" | "both"`; when omitted it inherits the traversal's global direction. Ranking is partitioned by the oriented source endpoint, so bidirectional windows have an unambiguous top N for each endpoint. Inside `batchOnce()`, the callback's batch-scoped `read.subgraph()` method accepts the same options and produces the same result as `store.subgraph()`, but compiles hydration and traversal into one embeddable statement. Both forms share the same validation, traversal, projection plan, and result assembly; only their physical execution strategy differs. Under the hood the traversal is a `WITH RECURSIVE` CTE and all filtering and hydration happen in the database. Direct `subgraph()` calls use a backend-tuned fixed cost: 2 statements on SQLite (nodes and edges, each embedding the CTE) and 3 on PostgreSQL (the closure ids once, then nodes and edges). Batch-scoped `read.subgraph()` uses 1 statement on both backends. Prefer direct `store.subgraph()` unless the read must compose with other independent reads in `batchOnce()`; PostgreSQL can execute the split hydration plan substantially faster for larger closures even though it uses more round trips. ```typescript store.subgraph( rootId: NodeId>, options: SubgraphOptions, ): Promise>; ``` **Options:** | Option | Type | Default | Description | |--------|------|---------|-------------| | `edges` | `readonly EK[]` | *(required)* | Edge kinds to follow during traversal | | `maxDepth` | `number` | `10` | Integer traversal depth from root, from 0 through `MAX_EXPLICIT_RECURSIVE_DEPTH` (1000); larger values are rejected | | `includeKinds` | `readonly NK[]` | all kinds | Node kinds to include in the result. Other kinds are traversed through but omitted from output | | `excludeRoot` | `boolean` | `false` | Exclude the root node from the result | | `direction` | `"out" \| "both"` | `"out"` | `"out"` follows edges in their defined direction; `"both"` treats edges as undirected | | `cyclePolicy` | `"prevent" \| "allow"` | `"prevent"` | Whether to detect and skip cycles during traversal | | `temporalMode` | `TemporalMode` | `graph.defaults.temporalMode` | Filter applied to both nodes and edges along the traversal — same semantics as `store.query()` and collection reads | | `asOf` | `string` (ISO-8601) | *(none)* | Snapshot timestamp, required when `temporalMode: "asOf"` | | `project` | `{ nodes?, edges? }` | *(none)* | Per-kind field projection — see [Projection](#subgraph-projection) below | **Result:** ```typescript type SubgraphResult = Readonly<{ root: SubgraphNodeResult | undefined; nodes: ReadonlyMap>; adjacency: ReadonlyMap[]>>; reverseAdjacency: ReadonlyMap[]>>; }>; ``` | Field | Description | |-------|-------------| | `root` | The root node, or `undefined` if it was not found or `excludeRoot` is set | | `nodes` | All reachable nodes keyed by string ID | | `adjacency` | Forward adjacency: `fromId → edgeKind → edges[]` | | `reverseAdjacency` | Reverse adjacency: `toId → edgeKind → edges[]` | Edges are only included when **both** endpoints appear in the result set. Nodes and edges are filtered by the resolved `temporalMode` — by default, only currently valid rows participate. Duplicate nodes (reachable via multiple paths) are deduplicated. **Example:** ```typescript const sg = await store.subgraph(run.id, { edges: ["has_task", "runs_agent", "uses_skill"], maxDepth: 4, }); // Root node (the traversal starting point) console.log(sg.root?.kind); // Lookup by ID const task = sg.nodes.get(taskId); // Forward adjacency: edges of a kind from a node const taskEdges = sg.adjacency.get(String(run.id))?.get("has_task") ?? []; const tasks = taskEdges.map((edge) => sg.nodes.get(String(edge.toId))); // Reverse adjacency: edges of a kind pointing to a node const parentEdges = sg.reverseAdjacency.get(taskId)?.get("has_task") ?? []; // Narrow by kind with a switch for (const node of sg.nodes.values()) { switch (node.kind) { case "Task": { console.log(node.title, node.status); break; } case "Agent": { console.log(node.model); break; } } } ``` **Filtering to specific node kinds:** ```typescript const tasksOnly = await store.subgraph(run.id, { edges: ["has_task", "depends_on"], includeKinds: ["Task"], excludeRoot: true, }); // tasksOnly.nodes values are typed as Node ``` **Bidirectional traversal:** ```typescript // Find all nodes connected to a skill, regardless of edge direction const neighborhood = await store.subgraph(skill.id, { edges: ["uses_skill", "has_task"], direction: "both", maxDepth: 3, }); ``` #### Subgraph Projection By default, `subgraph()` returns fully hydrated nodes and edges. The `project` option lets you specify which properties to keep per kind, reducing payload size and enabling SQL-level field extraction via `json_extract()` / JSONB paths. ```typescript const result = await store.subgraph(rootId, { edges: ["has_task", "uses_skill"], maxDepth: 2, project: { nodes: { Task: ["title", "meta"], Skill: ["name"], }, edges: { uses_skill: ["priority"], }, }, }); // Task → { kind, id, title, meta } — status omitted, compile-time error to access // Skill → { kind, id, name } // uses_skill → { id, kind, fromKind, fromId, toKind, toId, priority } ``` **Projection rules:** - Projected nodes always retain `kind` and `id`; projected edges always retain structural fields (`id`, `kind`, `fromKind`, `fromId`, `toKind`, `toId`). - Kinds omitted from `project` remain fully hydrated. - Include `"meta"` in the field list for the full metadata object, or omit it entirely. No partial metadata selection — the struct is small enough that subsetting adds complexity without savings. - Node projection keys must exist in `includeKinds` (or be any node kind when `includeKinds` is omitted). Edge projection keys must be in `edges`. Out-of-scope keys are a compile-time error. **Type narrowing:** Result types narrow per-kind based on the projection. Accessing an omitted field is a compile-time error: ```typescript for (const node of result.nodes.values()) { if (node.kind === "Task") { console.log(node.title); // OK console.log(node.status); // TypeScript error — status was not projected } } ``` #### `defineSubgraphProject()` When storing a projection config in a variable, TypeScript widens field arrays to `string[]`, defeating compile-time narrowing. Use `defineSubgraphProject()` to preserve literal types: ```typescript import { defineSubgraphProject } from "@nicia-ai/typegraph"; const agentProjection = defineSubgraphProject()({ nodes: { Task: ["title", "status"], Skill: ["name"], }, edges: { uses_skill: ["priority"], }, }); // Reuse across calls — types are preserved const result = await store.subgraph(rootId, { edges: ["has_task", "uses_skill"], project: agentProjection, }); ``` #### Choosing a query strategy TypeGraph offers several ways to load related data. The right choice depends on your access pattern: | Pattern | Best strategy | Why | |---------|--------------|-----| | Load entity with all relationships | `subgraph(maxDepth: 1)` | Fixed 2 SQLite / 3 PostgreSQL statements — recursive traversal cost does not grow with edge count | | Load entity with deep chain | `subgraph(maxDepth: N)` | Recursive CTE handles multi-hop without extra round trips per hop | | Filter/sort within a relationship | `.query().traverse()` | Fluent query supports WHERE/ORDER/LIMIT on target nodes, in one statement | | Several independent bounded reads in one round trip | `store.batchOnce()` | Embeds fluent queries, relations, neighbors, counts, or subgraphs in exactly one statement | | Multiple independent queries with per-query control | `store.batch()` | Typed tuple results, at most one query in flight — still at least a statement per query, and not a snapshot | | Check if an edge exists | `edges.X.findFrom()` | Lightweight — no node resolution needed; honors the graph's temporal mode by default | | Traverse + resolve one edge type | `edges.X.findFrom()` + `nodes.X.getByIds()` | Two queries, simple and explicit; pass `temporalMode` / `asOf` when reading history | | Shortest path, reachability, neighborhoods, degree | `store.algorithms.*` | Set-based BFS frontier or a single `COUNT` — see [Graph Algorithms](/graph-algorithms) | **Key insight:** `subgraph()` costs a fixed number of statements — 2 on SQLite (nodes, edges) and 3 on PostgreSQL (closure ids, then nodes and edges) — regardless of how many edge types it traverses or how much it returns. Parallel `findFrom` calls scale linearly instead: one per edge type, plus additional queries for node resolution. The gap widens as relationship count grows. For the common "load an entity and everything it touches" pattern (detail pages, config hydration, template instantiation), use `subgraph()` with `maxDepth: 1`. When one request needs several independent bounded reads, use [`store.batchOnce()`](#batch-query-execution) to keep them in one statement. Use `store.batch()` when a member cannot be embedded or sequential transaction execution is the intended contract; it still costs at least one statement per query. Reserve individual fluent queries for one-off operations. ### Graph Algorithms #### `store.algorithms` Lazy-initialized facade exposing the graph algorithms — `shortestPath`, `reachable`, `canReach`, `neighbors`, and `degree`. See [Graph Algorithms](/graph-algorithms) for the full API; this section is a quick reference. ```typescript // Shortest path between two nodes const path = await store.algorithms.shortestPath(alice, bob, { edges: ["knows"], }); // Every reachable node with its discovery depth const reachable = await store.algorithms.reachable(alice, { edges: ["knows"], maxHops: 5, }); // Fast boolean reachability check const connected = await store.algorithms.canReach(alice, bob, { edges: ["knows"], }); // k-hop neighborhood (source excluded) const twoHop = await store.algorithms.neighbors(alice, { edges: ["knows"], depth: 2, }); // Count incident edges const total = await store.algorithms.degree(alice, { edges: ["knows"] }); ``` Every traversal algorithm accepts `edges`, `maxHops` (default 10), `direction` (`"out" | "in" | "both"`, default `"out"`), and the compatibility-only `cyclePolicy`, plus `temporalMode` / `asOf` for temporal filtering — see [Temporal Behavior](/graph-algorithms#temporal-behavior). Traversal calls expand a de-duplicated BFS frontier one level at a time; `degree` compiles to a single `COUNT`. Node arguments accept either raw IDs or any object with an `id` field — `Node`, `NodeRef`, and the lightweight records returned by these algorithms all work. ### Query Builder #### `store.query()` Creates a query builder. See [Query Builder](/queries/overview) for full documentation. ```typescript const results = await store .query() .from("Person", "p") .whereNode("p", (p) => p.name.startsWith("A")) .select((ctx) => ctx.p) .execute(); ``` **Execution methods** (see [Execute](/queries/execute) for details): | Method | Returns | Description | |--------|---------|-------------| | `execute()` | `Promise` | Run query, return all results | | `first()` | `Promise` | Return first result or undefined | | `count()` | `Promise` | Count matching results | | `exists()` | `Promise` | Check if any results exist | | `paginate(options)` | `Promise>` | Cursor-based pagination | | `page(options)` | `CompiledOneStatementRead>` with `execute()` | Cold cursor page; executes alone or in `batchOnce()` | | `stream(options?)` | `AsyncIterable` | Stream results in batches | | `prepare()` | `PreparedQuery` | Validate query AST once for repeated execution with different parameters | #### `store.batchOnce(buildReads, options?)` and `store.batch(...queries)` Use `batchOnce()` to embed independent fluent and batch-scoped set-oriented reads in one statement. Use `batch()` for mixed fluent and queued collection reads that may run sequentially. See [Batch Query Execution](#batch-query-execution) for the exact contracts. ### Dynamic Collection Access The typed `store.nodes.*` and `store.edges.*` accessors require the kind name at compile time. When the kind is determined at runtime — iterating all kinds, resolving a node from edge metadata, building admin UIs or snapshot tools — use `getNodeCollection` and `getEdgeCollection` instead. #### `store.getNodeCollection(kind)` Returns the [`DynamicNodeCollection`](/types#dynamicnodecollection) for the given kind, or `undefined` if the kind is not registered in this graph. ```typescript import { getNodeKinds } from "@nicia-ai/typegraph"; // Count every node kind const counts: Record = {}; for (const kind of getNodeKinds(graph)) { const collection = store.getNodeCollection(kind); if (collection) { counts[kind] = await collection.count(); } } // Resolve a node from edge metadata const collection = store.getNodeCollection(edge.fromKind); const node = await collection?.getById(edge.fromId); ``` #### `store.getEdgeCollection(kind)` Returns the [`DynamicEdgeCollection`](/types#dynamicedgecollection) for the given kind, or `undefined` if the kind is not registered in this graph. ```typescript import { getEdgeKinds } from "@nicia-ai/typegraph"; // Snapshot all edges for (const kind of getEdgeKinds(graph)) { const collection = store.getEdgeCollection(kind); if (collection) { const edges = await collection.find({ limit: 10_000 }); snapshot.push(...edges); } } ``` The returned collections expose the full API (`create`, `getById`, `find`, `count`, `createFromRecord`, etc.) with widened generics — see [`DynamicNodeCollection`](/types#dynamicnodecollection) and [`DynamicEdgeCollection`](/types#dynamicedgecollection). Both edge lookups are also available inside `transaction`, `withTransaction`, and receipt-enabled transaction contexts. They resolve the transaction's own collections; lookups inside `measure` contribute to that scope's receipt. `getEdgeCollectionOrThrow` throws `KindNotFoundError` for an unknown kind and also accepts a Store-issued runtime edge token, using the same token validation as the Store. Known graph keys retain their edge property schema while endpoints are validated at runtime. For generic graph/kind helpers, use these lookups instead of casting `tx.edges[kind]`; see [dynamic edge types and migration](/types#dynamicedgecollection). Valid-time views expose `view.getEdgeCollection(kind)` for dynamic reads at their pinned coordinate. ### Dynamic Props Schema Access Returns the live `z.ZodObject` the store uses internally to validate `.create()` / `.update()` props. Same accessor for compile-time and graph-extension kinds. Useful for MCP tool wrappers that want to validate inputs against the same schema as the store, and for producing richer JSON Schema (refinements, formats, branded `searchable()` / `embedding()` types) than `introspect().properties` exposes. ```typescript store.getNodePropsSchema(kind: string): z.ZodObject | undefined; store.getNodePropsSchemaOrThrow(kind: string): z.ZodObject; store.getEdgePropsSchema(kind: string): z.ZodObject | undefined; store.getEdgePropsSchemaOrThrow(kind: string): z.ZodObject; ``` `Object.hasOwn`-gated lookup matches `getNodeCollection` (no prototype-name leakage). The `OrThrow` variants throw `KindNotFoundError` with `kindName`, `entity`, and host `graphId` when the kind is not registered. Identity holds for compile-time kinds: `store.getNodePropsSchema("Person") === Person.schema`. ```typescript import { z } from "zod"; const schema = store.getNodePropsSchemaOrThrow("Paper"); // Validate tool input with the same schema the store uses. const parsed = schema.parse(input); await store.getNodeCollectionOrThrow("Paper").create(parsed); // Produce JSON Schema for an MCP tool description. const jsonSchema = z.toJSONSchema(schema); ``` **Props-only contract.** These accessors return only the props validator. Failed `schema.parse()` throws `ZodError`; failed `collection.create()` wraps the same underlying issues in `ValidationError`. Operation-level checks — uniqueness, endpoint resolution (edges validate endpoints before props), temporal validity, backend constraints — still run only through `collection.create` / `update`. ### Registry Access #### `store.registry` Access to the type registry for ontology lookups. The registry is an internal type; use `store.registry` directly without importing its type. See [Ontology](/ontology) for registry methods. ### Search (`store.search`) Search operations are grouped under the `store.search` facade. The full guide lives in [Fulltext Search](/fulltext-search); this section is the signature reference. ```typescript store.search.fulltext(nodeKind, options): Promise>[]>; store.search.hybrid(nodeKind, options): Promise>[]>; store.search.rebuildFulltext(nodeKind?, options?): Promise; ``` #### `store.search.fulltext(nodeKind, options)` Runs a ranked fulltext query against nodes of the given kind. Requires at least one `searchable()` field on the node schema. `hit.node` is narrowed to the typed node for `nodeKind` — no cast required. | Option | Type | Default | Description | |--------|------|---------|-------------| | `query` | `string` | — *(required)* | Query string. Parsed according to `mode`. | | `limit` | `number` | — *(required)* | Max rows. Positive integer. | | `mode` | `"websearch" \| "phrase" \| "plain" \| "raw"` | `"websearch"` | Parser for `query`. | | `language` | `string` | per-row | Language override (Postgres only; throws on FTS5). | | `minScore` | `number` | — | Drop hits below this backend-native score. | | `includeSnippets` | `boolean` | `false` | Return a `…` snippet per hit. | #### `store.search.hybrid(nodeKind, options)` Runs a vector + fulltext hybrid query and fuses the two ranked lists with Reciprocal Rank Fusion. Requires both `vectorSearch` and `fulltextSearch` capabilities on the backend. | Option | Type | Default | Description | |--------|------|---------|-------------| | `limit` | `number` | — *(required)* | Final fused result count. | | `vector.fieldPath` | `string` | — *(required)* | Embedding field on the node. | | `vector.queryEmbedding` | `readonly number[]` | — *(required)* | Query vector. | | `vector.metric` | `"cosine" \| "l2" \| "inner_product"` | `"cosine"` | Distance metric. | | `vector.k` | `number` | `4 × limit` | Vector-side candidates to fuse. | | `vector.minScore` | `number` | — | Vector-side score floor. | | `fulltext.query` | `string` | — *(required)* | Fulltext query string. | | `fulltext.k` | `number` | `4 × limit` | Fulltext-side candidates to fuse. | | `fulltext.mode` | `FulltextQueryMode` | `"websearch"` | Parser mode. | | `fulltext.language` | `string` | per-row | Language override. | | `fulltext.minScore` | `number` | — | Fulltext-side score floor. | | `fulltext.includeSnippets` | `boolean` | `false` | Return snippets per fulltext sub-hit. | | `fusion.method` | `"rrf"` | `"rrf"` | Fusion method. | | `fusion.k` | `number` | `60` | RRF constant. | | `fusion.weights.vector` | `number` | `1` | Bias toward the vector retriever. | | `fusion.weights.fulltext` | `number` | `1` | Bias toward the fulltext retriever. | Each `HybridSearchHit` exposes `vector` and `fulltext` sub-results (each with its own `rank` and `score`) for ranking debugging. #### `store.search.rebuildFulltext(nodeKind?, options?)` Rebuilds the fulltext index from existing node data. Use after a schema change, a `DROP TABLE` / `TRUNCATE` of the fulltext table, or bulk inserts that bypassed the store. Run during a maintenance window for full consistency — concurrent hard-deletes between page fetches can be missed by a single pass. | Option | Type | Default | Description | |--------|------|---------|-------------| | `nodeKind` | `string \| undefined` | all kinds | Scope to a single kind. | | `options.pageSize` | `number` | `500` | Keyset page size. Positive integer. | | `options.maxSkippedIds` | `number` | `10_000` | Cap on returned `skippedIds`. Raise for forensic runs. | Returns `{ kinds, processed, upserted, cleared, skipped, skippedIds, skippedTruncated }`. See [Fulltext Search](/fulltext-search) for query modes, RRF tuning, `FulltextStrategy` customization, and troubleshooting. ### Temporal Views (`store.asOf` and `store.view`) A `StoreView` is a **read-only** lens that pins one temporal coordinate and routes every supported read through it — the as-of database value, in the style of Datomic `(d/as-of db t)` and SQL:2011 `FOR SYSTEM_TIME AS OF`. Use it when several reads should share the same temporal coordinate; reach for the per-query [`.temporal("asOf", T)`](/queries/temporal#point-in-time-queries-asof) when only one query needs it. ```typescript store.asOf(asOf: string): StoreView; store.view(coordinate: { mode: TemporalMode; asOf?: string }): StoreView; store.snapshot(): StoreView; ``` - **`store.asOf(T)`** pins valid-time `asOf` mode at timestamp `T`. - **`store.view({ mode, asOf })`** pins any public mode (`"current"`, `"asOf"`, `"includeEnded"`, `"includeTombstones"`). `asOf` is required for `"asOf"` mode. - **`store.snapshot()`** pins the current instant, captured once at construction — sugar for `store.asOf(new Date().toISOString())`. Unlike `store.view({ mode: "current" })` (which tracks "now" live and may read different surfaces against slightly different clocks), a snapshot is a stable point-in-time value where every surface observes the same instant. Mirrors Datomic's `(d/db conn)`. `asOf` must be a canonical UTC ISO-8601 timestamp (`YYYY-MM-DDTHH:mm:ss.sssZ`) — a date-only, zoned-offset, or natural-language string is rejected with a `ValidationError`, because the temporal filters compare it as text. ```typescript const past = store.asOf("2026-01-01T00:00:00.000Z"); const alice = await past.nodes.Person.getById(aliceId); const jobs = await past.edges.worksAt.findFrom(alice); const names = await past .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq("Alice")) .select((ctx) => ctx.p.name) .execute(); const reach = await past.reachable(aliceId, { edges: ["knows"] }); const sg = await past.subgraph(aliceId, { edges: ["knows"] }); ``` #### Surface The view exposes the read surface of the `Store`, each pinned to its coordinate: | Surface | Behavior | | --- | --- | | `view.nodes` / `view.edges`: `getById`, `getByIds`, `find`, `count` | pinned | | `view.edges`: `findFrom`, `findTo`, `bulkFindFrom`, `bulkFindTo`, `findByEndpoints` | pinned | | `view.bulkFindEdgesFrom(params, options?)` | pinned | | `view.bulkFindEdgesTo(params, options?)` | pinned | | `view.query()` | a pinned query builder with a **sealed** temporal axis — `.temporal(...)` throws | | `view.subgraph(rootId, options)` | pinned | | `view.reachable` / `canReach` / `shortestPath` / `neighbors` / `degree` | pinned | | `view.nodes`: `findByConstraint` / `bulkFindByConstraint` / `bulkFindByIndex` | current-only reads: delegate on a `"current"` view; reject on any temporal pin | | `view.search` reads (`fulltext` / `vector` / `hybrid`) | delegate to the live search on a `"current"` view; reject on any other pin | | `view.search.rebuildFulltext()` | rejected on every view (maintenance write) | | `view.mode` / `view.asOf` | the pinned coordinate | The algorithm and `subgraph` option objects are the same as on the live `Store` **minus** `temporalMode` / `asOf`, which the pin supplies. `view.query()` is a **capability-safe** pinned read context: the returned query builder seeds the view's coordinate and seals the temporal axis, so calling `.temporal(...)` on it (or on any builder derived from it) throws a `ConfigurationError`. To read at a different coordinate, construct a different view or use the live `store.query()`. #### Read-only A view is read-only by construction. Writes (`create` / `update` / `delete` / `upsert*` / `bulk*` / `getOrCreate*`) and temporally-unscoped reads on a view collection reject with a `ConfigurationError`, and the view exposes no `transaction`. Perform writes on the live `Store`. Constraint / index lookups (`findByConstraint`, `bulkFindByConstraint`, `bulkFindByIndex`) read current state only — they have no temporal axis — so a view **delegates** them on a `"current"` view and **rejects** them on any temporal pin (rather than silently returning current data while every sibling read is pinned). `search` is refused on a non-`"current"` view for the same reason: the fulltext / vector index reflects current state only. (Edge `findByEndpoints` *does* have a temporal axis and is pinned like `findFrom`.) See [Temporal queries](/queries/temporal#shared-coordinate-views-storeasof) for worked examples. #### Recorded time (`store.asOfRecorded`) With a store created with `{ history: true }` or an explicit `recordedRead` binding, `store.asOfRecorded(T)` returns a `RecordedStoreView` — a narrow read-only lens that reconstructs the graph as the recorded relation represented it at instant `T` (the system-time axis), composing with the valid-time coordinate above for bitemporal graph reads. ```typescript store.asOfRecorded(recordedAsOf: RecordedInstant): RecordedStoreView; // also: store.asOf(validT).asOfRecorded(recordedT) // store.view({ mode }).asOfRecorded(recordedT) store.recordedNow(): Promise; asRecordedInstant(value: string): RecordedInstant; // re-brand a persisted anchor recordedInstantRevision(value: RecordedInstant): number; recordedInstantWallTime(value: RecordedInstant): string; compareRecordedInstants(a: RecordedInstant, b: RecordedInstant): -1 | 0 | 1; ``` - **`store.asOfRecorded(T)`** is diagonal sugar — the recorded *and* valid axes both at `T`. Chain from `store.asOf(validT)` / `store.view({ mode })` to pin the two axes independently. - **`T` is a `RecordedInstant`**, a branded canonical string encoded as `r1:<16-digit revision>:`. It comes from `store.recordedNow()` or from `asRecordedInstant(...)` after the exact anchor has round-tripped through untyped storage. A raw wall-clock string (`new Date().toISOString()`) is a compile error because it cannot distinguish multiple commits in one millisecond. The logical revision orders commits; the timestamp is a non-decreasing physical wall-time high-water mark. Use `recordedInstantRevision(T)`, `recordedInstantWallTime(T)`, and `compareRecordedInstants(a, b)` instead of splitting or comparing anchor strings manually. Comparisons are meaningful only within one graph. See [Logical revision and physical time](/queries/temporal#logical-revision-and-physical-time). - **`store.recordedNow()`** returns the recorded high-water mark — the latest captured recorded instant. After guarding the `undefined` case, `store.asOfRecorded(checkpoint)` reconstructs everything committed so far. Use it as a deterministic anchor instead of the wall clock. Returns `undefined` before the first capture; throws if the store was not created with `{ history: true }`. - **`recordedRead`** binds an externally populated recorded relation for reads only. It does not capture TypeGraph writes, advance TypeGraph's recorded clock, or make `store.recordedNow()` available. It must be created with `recordedRelation({ schema })` using a `createSqlSchema(...)` schema and cannot be combined with `history: true`. - The view exposes only **reconstructing** reads: `nodes` / `edges` point reads (`getById` / `getByIds`) and bounded deterministic `scan()` pages, a sealed `query()`, `subgraph()`, and the graph algorithms (`reachable` / `canReach` / `shortestPath` / `degree`). Broad filtered collection reads, `search`, and fulltext / vector predicates reject — those indexes reflect current state only. - Built-in capture covers TypeGraph collection writes. Out-of-band database writes and row-returning raw SQL paths are not captured into the recorded relations. Adopt an external transaction under `history: true` with the callback form `store.withRecordedTransaction(externalTx, async (tx) => ...)`, which flushes capture before the caller commits. `store.withTransaction(...)` is a compile error on a history store, and the typed history transaction context omits raw `tx.sql`. Branch on `tx.sqlAvailability` (`"history"`) before accessing the SQL handle. See [Recorded time](/queries/temporal#recorded-time-bitemporal) for the full guide. ## Observability Hooks TypeGraph supports observability hooks for monitoring and logging store operations. Query hooks describe SQL statements submitted by Store read APIs, not logical API calls or backend-internal setup statements. Fluent queries, `batchOnce()`, `neighbors()`, `countNeighbors()`, and `subgraph()` all use this observed execution path. A logical read that submits more than one statement fires one start/end pair per statement: direct `subgraph()` emits two pairs on SQLite and three on PostgreSQL, while `tx.subgraph()` and the same subgraph embedded in `batchOnce()` emit one. A fluent query that retries with a different projection likewise fires a pair for each statement it submits. ### `StoreHooks` Configuration for observability callbacks: ```typescript import type { HookContext, QueryHookContext, OperationHookContext, StoreHooks, } from "@nicia-ai/typegraph"; ``` ```typescript type StoreHooks = Readonly<{ onQueryStart?: (ctx: QueryHookContext) => void; onQueryEnd?: (ctx: QueryHookContext, result: { rowCount: number; durationMs: number }) => void; onOperationStart?: (ctx: OperationHookContext) => void; onOperationEnd?: ( ctx: OperationHookContext, result: { durationMs: number; outcome: "written" | "unchanged" | "unknown" }, ) => void; onError?: (ctx: HookContext, error: Error) => void; }>; type HookContext = Readonly<{ operationId: string; graphId: string; startedAt: Date; /** 1-based attempt inside a retried transaction; absent means 1. */ attempt?: number; }>; type QueryHookContext = HookContext & Readonly<{ sql: string; params: readonly unknown[]; }>; type OperationHookContext = HookContext & Readonly<{ operation: "create" | "update" | "delete"; entity: "node" | "edge"; kind: string; id: string; }>; ``` > **Note:** Batch operations (`bulkCreate`, `bulkInsert`, `bulkUpsertById`, `bulkDelete`) skip per-item operation hooks for throughput, and the set-based bulk hooks (`onBulkOperationStart` / `onBulkOperationEnd`) do not stand in for them — those fire only for node `updateWhere`, so a batch method emits no hook events at all, neither per-item nor bulk. Call the single-item method to observe each write. Query hooks still fire normally. `onOperationEnd.result.outcome` is `"written"` when durable graph state changed. It is `"unchanged"` when an authoritative write attempt completed without a logical mutation—for example, when a one-statement durable edge get-or-create found the incumbent. Expected convergence is therefore a successful unchanged completion, never an `onError` event. This is the same decision TypeGraph uses to suppress revision/history churn. It is `"unknown"` when the backend command does not report an authoritative physical-write verdict; TypeGraph never guesses from a successful return alone. **Example:** ```typescript import { createStore, type StoreHooks } from "@nicia-ai/typegraph"; const hooks: StoreHooks = { onQueryStart: (ctx) => { console.log(`[${ctx.operationId}] SQL: ${ctx.sql}`); }, onQueryEnd: (ctx, result) => { console.log(`[${ctx.operationId}] ${result.rowCount} rows in ${result.durationMs}ms`); }, onOperationStart: (ctx) => { console.log(`[${ctx.operationId}] ${ctx.operation} ${ctx.entity}:${ctx.kind}`); }, onOperationEnd: (ctx, result) => { console.log( `[${ctx.operationId}] ${result.outcome} in ${result.durationMs}ms`, ); }, onError: (ctx, error) => { console.error(`[${ctx.operationId}] Error:`, error.message); }, }; const store = createStore(graph, backend, { hooks }); // CRUD operations trigger operation hooks; observed Store read statements // trigger query hooks. await store.nodes.Person.create({ name: "Alice" }); await store.query().from("Person", "p").select((ctx) => ctx.p).execute(); // Logs include: // [op-abc123] create node:Person // [op-abc123] Completed in 5ms // [query-def456] SQL: WITH ... SELECT ... // [query-def456] 1 rows in 2ms ``` # Semantic Search > Vector embeddings and similarity search for AI-powered retrieval TypeGraph supports semantic search using vector embeddings, enabling you to find semantically similar content using embedding models like OpenAI, Sentence Transformers, CLIP, or any model that produces fixed-dimension vectors. ## Overview Traditional search relies on exact keyword matching. Semantic search understands meaning—"machine learning" matches documents about "neural networks" and "AI algorithms" even without those exact words. **Key capabilities:** - Store embeddings as node properties alongside your graph data - Find the k most similar nodes using cosine, L2, or inner product distance - Combine semantic similarity with graph traversals and standard predicates - Automatic vector indexing for fast approximate nearest neighbor search ## Use Cases ### Retrieval-Augmented Generation (RAG) Build context-aware AI applications by retrieving relevant documents before generating responses: ```typescript async function ragQuery(question: string): Promise { const questionEmbedding = await embed(question); const context = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(questionEmbedding, 5, { metric: "cosine", minScore: 0.7, }) ) .select((ctx) => ({ title: ctx.d.title, content: ctx.d.content, })) .execute(); return await llm.chat({ messages: [ { role: "system", content: `Answer based on this context:\n${context.map((d) => d.content).join("\n\n")}`, }, { role: "user", content: question }, ], }); } ``` ### Semantic Document Search Find documents by meaning rather than keywords: ```typescript const results = await store .query() .from("Article", "a") .whereNode("a", (a) => a.embedding .similarTo(queryEmbedding, 20) .and(a.category.eq("technology")) ) .select((ctx) => ctx.a) .execute(); ``` ### Image Similarity Use CLIP or similar vision models for image search: ```typescript const similarImages = await store .query() .from("Image", "i") .whereNode("i", (i) => i.clipEmbedding.similarTo(queryImageEmbedding, 10)) .select((ctx) => ({ url: ctx.i.url, caption: ctx.i.caption, })) .execute(); ``` ### Product Recommendations Recommend products based on embedding similarity: ```typescript const recommendations = await store .query() .from("Product", "p") .whereNode("p", (p) => p.embedding .similarTo(referenceProductEmbedding, 10) .and(p.inStock.eq(true)) ) .select((ctx) => ctx.p) .execute(); ``` ## Database Setup Vector search requires database-specific extensions for storing and querying high-dimensional vectors efficiently. ### PostgreSQL with pgvector [pgvector](https://github.com/pgvector/pgvector) is the recommended extension for PostgreSQL. It provides: - Native `vector` column type - HNSW and IVFFlat indexes for fast approximate nearest neighbor search - Support for cosine, L2, and inner product distance **Installation:** ```sql -- Install the extension (requires superuser or database owner) CREATE EXTENSION vector; ``` **Docker setup:** ```yaml services: postgres: image: pgvector/pgvector:pg16 environment: POSTGRES_PASSWORD: password POSTGRES_DB: myapp ports: - "5432:5432" ``` **TypeGraph migration enables vector support:** ```typescript import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; // Generates DDL including `CREATE EXTENSION IF NOT EXISTS vector;`. // It does NOT create a single embeddings table — each embedding field gets // its own typed `vector(N)` table, provisioned by `createStoreWithSchema` // (the privileged migrator) at boot (see Storage Layout below). const migrationSQL = generatePostgresMigrationSQL(); ``` ### SQLite with sqlite-vec [sqlite-vec](https://github.com/asg017/sqlite-vec) provides vector search for SQLite. It offers: - `vec_f32` type for 32-bit float vectors - Cosine and L2 distance functions :::caution[sqlite-vec requires a native (better-sqlite3) connection] sqlite-vec is a loadable C extension. TypeGraph loads it through better-sqlite3's `loadExtension` in `createLocalSqliteBackend`, so it only applies to the **local, native** SQLite backend. It does **not** apply to the **libSQL / Turso** backend (`createLibsqlBackend`): `@libsql/client` does not expose `loadExtension`, and libSQL ships its **own** native vector engine (`F32_BLOB`, `vector_distance_cos`, `vector_top_k`) which is a different API than sqlite-vec. See [libSQL / Turso](#libsql--turso-native-vectors) below. ::: **Installation:** ```bash npm install sqlite-vec ``` **Loading the extension:** ```typescript import Database from "better-sqlite3"; import * as sqliteVec from "sqlite-vec"; const sqlite = new Database("myapp.db"); sqliteVec.load(sqlite); ``` **Limitations:** - sqlite-vec does not support inner product distance - Use `cosine` or `l2` metrics only ### libSQL / Turso (native vectors) The **libSQL / Turso** backend (`createLibsqlBackend`) does **not** use sqlite-vec. libSQL has a built-in vector engine — no extension to load — so vector and hybrid search work out of the box on local files, embedded replicas, and remote Turso databases: ```typescript import { createClient } from "@libsql/client"; import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql"; const client = createClient({ url: "libsql://my-db.turso.io", authToken: "..." }); const { backend } = await createLibsqlBackend(client); // backend.capabilities.vector?.supported === true ``` Under the hood it stores embeddings as `F32_BLOB` and searches with `vector_distance_cos` / `vector_distance_l2`, with optional approximate nearest-neighbor (DiskANN) indexes via `libsql_vector_idx` + `vector_top_k`. Supported metrics are `cosine` and `l2` (no `inner_product`), matching the sqlite-vec feature set. One caveat specific to DiskANN: `vector_top_k` is a table function with no filter pushdown, so the liveness filter every search applies (only non-deleted nodes may rank — see below) runs *after* ANN retrieval. TypeGraph over-fetches 4× `limit` neighbors to leave headroom; if more than 3×`limit` of those neighbors are filtered out, fewer than `limit` results return. pgvector and sqlite-vec apply the filter inside the index scan and do not share this bound. ### Supported Distance Metrics | Metric | PostgreSQL | SQLite (sqlite-vec) | libSQL / Turso | Description | |--------|------------|---------------------|----------------|-------------| | `cosine` | `<=>` | `vec_distance_cosine` | `vector_distance_cos` | Cosine distance (1 - similarity). Best for normalized embeddings. | | `l2` | `<->` | `vec_distance_l2` | `vector_distance_l2` | Euclidean distance. Good for unnormalized vectors. | | `inner_product` | `<#>` | Not supported | Not supported | Negative inner product. For maximum inner product search (MIPS). | ## Storage Layout & Maintenance Each embedding field is stored in its own typed, graph-scoped table named `tg_vec___`, carrying that field's fixed dimension (pgvector `vector(N)`, libSQL `F32_BLOB(N)`, sqlite-vec `vec0`). The privileged migrator (`createStoreWithSchema`, and `evolve()` for runtime-added fields) provisions each table plus a durable contribution marker at boot; the runtime hot path then asserts the marker (a cached SELECT) and never issues DDL, so a least-privilege, DML-only role can read and write embeddings. An embedding write against an un-provisioned slot throws `StoreNotInitializedError` rather than lazily creating the table; vector reads (`store.search.vector`, `store.search.hybrid`, and query-builder `.similarTo()` predicates) compile straight to SQL, so they surface the engine's missing-relation error instead — use `createVerifiedStore` to catch both at attach. See [Database roles & least privilege](/backend-setup#database-roles--least-privilege). Graph-scoping means several graphs in one database can declare the same `kind`+`field` at different dimensions without collision. This is transparent to queries — `.similarTo()`, `store.search.vector`, and `store.search.hybrid` read it for you. ### Deleted nodes never rank Every facade search (`store.search.vector` / `fulltext` / `hybrid`) computes its top-k over live nodes only: the search SQL constrains candidates to non-deleted node ids, so a stale embedding or fulltext row — one whose node was tombstoned by a writer that bypassed the store's cleanup — can neither surface in results nor crowd live rows out of the top-k. You always get `limit` results when at least `limit` live matches exist (on libSQL DiskANN, subject to the over-fetch bound above). ### Changing an embedding dimension Switching embedding models usually changes the vector dimension. Stored vectors can't be reinterpreted at a new dimension, so a stray write at the old dimension throws `EmbeddingDimensionChangedError`. Update the field's `embedding(N)` declaration, then recompute the stored vectors with `store.reembedVectorField()`, which recreates the field's storage at the new dimension: ```typescript // embedding(1536) → embedding(3072): recreate storage and re-embed in batches. // `embed` receives a page of nodes and returns a Map from node id to vector. await store.reembedVectorField("Document", "embedding", { embed: async (nodes) => { const vectors = await batchEmbed(nodes.map((node) => node.content)); return new Map(nodes.map((node, index) => [node.id, vectors[index]])); }, }); // → { recreated: true, reembedded: } ``` Between the declaration change and the `reembedVectorField()` call, the slot is in a deliberate limbo: boot (`createStoreWithSchema` / `evolve()`) detects that the provisioned storage no longer matches the declared shape, warns, and leaves it untouched — it never recreates the table implicitly, because that would silently drop every stored vector. Embedding writes to the field fail with a `StoreNotInitializedError` whose reason is `stale` (its message points here) until `reembedVectorField()` recreates the storage and re-stamps its durable marker. Without an `embed` callback the storage is recreated empty and you re-embed via normal `update()` writes. ### Reclaiming removed embedding fields Removing an embedding field from a kind that still exists orphans its `tg_vec_*` table. `store.materializeRemovals()` reclaims it — it drops per-field tables for embedding fields no longer in the active schema and reports them in `reclaimedVectorFields`: ```typescript const { reclaimedVectorFields } = await store.materializeRemovals(); // → [{ kind: "Document", fieldPath: "embedding", status: "reclaimed" }] ``` The active schema is the source of truth, so a removed-then-re-added field is never dropped. The pass is idempotent. ### Migrating from the legacy shared table Earlier versions stored every embedding in a single shared `typegraph_node_embeddings` table. If you have existing data there, run the one-time, idempotent `migrateLegacyEmbeddings()` utility to copy it into the new per-field tables (new deployments need no action): ```typescript import { migrateLegacyEmbeddings } from "@nicia-ai/typegraph"; const result = await migrateLegacyEmbeddings({ backend }); // → { migrated, perField, skippedDimensionMismatch, legacyTablePresent } ``` ## Schema Design ### Defining Embedding Properties Use the `embedding()` function to define vector properties with a specific dimension: ```typescript import { defineNode, embedding } from "@nicia-ai/typegraph"; import { z } from "zod"; const Document = defineNode("Document", { schema: z.object({ title: z.string(), content: z.string(), embedding: embedding(1536), // OpenAI ada-002 dimension }), }); const Image = defineNode("Image", { schema: z.object({ url: z.string(), caption: z.string().optional(), clipEmbedding: embedding(512), // CLIP ViT-B/32 dimension }), }); ``` ### Common Embedding Dimensions | Model | Dimensions | Use Case | |-------|------------|----------| | all-MiniLM-L6-v2 | 384 | Fast, lightweight text embeddings | | CLIP ViT-B/32 | 512 | Image-text multimodal | | BERT base | 768 | General text embeddings | | OpenAI ada-002 | 1536 | High-quality text embeddings | | OpenAI text-embedding-3-small | 1536 | Efficient, high-quality | | OpenAI text-embedding-3-large | 3072 | Maximum quality | | Cohere embed-v3 | 1024 | Multilingual support | ### Optional Embeddings Embedding properties can be optional for gradual population: ```typescript const Article = defineNode("Article", { schema: z.object({ title: z.string(), content: z.string(), embedding: embedding(1536).optional(), }), }); // Create without embedding const article = await store.nodes.Article.create({ title: "Draft Article", content: "...", }); // Add embedding later via background job await store.nodes.Article.update(article.id, { embedding: await generateEmbedding(article.content), }); ``` ### Multiple Embeddings per Node Nodes can have multiple embedding fields for different purposes: ```typescript const Product = defineNode("Product", { schema: z.object({ name: z.string(), description: z.string(), imageUrl: z.string(), // Text embedding for description search textEmbedding: embedding(1536).optional(), // Image embedding for visual similarity imageEmbedding: embedding(512).optional(), }), }); ``` ## Storing Embeddings Embeddings are stored when creating or updating nodes: ```typescript // Using OpenAI import OpenAI from "openai"; const openai = new OpenAI(); async function generateEmbedding(text: string): Promise { const response = await openai.embeddings.create({ model: "text-embedding-ada-002", input: text, }); return response.data[0].embedding; } // Store with embedding const embedding = await generateEmbedding("Machine learning fundamentals"); await store.nodes.Document.create({ title: "ML Guide", content: "Machine learning fundamentals...", embedding: embedding, }); ``` ### Batch Embedding For bulk operations, batch your embedding API calls: ```typescript async function batchEmbed(texts: string[]): Promise { const response = await openai.embeddings.create({ model: "text-embedding-ada-002", input: texts, }); return response.data.map((d) => d.embedding); } // Process in batches const documents = await fetchDocumentsWithoutEmbeddings(); const batchSize = 100; for (let i = 0; i < documents.length; i += batchSize) { const batch = documents.slice(i, i + batchSize); const embeddings = await batchEmbed(batch.map((d) => d.content)); await store.transaction(async (tx) => { for (const [index, doc] of batch.entries()) { await tx.nodes.Document.update(doc.id, { embedding: embeddings[index], }); } }); } ``` ## Querying ### Basic Similarity Search Use `.similarTo()` to find the k most similar nodes: ```typescript const queryEmbedding = await generateEmbedding("neural networks"); const similar = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 10) // Top 10 most similar ) .select((ctx) => ({ title: ctx.d.title, content: ctx.d.content, })) .execute(); ``` ### Approximate retrieval for `.similarTo()` (opt-in) By default `.similarTo()` ranks with an exact distance scan — correct at any scale, and index-served by the PostgreSQL planner where the plan shape allows. When a kind declares an ANN index (`embedding(n)` defaults to `hnsw`), you can opt the predicate into the engine's native approximate retrieval: ```typescript const similar = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding .similarTo(queryEmbedding, 10, { approximate: true }) .and(d.status.eq("published")), ) .select((ctx) => ctx.d) .execute(); ``` This is a semantic change, never applied silently: results are subject to the index's recall. Composed predicates constrain the ANN candidate set — exactly on pgvector and sqlite-vec, bounded by over-fetch on libSQL DiskANN. A kind declared with `indexType: "none"` keeps the exact scan even with the opt-in — a declared degradation, since there is no ANN structure for `approximate` to opt into and the results are exactly what was asked for. **A `metric` override that differs from the field's declared metric is refused when `approximate: true` is stated.** Every engine materializes metric-specific ANN structures — `vec0` bakes `distance_metric` into the virtual table, libSQL's DiskANN index is built with `metric=…`, pgvector's index carries a per-metric operator class — so an ANN structure only retrieves under the metric it was built for. Retrieving by the declared metric and re-scoring under the override returns the declared metric's neighbors wearing the override's scores: the wrong rows, silently. The two options state something that cannot both hold, so the call throws a `ConfigurationError` naming both metrics rather than quietly serving the exact scan: ```typescript // Refused: the HNSW index is built for cosine. d.embedding.similarTo(queryEmbedding, 10, { approximate: true, metric: "l2", }); ``` `details` carries `nodeKind`, `fieldPath`, `requestedMetric`, `declaredMetric`, and `indexType`. Omit `metric` (or pass the declared one) to keep approximate retrieval, or drop `approximate` to scan exactly under the overriding metric. A slot declared `indexType: "none"` is not refused: there is no ANN structure to be bound to a metric. Note the deliberate asymmetry with the facade. `store.search.vector` and `store.search.hybrid` refuse **every** metric override that differs from the declared one, whether or not `approximate` was stated — their rule is broader because vector storage is built for the declared metric and the facade is the guided surface. The query builder's *exact* path stays wider on purpose: an exact scan computes any metric over the stored vectors correctly, and nothing was stated there that the engine cannot honor. Only the silent half — the combination that cannot be served — is closed here. ### Scoped facade search: filters, pagination, subclasses `store.search.vector` (and `fulltext` / `hybrid`) accept a `where` predicate, an `offset`, and `includeSubClasses` — all compiled into the search statement itself, so the engine ranks only eligible rows. A filter never costs you results: you get `limit` hits whenever `limit` matching nodes exist (on libSQL DiskANN, subject to the over-fetch bound above). ```typescript // Top 10 most similar *published* documents, second page. const hits = await store.search.vector("Document", { fieldPath: "embedding", queryEmbedding, limit: 10, offset: 10, where: (d) => d.status.eq("published"), }); // Search a kind and all of its subClassOf descendants; per-kind results // merge into one globally ordered ranking. Kinds that don't declare the // embedding field are skipped. const acrossKinds = await store.search.vector("Content", { fieldPath: "embedding", queryEmbedding, limit: 10, includeSubClasses: true, }); ``` The `where` predicate is compiled by the same query compiler as `store.query()` — property predicates behave identically, use the same declared indexes, and apply the same current-read semantics (tombstoned nodes and nodes outside their validity window never rank). Kinds expanded via `includeSubClasses` must share one declared metric: scores from different metrics cannot merge into one ranking (and a per-call `metric` cannot bridge the gap — each kind's storage is validated against its declared metric), so mixed-metric expansions throw; search those kinds separately. ### Choosing a Distance Metric ```typescript // Cosine similarity (default) - best for normalized embeddings d.embedding.similarTo(queryEmbedding, 10, { metric: "cosine" }) // L2 (Euclidean) distance - for unnormalized embeddings d.embedding.similarTo(queryEmbedding, 10, { metric: "l2" }) // Inner product - for maximum inner product search (PostgreSQL only) d.embedding.similarTo(queryEmbedding, 10, { metric: "inner_product" }) ``` **When to use each:** - **Cosine**: Most common choice. Works well with normalized embeddings (OpenAI, Sentence Transformers). Focuses on direction, not magnitude. - **L2**: Use when vector magnitude matters. Good for detecting exact duplicates. - **Inner product**: For MIPS (maximum inner product search). Useful when embeddings encode both relevance and importance in magnitude. ### Minimum Score Filtering Filter results below a similarity threshold: ```typescript const highQualityMatches = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 100, { metric: "cosine", minScore: 0.8, // Only results with similarity >= 0.8 }) ) .select((ctx) => ctx.d) .execute(); ``` The `minScore` parameter filters results using **similarity** (not distance): - **Cosine**: 1.0 = identical, 0.0 = orthogonal. Typical thresholds: 0.7-0.9 - **L2**: Maximum distance to include (lower = more similar) - **Inner product**: Minimum inner product value :::note[Similarity vs Distance] While the underlying database operators use distance (where 0 = identical for cosine), `minScore` uses similarity semantics for intuitive usage. TypeGraph converts internally: `distance_threshold = 1 - minScore` for cosine. ::: ### Combining with Predicates Semantic search integrates with all standard query predicates: ```typescript const filteredSearch = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding .similarTo(queryEmbedding, 20) .and(d.category.eq("technology")) .and(d.publishedAt.gte("2024-01-01")) .and(d.status.eq("published")) ) .select((ctx) => ctx.d) .execute(); ``` ### Combining with Graph Traversals Search within graph relationships: ```typescript // Find similar documents by authors I follow const personalizedSearch = await store .query() .from("Person", "me") .whereNode("me", (p) => p.id.eq(currentUserId)) .traverse("follows", "f") .to("Person", "author") .traverse("authored", "a", { direction: "in" }) .to("Document", "d") .whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 10) ) .select((ctx) => ({ title: ctx.d.title, author: ctx.author.name, })) .execute(); ``` ## Best Practices ### Normalize Your Embeddings Most embedding models produce normalized vectors (unit length). If yours doesn't, normalize before storing: ```typescript function normalize(vector: number[]): number[] { const magnitude = Math.sqrt(vector.reduce((sum, v) => sum + v * v, 0)); return vector.map((v) => v / magnitude); } await store.nodes.Document.create({ title: "Example", content: "...", embedding: normalize(rawEmbedding), }); ``` ### Use Consistent Embedding Models Always use the same model for both storing and querying: ```typescript // Bad: Mixing models const docEmbedding = await embed("text-embedding-ada-002", content); const queryEmbedding = await embed("text-embedding-3-small", query); // Different! // Good: Same model throughout const MODEL = "text-embedding-ada-002"; const docEmbedding = await embed(MODEL, content); const queryEmbedding = await embed(MODEL, query); ``` ### Handle Missing Embeddings Not all nodes may have embeddings. Handle gracefully: ```typescript // Only search nodes with embeddings const results = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding .isNotNull() .and(d.embedding.similarTo(queryEmbedding, 10)) ) .select((ctx) => ctx.d) .execute(); ``` ### Choose Appropriate k Values The `k` parameter (number of results) affects performance: ```typescript // For RAG: Small k (3-10) for focused context d.embedding.similarTo(query, 5) // For exploration: Larger k with pagination d.embedding.similarTo(query, 100) ``` ### Index Considerations Vector indexes (HNSW, IVFFlat) trade accuracy for speed: - **Small datasets (< 10K)**: Exact search is fast enough - **Medium datasets (10K-1M)**: HNSW provides good recall with fast queries - **Large datasets (> 1M)**: Consider IVFFlat with appropriate parameters TypeGraph creates HNSW indexes by default for optimal balance. ### Filtered vector search needs a node index on the filter field Combining `similarTo` with a property predicate is the shape that degrades first at scale — and the vector index is not the reason. The candidates side (`d.category.eq(...)`) is a JSON property predicate over the nodes table, and rows that carry an embedding field have LARGE props: on PostgreSQL the predicate scan detoasts every row, so at 50k documents (384-dim embeddings) the filter alone costs ~375ms regardless of how the vector side is executed. SQLite pays the same class of cost parsing large JSON props per row. Declare a node index on the filter field and materialize it — the candidates predicate becomes an index lookup: ```typescript import { defineNodeIndex } from "@nicia-ai/typegraph/indexes"; const categoryIndex = defineNodeIndex(Document, { fields: ["category"] }); const graph = defineGraph({ id: "docs", nodes: { Document: { type: Document } }, edges: {}, indexes: [categoryIndex], }); await store.materializeIndexes(); ``` Measured at 50k documents on PostgreSQL: the filtered exact search drops from ~375ms to ~19ms and the filtered approximate search to ~20ms — a ~20× difference from one declared index. The `bench:vector` lane tracks both forms (`vector:exact-filtered` before the index, `vector:exact-filtered-postindex` after). ### Approximate search under selective filters `approximate: true` combined with a highly selective property filter is the shape where approximate means it. The index scan walks neighbors best-first and keeps going until enough filtered rows surface (TypeGraph applies pgvector's `hnsw.iterative_scan = strict_order` automatically on transaction-capable Postgres drivers with pgvector ≥ 0.8 — the setting is transaction-scoped, so non-transactional backends such as `neon-http` keep the plain bounded scan), but the scan is still bounded by pgvector's `hnsw.max_scan_tuples` (default 20,000). If the nearest rows matching the filter live far from the query — a filter *correlated* with embedding geometry, like "category X" when category X's documents form their own distant cluster — the scan can exhaust its budget and return plausible-but-distant rows. For filters independent of the embedding space (the common case), filtered approximate recall stays near 1.0. When the filter is known to be geometry-correlated and selective, drop `approximate` (the exact path is index-assisted on the candidates side by a node index on the filter field) or raise `hnsw.max_scan_tuples`. ### Tuning recall per query with `efSearch` pgvector's HNSW index searches a dynamic candidate list whose size is the `hnsw.ef_search` GUC — **default 40**. That frontier caps how many neighbors a single scan can surface, so on corpora past a few million vectors recall@k flattens well below 1.0 at the default. TypeGraph exposes it as a per-search `efSearch` knob on `store.search.vector` and the vector half of `store.search.hybrid`: ```typescript const hits = await store.search.hybrid("Document", { limit: 20, vector: { fieldPath: "embedding", queryEmbedding, k: 80, // over-fetch 80 candidates from the vector side efSearch: 240, // ~3× k — high-recall frontier for this query }, fulltext: { query: "renewable energy" }, }); ``` Sizing guidance: - **Floor — `efSearch >= k`.** Hybrid over-fetches `k` candidates from the vector side (default `4 * limit`). If `efSearch` is below `k` the scan can't fill the candidate set, so the over-fetch silently under-delivers — RRF papers over this on head queries (the fulltext half covers the miss) but drops tail queries only the vector side knows about. - **Target — ~2–4× `k`.** On million-scale corpora this clears roughly 0.95 recall@10, versus ~0.82–0.85 at the default 40. Verify the curve against your own corpus rather than hard-coding a multiplier. - **Ceiling — 1000.** pgvector caps `hnsw.ef_search` at 1000; TypeGraph rejects a larger `efSearch` with a clear error. Because it's per-search, one connection pool can serve both a latency-sensitive interactive path (omit `efSearch`, inherit the session default) and a recall-sensitive batch/ETL path (raise it) — a session GUC can't, a per-call override can. **Mechanics and limits.** The override is applied transaction-locally (`SET LOCAL hnsw.ef_search`) around the vector `SELECT`, so it never leaks to the next query on a pooled connection. Omitting it preserves today's behavior exactly — no transaction is opened. It applies to the **Postgres HNSW** path only: - **SQLite backends refuse it.** Neither `sqlite-vec` — whose `vec0` KNN takes only `k`, the page size — nor `libsql-native`, whose DiskANN `vector_top_k` fixes `search_l` at index-creation time, has a per-search frontier to set. Supplying `efSearch` to either, on the vector path or the hybrid path, is refused with `UnsupportedBackendCapabilityError` (`details.capability` `vector.searchFrontierTuning`, `details.reason` naming the engine's limitation) rather than searching as if the option had not been passed. - **Transaction-less Postgres drivers** (`drizzle-orm/neon-http`) can't scope `SET LOCAL`, so a search that supplies `efSearch` is refused with `UnsupportedBackendCapabilityError`. Use a transactional driver (`node-postgres` / `neon-serverless` / `postgres-js`) to apply it. - It tunes HNSW only. Supplying it for IVFFlat is refused with a `ConfigurationError`; IVFFlat's analogous knob (`ivfflat.probes`) is not yet exposed. ## Troubleshooting ### "Extension not found" errors **PostgreSQL:** ```sql -- Check if pgvector is installed SELECT * FROM pg_extension WHERE extname = 'vector'; -- Install it CREATE EXTENSION vector; ``` **SQLite:** ```typescript // Ensure sqlite-vec is loaded before queries import * as sqliteVec from "sqlite-vec"; sqliteVec.load(sqlite); ``` ### "Inner product not supported" (SQLite) sqlite-vec only supports `cosine` and `l2` metrics. Use one of those instead: ```typescript // Instead of: d.embedding.similarTo(query, 10, { metric: "inner_product" }) // Use: d.embedding.similarTo(query, 10, { metric: "cosine" }) ``` ### Dimension mismatch errors Ensure query embedding has the same dimension as stored embeddings: ```typescript const Document = defineNode("Document", { schema: z.object({ embedding: embedding(1536), // 1536 dimensions }), }); // Query embedding must also be 1536 dimensions const queryEmbedding = await embed(text); // Verify this returns 1536-dim vector ``` ### Slow queries 1. **Check index creation**: Vector indexes may not exist 2. **Reduce k**: Smaller k = faster queries 3. **Add filters**: Pre-filter with standard predicates before similarity search 4. **Consider approximate search**: HNSW indexes sacrifice some accuracy for speed ## Hybrid Search: Combining with Fulltext Vector search excels at semantic similarity but misses exact matches — proper nouns, SKUs, code identifiers, rare technical terms. **Hybrid search** fuses vector and fulltext results with Reciprocal Rank Fusion and typically beats either approach alone. ```typescript // One query, both signals — fused with RRF at the SQL layer const results = await store .query() .from("Document", "d") .whereNode("d", (d) => d.$fulltext .matches("renewable energy", 50) .and(d.embedding.similarTo(queryVec, 50)) ) .select((ctx) => ctx.d) .limit(10) .execute(); ``` For tunable per-source weights and RRF parameters, use the store-level `store.search.hybrid()` API. See the [Fulltext Search guide](/fulltext-search) for the complete hybrid workflow. ## API Reference See the [Predicates documentation](/queries/predicates#embedding) for complete API reference of the `similarTo()` predicate and related options. See [Fulltext Search](/fulltext-search) for the `n.$fulltext.matches()` predicate and `searchable()` schema brand. # Testing > Patterns for testing code that uses TypeGraph TypeGraph's in-memory SQLite backend makes tests fast and isolated — each test gets a fresh database with zero setup cost. This guide covers test utilities, common patterns, and strategies for testing at different levels. ## Test Setup ### In-memory backend (recommended) `createLocalSqliteBackend()` creates an in-memory SQLite database with TypeGraph tables pre-configured. Each call returns a completely isolated database. ```typescript import { beforeEach, describe, expect, it } from "vitest"; import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; import { createStore } from "@nicia-ai/typegraph"; import { graph } from "../src/graph"; // your graph definition describe("Person queries", () => { let store: ReturnType>; beforeEach(() => { const { backend } = createLocalSqliteBackend(); store = createStore(graph, backend); }); it("creates and retrieves a person", async () => { const alice = await store.nodes.Person.create({ name: "Alice", email: "alice@example.com", }); const found = await store.nodes.Person.getById(alice.id); expect(found?.props.name).toBe("Alice"); }); }); ``` No teardown is needed — the in-memory database is garbage collected when the backend goes out of scope. ### Shared test helper If many test files use the same setup, extract a helper: ```typescript // tests/test-helpers.ts import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; import { createStore } from "@nicia-ai/typegraph"; import { graph } from "../src/graph"; export function createTestStore() { const { backend } = createLocalSqliteBackend(); return createStore(graph, backend); } ``` ```typescript // tests/person.test.ts import { beforeEach, describe, expect, it } from "vitest"; import { createTestStore } from "./test-helpers"; describe("Person", () => { let store: ReturnType; beforeEach(() => { store = createTestStore(); }); // tests... }); ``` ### createStore vs createStoreWithSchema | Factory | Sync? | Schema management | Use for | |---------|-------|-------------------|---------| | `createStore(graph, backend)` | Yes | None | Tests, local dev (no `searchable()` fields) | | `createStoreWithSchema(graph, backend)` | No | Auto-init, auto-migrate | Production, fulltext, schema evolution tests | Use `createStore` for most tests — it's synchronous and avoids async setup. Use `createStoreWithSchema` when you're specifically testing schema migrations or evolution, **or whenever the graph under test has `searchable()` fields**: a fulltext write/search through a bare `createStore()` test throws `StoreNotInitializedError` because the durable fulltext marker was never written. ```typescript import { createStoreWithSchema } from "@nicia-ai/typegraph"; it("migrates from v1 to v2", async () => { const { backend } = createLocalSqliteBackend(); // Initialize with v1 schema const [storeV1] = await createStoreWithSchema(graphV1, backend); await storeV1.nodes.Person.create({ name: "Alice" }); // Migrate to v2 schema const [storeV2, result] = await createStoreWithSchema(graphV2, backend); expect(result.status).toBe("migrated"); }); ``` ## Testing Queries ### Seed data, then query The typical pattern is: create data through the collection API, then verify queries return the expected results. ```typescript it("finds friends-of-friends", async () => { // Seed const alice = await store.nodes.Person.create({ name: "Alice" }); const bob = await store.nodes.Person.create({ name: "Bob" }); const carol = await store.nodes.Person.create({ name: "Carol" }); await store.edges.knows.create(alice, bob, {}); await store.edges.knows.create(bob, carol, {}); // Query const fof = await store .query() .from("Person", "p") .whereNode("p", (p) => p.id.eq(alice.id)) .traverse("knows", "e") .recursive({ minHops: 2, maxHops: 2 }) .to("Person", "friend") .select((ctx) => ctx.friend.name) .execute(); expect(fof).toEqual(["Carol"]); }); ``` ### Bulk seeding For tests that need a larger dataset, use `bulkCreate` for speed: ```typescript beforeEach(async () => { const people = Array.from({ length: 100 }, (_, i) => ({ props: { name: `Person ${i}`, email: `person${i}@example.com` }, })); await store.nodes.Person.bulkCreate(people); }); ``` ### Testing query shapes with toSQL() You can inspect the generated SQL without executing to verify query structure: ```typescript it("compiles a traversal to a single statement", () => { const query = store .query() .from("Person", "p") .traverse("worksAt", "e") .to("Company", "c") .select((ctx) => ({ person: ctx.p.name, company: ctx.c.name })); const { sql } = query.toSQL(); expect(sql).toContain("WITH"); expect(sql).not.toContain(";"); // single statement }); ``` ### Testing prepared queries ```typescript it("executes prepared queries with different bindings", async () => { await store.nodes.Person.create({ name: "Alice" }); await store.nodes.Person.create({ name: "Bob" }); const prepared = store .query() .from("Person", "p") .whereNode("p", (p) => p.name.eq(p.name.bind("targetName"))) .select((ctx) => ctx.p.name) .prepare(); const alice = await prepared.execute({ targetName: "Alice" }); const bob = await prepared.execute({ targetName: "Bob" }); expect(alice).toEqual(["Alice"]); expect(bob).toEqual(["Bob"]); }); ``` ## Testing Transactions Verify atomicity by asserting that failed transactions leave no partial data: ```typescript it("rolls back on error", async () => { try { await store.transaction(async (tx) => { await tx.nodes.Person.create({ name: "Alice" }); throw new Error("abort"); }); } catch { // expected } const all = await store .query() .from("Person", "p") .select((ctx) => ctx.p) .execute(); expect(all).toHaveLength(0); // Alice was rolled back }); ``` ## Testing with the Query Profiler Use the [Query Profiler](/performance/profiler) in tests to catch unindexed filter patterns before they reach production. ```typescript import { QueryProfiler } from "@nicia-ai/typegraph/profiler"; import { toDeclaredIndexes } from "@nicia-ai/typegraph/indexes"; import { personEmail } from "../src/indexes"; describe("Index coverage", () => { it("all query filters have index coverage", async () => { const profiler = new QueryProfiler({ declaredIndexes: toDeclaredIndexes([personEmail]), }); const profiledStore = profiler.attachToStore(store); // Run representative queries await profiledStore .query() .from("Person", "p") .whereNode("p", (p) => p.email.eq("alice@example.com")) .select((ctx) => ctx.p.name) .execute(); // Fails if any filter property lacks an index profiler.assertIndexCoverage(); }); }); ``` This is particularly effective when run against your full test suite — it catches filter patterns across all tests, not just the ones you remember to check manually. ## PGlite Tests (Postgres dialect, no Docker) Use `createLocalPgliteBackend()` when you want the PostgreSQL dialect and pgvector path in ordinary test runs without starting Docker. It creates an in-process PGlite engine, runs TypeGraph DDL, and returns a backend that should be closed after each test. ```typescript import { afterEach, beforeEach, describe, expect, it } from "vitest"; import { createLocalPgliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres/pglite"; import { createStore } from "@nicia-ai/typegraph"; import { graph } from "../src/graph"; describe("Postgres dialect behavior", () => { let cleanup: (() => Promise) | undefined; let store: ReturnType>; beforeEach(async () => { const { backend } = await createLocalPgliteBackend(); cleanup = () => backend.close(); store = createStore(graph, backend); }); afterEach(async () => { await cleanup?.(); }); it("runs against the Postgres query compiler", async () => { const alice = await store.nodes.Person.create({ name: "Alice" }); expect(await store.nodes.Person.getById(alice.id)).toBeDefined(); }); }); ``` PGlite is a good fit for SQL dialect coverage, pgvector behavior, and embedded Postgres workflows. Use a real PostgreSQL server for driver-specific behavior such as node-postgres named statements, postgres-js, pgbouncer, connection pool behavior, isolation levels, and true concurrent sessions. ## PostgreSQL Integration Tests For tests that verify PostgreSQL-specific behavior (JSONB operators, GIN indexes, concurrent writes), connect to a real database: ```typescript import { Pool } from "pg"; import { drizzle } from "drizzle-orm/node-postgres"; import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; describe("PostgreSQL integration", () => { let pool: Pool; let store: ReturnType>; beforeAll(async () => { pool = new Pool({ connectionString: process.env.TEST_DATABASE_URL }); await pool.query(generatePostgresMigrationSQL()); const db = drizzle(pool); const backend = createPostgresBackend(db); store = createStore(graph, backend); }); afterAll(async () => { await pool.end(); }); beforeEach(async () => { await pool.query("TRUNCATE typegraph_nodes, typegraph_edges CASCADE"); }); it("handles concurrent writes", async () => { const creates = Array.from({ length: 100 }, (_, i) => store.nodes.Person.create({ name: `Person ${i}` }), ); await Promise.all(creates); const count = await store.nodes.Person.count(); expect(count).toBe(100); }); }); ``` ### Skipping when no database is available Guard PostgreSQL tests so they're skipped in environments without a database: ```typescript const describePostgres = process.env.TEST_DATABASE_URL ? describe : describe.skip; describePostgres("PostgreSQL-specific", () => { // ... }); ``` ## Testing Pyramid | Level | Backend | Speed | Isolation | When to use | |-------|---------|-------|-----------|-------------| | Unit | In-memory SQLite | Fast (~1ms setup) | Full (fresh DB per test) | Collection API, query logic, business rules | | Integration | PGlite | Fast-medium | Full (fresh DB per test) | Postgres SQL dialect, pgvector, embedded workflows | | Integration | SQLite file or real PostgreSQL | Medium | Shared (truncate between tests) | Concurrency, driver behavior, isolation levels | | Profiler | In-memory SQLite | Fast | Full | Index coverage, query pattern verification | Most tests should be unit tests with in-memory SQLite. Use PGlite when you need the Postgres dialect or pgvector path in normal test runs. Reserve real PostgreSQL integration tests for behavior PGlite cannot model, such as connection pooling, driver quirks, concurrent sessions, and isolation levels. ## Next Steps - [Backend Setup](/backend-setup) — Configure SQLite and PostgreSQL backends - [Query Profiler](/performance/profiler) — Automatic index recommendations - [Schemas & Stores](/schemas-stores) — Collection API reference # Types > TypeScript type definitions and utilities This reference documents TypeGraph's TypeScript types and utility functions. ## Node Types ### `Node` The full node type returned from store operations. ```typescript type Node = Readonly<{ id: NodeId; // Branded ID type kind: N["kind"]; // Node kind name meta: { version: number; // Monotonic version counter validFrom: string | undefined; // Temporal validity start (ISO string) validTo: string | undefined; // Temporal validity end (ISO string) createdAt: string; // Created timestamp (ISO string) updatedAt: string; // Updated timestamp (ISO string) deletedAt: string | undefined; // Soft delete timestamp (ISO string) }; }> & z.infer; // Schema properties are flattened ``` ### `NodeId` Branded string type for type-safe node IDs. Prevents accidentally mixing IDs from different node types. ```typescript type NodeId = string & { readonly [__nodeId]: N }; ``` **Example:** ```typescript import { type NodeId } from "@nicia-ai/typegraph"; type PersonId = NodeId; type CompanyId = NodeId; function getPersonById(id: PersonId): Promise> { // TypeScript prevents passing a CompanyId here return store.nodes.Person.getById(id); } ``` ### `NodeProps` Extracts just the property types from a node definition. Use this when you only need the schema data without node metadata. ```typescript type NodeProps = z.infer; ``` **Example:** ```typescript import { type NodeProps } from "@nicia-ai/typegraph"; type PersonProps = NodeProps; // { name: string; email?: string; age?: number } // Useful for form data, API payloads, or validation function validatePersonData(data: PersonProps): boolean { return data.name.length > 0; } ``` ### `NodeRef` Type-safe reference to a node of a specific kind. Used for edge collection methods to enforce that endpoints match the allowed node types. Defaults to `NodeType` when no type parameter is given. ```typescript type NodeRef = Node | Readonly<{ kind: N["kind"]; id: string }>; ``` Accepts either: - A `Node` instance (e.g., the result of `store.nodes.Person.create()`) - An explicit object with the correct type name and ID ### `SelectableNode` The node type available in `select()` context. Properties are flattened (not nested under `props`). ```typescript type SelectableNode = Readonly<{ id: NodeId; kind: N["kind"]; meta: { version: number; validFrom: string | undefined; validTo: string | undefined; createdAt: string; updatedAt: string; deletedAt: string | undefined; }; }> & z.infer; // Properties are flattened ``` `id` carries the same `NodeId` brand as `Node`, so a projected id can be passed straight into `getById`/`getByIds` without a cast. **Example:** ```typescript // In select context, access properties directly .select((ctx) => ({ id: ctx.p.id, // NodeId name: ctx.p.name, // Direct property access (not ctx.p.props.name) email: ctx.p.email, created: ctx.p.meta.createdAt, })) ``` ## Edge Types ### `Edge` The full edge type returned from store operations. The `From` and `To` type parameters carry compile-time node type information for the edge endpoints. ```typescript type Edge< E extends EdgeType = EdgeType, From extends NodeType = NodeType, To extends NodeType = NodeType, > = Readonly<{ id: EdgeId; // Branded ID type kind: E["kind"]; fromKind: From["kind"]; fromId: NodeId; toKind: To["kind"]; toId: NodeId; meta: { validFrom: string | undefined; // Temporal validity start (ISO string) validTo: string | undefined; // Temporal validity end (ISO string) createdAt: string; // Created timestamp (ISO string) updatedAt: string; // Updated timestamp (ISO string) deletedAt: string | undefined; // Soft delete timestamp (ISO string) }; }> & z.infer; // Schema properties are flattened ``` ### `EdgeId` Branded string type for type-safe edge IDs. Prevents accidentally mixing IDs from different edge types. ```typescript type EdgeId = string & { readonly [__edgeId]: E }; ``` **Example:** ```typescript import { type EdgeId } from "@nicia-ai/typegraph"; type WorksAtId = EdgeId; function getEdgeById(id: WorksAtId): Promise> { return store.edges.worksAt.getById(id); } ``` ### `EdgeProps` Extracts just the property types from an edge definition. ```typescript type EdgeProps = z.infer; ``` **Example:** ```typescript import { type EdgeProps } from "@nicia-ai/typegraph"; type WorksAtProps = EdgeProps; // { role: string; startDate?: string } ``` ### `SelectableEdge` The edge type available in `select()` context. Properties are flattened. ```typescript type SelectableEdge = Readonly<{ id: string; kind: E["kind"]; fromId: string; toId: string; meta: { validFrom: string | undefined; validTo: string | undefined; createdAt: string; updatedAt: string; deletedAt: string | undefined; }; }> & z.infer; // Edge properties are flattened ``` `traverse()` defaults to `expand: "inverse"`, which can match the graph's *registered inverse* edge kind alongside the one you asked for — and `inverseOf(edgeA, edgeB)` doesn't require `edgeA`/`edgeB` to share a props schema — so the row behind an edge alias isn't guaranteed to be the requested kind, or to have the requested kind's schema. This affects `SelectableEdge` in three different ways: - **`kind: E["kind"]` can already be wrong today.** It's a literal type (e.g. `"manages"`), not `string` — but the runtime value can be the registered inverse kind (e.g. `"managedBy"`) under the default expansion mode. This isn't a missing brand, it's an existing type-accuracy gap: don't trust `ctx.e.kind` without knowing the traversal can't have expanded into a different kind. - **The flattened schema properties have the same gap.** `ctx.e.role` is typed against the requested edge kind's schema, but an inverse-branch row's real props came from a different schema and may not have a `role` field at all — reading it returns `undefined`, not a type error. - **`id`/`fromId`/`toId` stay plain `string`**, unlike `SelectableNode.id` — deliberately not branded `EdgeId`/`NodeId`/`NodeId`. This one's just an ergonomics gap (`string` never overclaims), but branding these fields would compile while being actively wrong for the same reason: a mismatched-kind id would compile straight into `getById` and silently return `undefined` instead of erroring. When you know a traversal is single-kind, re-brand explicitly: ```typescript import { asEdgeId, asNodeId } from "@nicia-ai/typegraph"; const rows = await store .query() .from("Person", "p") .traverse("worksAt", "e", { expand: "none" }) .to("Company", "c") .select((ctx) => ({ edgeId: ctx.e.id, companyId: ctx.e.toId })) .execute(); const edge = await store.edges.worksAt.getById(asEdgeId(rows[0]!.edgeId)); const company = await store.nodes.Company.getById(asNodeId(rows[0]!.companyId)); ``` **Example:** ```typescript // Access edge properties in select context. expand: "none" makes the // schema (and kind/id) trustworthy — see the warning above. .traverse("worksAt", "e", { expand: "none" }) .select((ctx) => ({ role: ctx.e.role, // Direct edge property access salary: ctx.e.salary, edgeId: ctx.e.id, startedAt: ctx.e.meta.createdAt, })) ``` ### `TypedEdgeCollection` A type-safe edge collection derived from the edge registration. This is what `store.edges.*` returns. With array-valued `to`, each source type can connect to each target type. With a [source-dependent target map](/core-concepts#source-dependent-targets), typed writes preserve the allowed pairs instead of accepting independent endpoint unions. ```typescript // assignedTo allows Employee -> Department and Student -> Course. await store.edges.assignedTo.create(employee, department, {}); await store.edges.assignedTo.create(student, course, {}); // @ts-expect-error Employee cannot be assigned to Course. await store.edges.assignedTo.create(employee, course, {}); ``` Keep the inferred edge and graph types to retain this information. Widening a declaration to a generic endpoint union can lose compile-time correlation; runtime validation still enforces the graph's allowed pairs. The same runtime validation applies to dynamic collections whose kinds are only known at runtime. This write-side guarantee does not make query results a correlated union of source/target pairs; narrow result kinds explicitly when consuming them. ### `DynamicNodeCollection` A node collection with widened generics for runtime string-keyed access via [`store.getNodeCollection(kind)`](/schemas-stores#storegetnodecollectionkind). Exposes the full `NodeCollection` API (`create`, `getById`, `find`, `count`, `createFromRecord`, etc.) but accepts `Record` for schema-typed parameters since the concrete node type is not known at compile time. ID parameters (`getById`, `getByIds`, `update`, `delete`, `hardDelete`, `bulkDelete`) accept plain `string` instead of branded `NodeId`, since the dynamic path typically receives IDs from edge metadata, snapshots, or external input where the brand is not available. ```typescript import type { DynamicNodeCollection } from "@nicia-ai/typegraph"; // Derived from NodeCollection with ID parameters widened to string ``` ### `DynamicNodeKind`, `DynamicNode`, and `DynamicNodeReference` `DynamicNodeKind` preserves and nominally marks the requested collection key. `DynamicNode` is the node value returned by a `DynamicNodeCollection`, and `DynamicNodeReference` is the nominal lightweight `{ kind, id }` form returned by runtime-aware identity reads. The markers let those results flow back into identity operations without making arbitrary string kinds valid compile-time inputs. ```typescript import type { DynamicNode, DynamicNodeKind, DynamicNodeReference, } from "@nicia-ai/typegraph"; ``` Identity reads return `IdentityNodeReference`, the union of the graph's compile-time node references and `DynamicNodeReference`, because a class can contain both after runtime evolution. ### `DynamicEdgeCollection` `DynamicEdgeCollection` is an edge collection for runtime endpoint dispatch. Obtain it from [`store.getEdgeCollection(kind)`](/schemas-stores#storegetedgecollectionkind), `store.getEdgeCollectionOrThrow(kind)`, or either method on a transaction context. When `kind` belongs to the graph's TypeScript definition, `E` retains that edge's property schema and result type. An arbitrary `string` uses the default `AnyEdgeType`, so properties are checked at runtime. Endpoints accept `{ kind: string; id: string }`. Writes validate endpoint domains and source-dependent pairs against the graph's schema. ID parameters accept plain `string`; returned edges retain their typed IDs. This surface exposes the full collection API, including bulk operations. Use it for helpers generic over a graph and edge kind: ```typescript import type { EdgeKinds, GraphDef, NodeRef, TransactionContext } from "@nicia-ai/typegraph"; import type { z } from "zod"; async function connect>( tx: TransactionContext, kind: K, from: NodeRef, to: NodeRef, props: z.input, ) { return tx.getEdgeCollectionOrThrow(kind).getOrCreateByEndpoints( from, to, props, { ifExists: "update" }, ); } ``` For a reusable collection parameter, use `DynamicEdgeCollection`. No collection cast or dependency on generated declaration filenames is needed. **Upgrading from 0.54/0.55:** source-dependent targets in 0.55 made endpoint arguments a union of valid pairs. TypeScript cannot resolve that union inside some generic `G`/`K` helpers, even for array-valued targets. Migrate dynamic calls from `tx.edges[kind]` to `tx.getEdgeCollectionOrThrow(kind)` (or check the optional lookup result). Generic `EdgeCollection` and `TypedEdgeCollection>` annotations can encounter the same deferred-type limitation; use `DynamicEdgeCollection` when endpoints are runtime data. Keep `store.edges.specificKind` for compile-time endpoint checking. `EdgeRegistration` now permits array or map targets in broad annotations. Code inspecting `.to` must narrow with `isEdgeTargetMap`; an explicitly Cartesian registration can specify `readonly To[]` as its fourth type argument. Known-kind dynamic lookups now check properties at compile time. Callers passing an unvalidated record should parse it with the selected schema first. An `unknown` property value was not accepted by the typed 0.54 API either. ### `DynamicStoreViewEdgeCollection` `view.getEdgeCollection(kind)` returns an optional dynamic collection bound to the view's valid-time coordinate. It retains known edge property types and accepts runtime endpoint references and string IDs. It exposes only the view's reads: there are no writes, deferred batch reads, or per-call temporal overrides. Recorded-time views retain their narrower reconstructing-read API. ## Subgraph Types These types are used with [`store.subgraph()`](/schemas-stores#storesubgraphrootid-options) for typed neighborhood extraction. ### `AnyNode` Discriminated union of all runtime node types in a graph. Each member carries its own `kind` literal, so `switch (node.kind)` narrows the type automatically. ```typescript import type { AnyNode } from "@nicia-ai/typegraph"; type MyNode = AnyNode; // = Node | Node | ... ``` ### `AnyEdge` Discriminated union of all runtime edge types in a graph. ```typescript import type { AnyEdge } from "@nicia-ai/typegraph"; type MyEdge = AnyEdge; // = Edge | Edge | ... ``` ### `SubsetNode` Narrows `AnyNode` to a subset of node kinds. Useful when `store.subgraph()` is called with `includeKinds`. ```typescript import type { SubsetNode } from "@nicia-ai/typegraph"; type TaskOrAgent = SubsetNode; // = Node | Node ``` ### `SubsetEdge` Narrows `AnyEdge` to a subset of edge kinds. ```typescript import type { SubsetEdge } from "@nicia-ai/typegraph"; type TraversedEdges = SubsetEdge; ``` ### `SubgraphOptions` Options for `store.subgraph()`. See the [store reference](/schemas-stores#storesubgraphrootid-options) for the full parameter table. ```typescript type SubgraphOptions = Readonly<{ edges: readonly EK[]; maxDepth?: number; includeKinds?: readonly NK[]; excludeRoot?: boolean; direction?: "out" | "both"; cyclePolicy?: "prevent" | "allow"; }>; ``` ### `SubgraphResult` The return type of `store.subgraph()`. Contains the root node, a node index, and forward/reverse adjacency maps for immediate traversal. ```typescript type SubgraphResult = Readonly<{ root: SubgraphNodeResult | undefined; nodes: ReadonlyMap>; adjacency: ReadonlyMap[]>>; reverseAdjacency: ReadonlyMap[]>>; }>; ``` ## Graph Configuration Types ### `DeleteBehavior` Controls what happens to edges when a node is deleted. ```typescript type DeleteBehavior = "restrict" | "cascade" | "disconnect"; ``` | Value | Description | |-------|-------------| | `"restrict"` | Prevent deletion if edges exist | | `"cascade"` | Delete connected edges | | `"disconnect"` | Remove edges without error | ### `Cardinality` Controls how many edges of a type can connect from/to a node. ```typescript type Cardinality = "many" | "one" | "unique" | "oneActive"; ``` | Value | Description | |-------|-------------| | `"many"` | No limit on edges | | `"one"` | At most one edge per source node | | `"unique"` | At most one edge per source-target pair | | `"oneActive"` | At most one active edge (`validTo` is `undefined`) per source node | ### `InferenceType` Controls how ontology relationships affect queries. ```typescript type InferenceType = | "subsumption" // Query for X includes subclass instances | "hierarchy" // Enables broader/narrower traversal | "substitution" // Can substitute equivalent types | "constraint" // Validation rules | "composition" // Part-whole navigation | "association" // Discovery/recommendation | "none"; // No automatic inference ``` ## Query Types ### `VariableLengthSpec` Configuration for variable-length (recursive) traversals. ```typescript type VariableLengthSpec = Readonly<{ minDepth: number; // Minimum hops (default: 1) maxDepth: number; // Maximum hops (-1 = unlimited) cyclePolicy: "prevent" | "allow"; // Cycle handling mode pathAlias?: string; // Column alias for projected path depthAlias?: string; // Column alias for projected depth }>; ``` ### `SetOperationType` Available set operations for combining queries. ```typescript type SetOperationType = "union" | "unionAll" | "intersect" | "except"; ``` ### `PaginateOptions` Options for cursor-based pagination. ```typescript type PaginateOptions = Readonly<{ first?: number; // Items to fetch (forward) after?: string; // Cursor to start after (forward) last?: number; // Items to fetch (backward) before?: string; // Cursor to start before (backward) }>; ``` ### `PaginatedResult` Result of a paginated query. ```typescript type PaginatedResult = Readonly<{ data: readonly R[]; nextCursor: string | undefined; prevCursor: string | undefined; hasNextPage: boolean; hasPrevPage: boolean; }>; ``` ### `StreamOptions` Options for streaming results. ```typescript type StreamOptions = Readonly<{ batchSize?: number; // Items per batch (default: 1000) }>; ``` ## Utility Functions ### `generateId()` Generates a unique ID using nanoid. ```typescript import { generateId } from "@nicia-ai/typegraph"; function generateId(): string; const id = generateId(); // "V1StGXR8_Z5jdHi6B-myT" ``` ## Constants ### `MAX_RECURSIVE_DEPTH` Maximum depth for unbounded recursive traversals (10). ```typescript import { MAX_RECURSIVE_DEPTH } from "@nicia-ai/typegraph"; // MAX_RECURSIVE_DEPTH = 10 ``` Recursive traversals are capped at this depth when no `maxHops` is specified in the `recursive()` options object. Explicit `maxHops` values are validated against `MAX_EXPLICIT_RECURSIVE_DEPTH` (1000). Cycle prevention is enabled by default. To allow revisits for maximum performance, use `cyclePolicy: "allow"`. ### `MAX_EXPLICIT_RECURSIVE_DEPTH` Maximum allowed value for the `maxHops` option in recursive traversals (1000). ```typescript import { MAX_EXPLICIT_RECURSIVE_DEPTH } from "@nicia-ai/typegraph"; // MAX_EXPLICIT_RECURSIVE_DEPTH = 1000 ``` # Ejecting > How to migrate away from TypeGraph if you need to TypeGraph is designed with zero lock-in. If you decide to move on, you're left with a clean, conventional database schema that works with any SQL tooling. ## What You're Left With When you eject TypeGraph, your database contains two well-structured tables: ```sql -- Your nodes SELECT * FROM typegraph_nodes; ┌──────────┬─────────┬──────────────────┬─────────────────────────────────┐ │ kind │ id │ props │ created_at │ ├──────────┼─────────┼──────────────────┼─────────────────────────────────┤ │ Person │ p-001 │ {"name": "Ada"} │ 2024-01-15T10:30:00Z │ │ Company │ c-001 │ {"name": "Acme"} │ 2024-01-15T10:30:00Z │ └──────────┴─────────┴──────────────────┴─────────────────────────────────┘ -- Your relationships SELECT * FROM typegraph_edges; ┌──────────┬──────────┬─────────┬──────────┬─────────┐ │ kind │ from_id │ to_id │ props │ ... │ ├──────────┼──────────┼─────────┼──────────┼─────────┤ │ worksAt │ p-001 │ c-001 │ {} │ ... │ └──────────┴──────────┴─────────┴──────────┴─────────┘ ``` This is exactly the schema you'd design yourself for a flexible entity-relationship system. ## Querying Without TypeGraph All your data is accessible with plain SQL. No special drivers, no proprietary formats. ### Find all people at a company ```sql SELECT n.props->>'name' as person_name FROM typegraph_nodes n JOIN typegraph_edges e ON e.from_id = n.id WHERE e.kind = 'worksAt' AND e.to_id = 'c-001' AND n.deleted_at IS NULL; ``` ### Traverse a relationship ```sql SELECT p.props->>'name' as person, c.props->>'name' as company FROM typegraph_nodes p JOIN typegraph_edges e ON e.from_id = p.id AND e.kind = 'worksAt' JOIN typegraph_nodes c ON c.id = e.to_id WHERE p.kind = 'Person' AND c.kind = 'Company' AND p.deleted_at IS NULL AND c.deleted_at IS NULL; ``` ### Point-in-time query ```sql SELECT * FROM typegraph_nodes WHERE kind = 'Article' AND valid_from <= '2024-06-01' AND (valid_to IS NULL OR valid_to > '2024-06-01'); ``` ## Using Your Own Tools The schema works with everything in the SQL ecosystem: - **ORMs**: Drizzle, Prisma, Knex, TypeORM, Sequelize - **Query builders**: Kysely, Slonik - **Raw SQL**: Any PostgreSQL or SQLite client - **BI tools**: Metabase, Superset, Tableau - **Migration tools**: dbmate, Flyway, Liquibase ### Example: Drizzle ORM ```typescript import { pgTable, text, jsonb, timestamp } from "drizzle-orm/pg-core"; const nodes = pgTable("typegraph_nodes", { graphId: text("graph_id").notNull(), kind: text("kind").notNull(), id: text("id").notNull(), props: jsonb("props").notNull(), createdAt: timestamp("created_at").notNull(), updatedAt: timestamp("updated_at").notNull(), deletedAt: timestamp("deleted_at"), }); // Query as usual const people = await db .select() .from(nodes) .where(eq(nodes.kind, "Person")); ``` ### Example: Prisma ```prisma model TypegraphNode { graphId String @map("graph_id") kind String id String props Json createdAt DateTime @map("created_at") deletedAt DateTime? @map("deleted_at") @@id([graphId, kind, id]) @@map("typegraph_nodes") } ``` ## Migration Strategies ### Option 1: Keep the schema as-is The TypeGraph schema is production-ready. Continue using it directly with your preferred SQL tools. ### Option 2: Normalize into separate tables If you want traditional per-entity tables: ```sql -- Create a typed table CREATE TABLE people AS SELECT id, props->>'name' as name, props->>'email' as email, created_at, updated_at FROM typegraph_nodes WHERE kind = 'Person' AND deleted_at IS NULL; -- Add constraints ALTER TABLE people ADD PRIMARY KEY (id); ``` ### Option 3: Create views for compatibility Keep the original tables but add typed views: ```sql CREATE VIEW people AS SELECT id, props->>'name' as name, props->>'email' as email, created_at FROM typegraph_nodes WHERE kind = 'Person' AND deleted_at IS NULL; ``` ## What About the Ontology? The ontology (type hierarchies, edge constraints) exists only in your TypeScript code. The database stores raw data without semantic constraints. After ejecting: - You lose automatic subclass queries (`includeSubClasses`) - You lose edge validation (ensuring valid from/to kinds) - You keep all your data exactly as stored If you need these features, you'll implement them in application code—which is what any alternative would require anyway. ## Summary TypeGraph adds a type-safe API layer over a conventional SQL schema. Remove the library and you still have: - Standard SQL tables - JSON properties (supported natively by SQLite and PostgreSQL) - Full temporal history - Soft deletes - No proprietary formats - No data migration required Your data is always yours. # Agent Decision Replay > Reconstruct the graph an agent actually saw and replay the same reasoning code This example shows recorded-time capture as an agent-debugging tool. An agent chooses the best source for a claim by traversing a knowledge graph and ranking papers by citation authority. Later, the graph changes: a paper is removed and new citations arrive. A live rerun gives a different answer, but `store.asOfRecorded(decisionTime)` reconstructs the exact graph the agent saw. :::tip[Just want the code?] Full source on GitHub: [`packages/typegraph/examples/21-agent-decision-replay.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/21-agent-decision-replay.ts) ::: ## What It Demonstrates - `history: true` capture on a knowledge graph used by an AI agent. - `store.recordedNow()` as the durable decision-time checkpoint. - One reasoning function that runs unchanged against live and recorded views. - A sealed recorded `query()` plus `degree()` graph algorithm over reconstructed history. - Why live reruns cannot explain old decisions once the evidence graph changes. ## Run It From the repository root: ```bash pnpm --filter @nicia-ai/typegraph exec tsx examples/21-agent-decision-replay.ts ``` Or from `packages/typegraph`: ```bash npx tsx examples/21-agent-decision-replay.ts ``` ## The Reusable Shape The agent takes a read surface, not a concrete store. A live `StoreView` and a recorded `RecordedStoreView` both satisfy the parts it needs: ```typescript type EvidenceView = Pick< RecordedStoreView, "query" | "degree" | "nodes" >; async function recommendSource(view: EvidenceView, claimId: string) { const supporterIds = await view .query() .from("Claim", "c") .whereNode("c", (c) => c.id.eq(claimId)) .traverse("supports", "e", { direction: "in" }) .to("Paper", "p") .select((ctx) => ctx.p.id) .execute(); const scored = await Promise.all( supporterIds.map(async (id) => ({ id, citations: await view.degree(id, { edges: ["cites"], direction: "in", }), })), ); return scored.sort((a, b) => b.citations - a.citations)[0]; } ``` At decision time, save the recorded anchor: ```typescript const decisionTime = await store.recordedNow(); if (decisionTime === undefined) throw new Error("expected recorded history"); const original = await recommendSource( store.view({ mode: "current" }), claim.id, ); ``` During audit, replay the same code against the old graph: ```typescript const replay = await recommendSource( store.asOfRecorded(decisionTime), claim.id, ); ``` ## Sample Output ```text Agent answers: 'best source for the claim?' -> Kaplan - Scaling Laws (3 citations) recorded at: 2026-... Re-run on the CURRENT graph (the evidence moved): -> Hoffmann - Compute-Optimal LLMs (3 citations) x This is a different answer. Replay on `store.asOfRecorded(decisionTime)` (same code): -> Kaplan - Scaling Laws (3 citations) ok Reproduced exactly. ``` ## When to Use This Pattern Use recorded decision replay when agent output must be explainable after the graph changes: - Eval replay for graph-grounded agents - Audit trails for automated decisions - Debugging "why did the agent say this?" incidents - Reproducible RAG and knowledge-graph experiments See [Temporal queries](/queries/temporal#recorded-time-bitemporal) for the recorded-time rules and [Graph Algorithms](/graph-algorithms) for `degree()`. # Agent-Driven Schema > Runtime schema evolution end-to-end — an agent proposes new kinds and edges, the validator gates them, store.evolve commits, the live graph absorbs them with no restart and no codegen. A runnable end-to-end demonstration of [graph extensions](/graph-extensions) — the 0.25.0 feature that lets you grow the schema **at runtime** when something new shows up in the world. An agent (LLM, scraper, ETL pipeline) proposes new node and edge kinds from observed data, the validator gates the proposal, and `store.evolve()` commits the new schema as a durable version. The live graph starts ingesting under the new schema in the same process, with full Zod validation, fulltext indexing, unique constraints, and cross-kind edge enforcement — no restart, no codegen, no `any` cast on the read side. :::tip[Just want the code?] Full source on GitHub: [`packages/typegraph/examples/16-graph-extensions.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/16-graph-extensions.ts) ::: :::note[Want to see this against real data with a real LLM?] A separate demo repo runs the same loop against public-record clinical research data with an open-weight LLM proposing schema after each stage of the corpus arrives: [`pdlug/typegraph-clinical-demo`](https://github.com/pdlug/typegraph-clinical-demo). ::: ## The shape of the loop Runtime schema evolution is one verb (`store.evolve`) sitting between two deliberate sources of friction: 1. **The validator.** Every proposal — even one built in TypeScript by your own code — flows through `validateGraphExtension`, which returns structured `{path, code, message}` issues. Typos, unknown property refinements, malformed edge endpoints all come back as routable errors instead of throwing deep in `evolve`. 2. **The incompatibility classifier.** Once data exists, `evolve` rejects any change that would corrupt it — narrowing a type, removing a required field on a populated kind, dropping an edge with live rows. A misbehaving agent can't silently break the graph. You provide the third moving part — the operator gate. Nothing in TypeGraph auto-applies what an agent proposes. Your code decides whether to call `evolve` on a validated proposal; the library makes sure the call is safe when you do. ## What you get The example walks through nine steps against a fresh SQLite database, starting with a single compile-time kind (`Document`) and ending with three kinds the agent invented at runtime. Sample output: ```text [1] Booted with compile-time kind: Document Active schema version: 1 Materialized 1 compile-time index [2] Agent proposal validated: nodes: Paper, Author edges: authoredBy [3] evolve() committed and materialized Active schema version: 2 registry.hasNodeType('Paper'): true registry.hasEdgeType('authoredBy'): true [4] Ingested 2 Paper, 2 Author, 2 authoredBy edges fulltext("Paper", "transformer architecture") -> 1 hit(s) score=0.61 title="Attention is all you need" Duplicate doi rejected: true [5] Dynamic multi-hop traversal Paper -> Author: Language models are unsupervised multitask learners (#1) by Alec Radford [6] Incompatible re-proposal rejected: Paper.year: TYPE_CHANGE (number -> string) [7] Deprecated kinds: [Document] [8] removeKinds(['Author']) — cascading edge cleanup registry.hasNodeType('Author'): false registry.hasEdgeType('authoredBy'): false Active schema version: 5 [9] Restart parity: validation.status: VALID registry.hasNodeType('Paper'): true registry.hasNodeType('Author') (removed in step 7): false deprecatedKinds: [Document] Found 2 Paper nodes after restart ``` The interesting moments are step 4 (fulltext + unique-constraint enforcement against a runtime kind), step 6 (the incompatibility gate), and step 9 (everything survives a fresh `createStoreWithSchema` against the same database — kinds, deprecation flags, indexes). ## The boot graph Start with one compile-time kind so the live graph has structure before the agent shows up. Everything compile-time stays type-safe end-to-end even as extension kinds accumulate around it: ```typescript const Document = defineNode("Document", { schema: z.object({ title: z.string(), body: z.string(), }), }); const documentTitle = defineNodeIndex(Document, { fields: ["title"] }); const baseGraph = defineGraph({ id: "research_corpus", nodes: { Document: { type: Document } }, edges: {}, indexes: [documentTitle], }); const { backend } = createLocalSqliteBackend(); const [store] = await createStoreWithSchema(baseGraph, backend); ``` ## Scene by scene ### 1. The agent returns a JSON proposal What an LLM or scraper actually hands back is a JSON document, not a typed value. `validateGraphExtension(unknown, { strict })` is the Result-style entry point: it walks the document, collects every structural issue, and returns a typed `GraphExtension` on success. Strict mode rejects unknown sibling keys so a field typo (`node` instead of `nodes`) fails loudly instead of silently producing an empty extension. ```typescript const agentJson: unknown = JSON.parse(`{ "nodes": { "Paper": { "description": "Academic paper inferred from the corpus", "properties": { "title": { "type": "string", "minLength": 1, "searchable": {} }, "abstract": { "type": "string", "searchable": {}, "optional": true }, "doi": { "type": "string", "minLength": 1 }, "year": { "type": "number", "int": true, "min": 1900, "max": 2100 } }, "unique": [{ "name": "paper_doi_unique", "fields": ["doi"] }] }, "Author": { "properties": { "name": { "type": "string", "minLength": 1, "searchable": {} }, "affiliation": { "type": "string", "optional": true } } } }, "edges": { "authoredBy": { "from": ["Paper"], "to": ["Author"], "properties": { "order": { "type": "number", "int": true, "min": 1 } } } }, "indexes": [ { "entity": "node", "kind": "Paper", "name": "paper_by_year", "fields": ["year"] } ] }`); const result = validateGraphExtension(agentJson, { strict: true }); if (!result.success) { // result.error.issues is a structured list — route it back to the agent. throw result.error; } const extension: GraphExtension = result.data; ``` Each issue carries a stable `code` (`UNKNOWN_DOCUMENT_KEY`, `INVALID_PROPERTY_REFINEMENT`, `MISSING_REQUIRED_FIELD`, …) — useful both for routing failures back to the model in a repair loop and for treating validation outcomes as data instead of exceptions. ### 2. Commit + materialize atomically `evolve` runs the incompatibility check, commits a new schema version, and returns a new `Store` carrying the merged registry. `eager: {}` turns it into a one-call "schema committed AND indexes materialized" verb so the returned store is ready to ingest by the time the promise resolves: ```typescript const evolved = await store.evolve(extension, { eager: {} }); ``` ### 3. Ingest under the new schema Extension kinds get the same CRUD surface as compile-time kinds — there's no separate dynamic API to learn. The difference is at the type system: extension kinds aren't visible at compile time, so the collection's type is the generic `NodeCollection` rather than a kind-specific one. Property validation, unique constraints, fulltext indexing, and edge endpoint enforcement all run live: ```typescript const papers = evolved.getNodeCollectionOrThrow("Paper"); const authors = evolved.getNodeCollectionOrThrow("Author"); const authoredBy = evolved.getEdgeCollectionOrThrow("authoredBy"); const attention = await papers.create({ title: "Attention is all you need", abstract: "We propose a new simple network architecture, the Transformer.", doi: "10.5555/3295222.3295349", year: 2017, }); await authoredBy.create(attention, vaswani, { order: 1 }); // The `searchable: {}` brand on Paper.title flows through to the // backend's fulltext index — same BM25 retrieval extension kinds get // at compile time. const hits = await evolved.search.fulltext("Paper", { query: "transformer architecture", limit: 3, }); // Unique-constraint enforcement is live for extension kinds too. const duplicate = await papers.create({ title: "Duplicate doi", doi: "10.5555/3295222.3295349", // same doi as `attention` above year: 2024, }).catch((err) => err); // → UniqueConstraintError ``` ### 4. Multi-hop traversals over runtime kinds The typed query builder methods (`from`, `traverse`, `to`) require compile-time kind literals so they can give you full intellisense. For runtime kinds, use the `*Dynamic` siblings — they accept arbitrary strings, validate them against the registry (typo → `KindNotFoundError`), and surface a `.field("name").number().gte(...)` predicate API so an MCP server can traverse a runtime graph without an `as any`: ```typescript const rows = await evolved .query() .fromDynamic("Paper", "p") .traverseDynamic("authoredBy", "a") .toDynamic("Author", "u") .whereNode("p", (p) => p.field("year").number().gte(2018)) .select((ctx) => ({ paperTitle: ctx.p.title, authorName: ctx.u.name, order: ctx.a.order, })) .execute(); ``` Dynamic and typed methods mix freely — `from("Document", ...)` chained to `traverseDynamic("authoredBy", ...)` is well-formed. ### 5. The incompatibility gate The agent (or any caller) eventually proposes something that would corrupt existing data. Against an empty kind some changes are allowed; against a populated one, the classifier rejects them: ```typescript const breakingProposal: GraphExtension = { ...extension, nodes: { ...extension.nodes, Paper: { ...extension.nodes!.Paper!, properties: { ...extension.nodes!.Paper!.properties, year: { type: "string" }, // was number — TYPE_CHANGE }, }, }, }; const rejection = await evolved .evolve(breakingProposal) .catch((err) => err); if (rejection instanceof IncompatibleChangeError) { for (const change of rejection.changes) { console.log(`${change.kind}.${change.field}: ${change.type}`); } // → Paper.year: TYPE_CHANGE } ``` This is the "data corruption" backstop. A misbehaving agent in a tight loop cannot evolve the schema into a shape that breaks existing rows. ### 6. Deprecate and remove Deprecation is a **signal**, not a gate. It surfaces in `store.introspect().deprecatedKinds` for codegen tools and lint rules, but reads and writes against the deprecated kind continue to work: ```typescript const deprecated = await evolved.deprecateKinds(["Document"]); // Document is now flagged but still fully usable. await deprecated.nodes.Document.create({ title: "Legacy doc", body: "Still readable, just flagged", }); ``` Removal is the harder verb. `removeKinds(["Author"], { eager: {} })` commits a new schema version that drops `Author` and cascades to any edge whose endpoints depend on it (here, `authoredBy`). With `eager: {}` the data-cleanup phase runs inline — rows and edge data are deleted before the verb returns: ```typescript const trimmed = await deprecated.removeKinds(["Author"], { eager: {} }); // registry.hasNodeType('Author') -> false // registry.hasEdgeType('authoredBy') -> false (cascaded) ``` ### 7. Restart parity The whole point of "durable schema versions" is that nothing above is in-memory state. A fresh process opening the same database sees every accepted extension, deprecation flag, and materialized index without re-running any verb: ```typescript const [restored, validation] = await createStoreWithSchema(baseGraph, backend); // validation.status === "VALID" // restored.registry.hasNodeType("Paper") -> true // restored.registry.hasNodeType("Author") -> false (removed in step 6) // restored.introspect().deprecatedKinds -> ["Document"] const papers = restored.getNodeCollectionOrThrow("Paper"); await papers.find({}); // returns the rows from step 3 ``` This is what makes runtime evolution safe for production: agent-proposed schema is regular schema by the time the next process starts up. ## Run it ```bash git clone https://github.com/nicia-ai/typegraph cd typegraph pnpm install npx tsx packages/typegraph/examples/16-graph-extensions.ts ``` The example builds the graph, walks every step against an in-memory SQLite database, and prints output for each. To persist it, point `createLocalSqliteBackend()` at a file path. To run on Postgres, swap the import to `createPostgresBackend` — see [Backend Setup](/backend-setup). ## Next steps - [Graph Extensions](/graph-extensions) — the full reference for `defineGraphExtension`, `validateGraphExtension`, `evolve`, `deprecateKinds`, `removeKinds`, and the structured-issue codes - [`pdlug/typegraph-clinical-demo`](https://github.com/pdlug/typegraph-clinical-demo) — the same loop driven by an open-weight LLM against public-record clinical data, with a repair loop and a smoke-test pattern that catches latent shape mismatches the static validator can't - [Schema Migrations](/schema-management) — the lower-level primitives graph extensions ride on top of (`SchemaVersion`, change classification, reconciliation watermarks) - [Dynamic Queries](/queries/source) — `fromDynamic`, `traverseDynamic`, and the predicate accessor for runtime kinds # Audit Trail > Complete change tracking with user attribution and diff generation This example shows how to build a comprehensive audit system that: - **Tracks all changes** using TypeGraph's temporal model - **Attributes changes** to users and sessions - **Generates diffs** between versions - **Supports compliance queries** (who changed what, when) - **Exports audit logs** for external systems ## How TypeGraph Enables Auditing TypeGraph's temporal model provides built-in auditing capabilities: 1. **Every update creates a new version** - Old data is preserved with `valid_to` timestamp 2. **Temporal queries** - Query any point in time with `asOf` or get full history with `includeEnded` 3. **Metadata fields** - `createdAt`, `updatedAt`, `version` are tracked automatically This example extends the built-in capabilities with: - User attribution (who made the change) - Change descriptions (why the change was made) - Structured diffs (what exactly changed) ## Schema Definition ```typescript import { z } from "zod"; import { defineNode, defineEdge, defineGraph, searchable } from "@nicia-ai/typegraph"; // Audited entity (example: Settings) const Setting = defineNode("Setting", { schema: z.object({ key: z.string(), value: z.string(), category: z.string(), description: z.string().optional(), }), }); // Audit queries frequently include free-text predicates: "find all // changes mentioning GDPR" or "find changes where the reason contains // 'security patch'". Marking `reason` as `searchable()` enables BM25 // ranked fulltext over that column without an external search service. // Users making changes const User = defineNode("User", { schema: z.object({ email: z.string().email(), name: z.string(), role: z.enum(["admin", "editor", "viewer"]), }), }); // Explicit audit log entries (for cross-cutting concerns) const AuditEntry = defineNode("AuditEntry", { schema: z.object({ entityType: z.string(), entityId: z.string(), action: z.enum(["create", "update", "delete", "restore"]), timestamp: z.string().datetime(), changes: z.record(z.object({ before: z.unknown().optional(), after: z.unknown().optional(), })).optional(), reason: searchable({ language: "english" }).optional(), ipAddress: z.string().optional(), userAgent: z.string().optional(), }), }); // Sessions for grouping changes const Session = defineNode("Session", { schema: z.object({ startedAt: z.string().datetime(), endedAt: z.string().datetime().optional(), ipAddress: z.string().optional(), userAgent: z.string().optional(), }), }); // Edges const performedBy = defineEdge("performedBy"); const inSession = defineEdge("inSession"); const hasSession = defineEdge("hasSession"); const graph = defineGraph({ id: "audit_trail", nodes: { Setting: { type: Setting, unique: [ { name: "setting_key", fields: ["key"], scope: "kind", collation: "binary", }, ], }, User: { type: User, unique: [ { name: "user_email", fields: ["email"], scope: "kind", collation: "caseInsensitive", }, ], }, AuditEntry: { type: AuditEntry }, Session: { type: Session }, }, edges: { performedBy: { type: performedBy, from: [AuditEntry], to: [User] }, inSession: { type: inSession, from: [AuditEntry], to: [Session] }, hasSession: { type: hasSession, from: [User], to: [Session] }, }, }); ``` ## Audit Context Create a context object to track the current user and session: ```typescript interface AuditContext { userId: string; sessionId?: string; ipAddress?: string; userAgent?: string; reason?: string; } // Thread-local storage for audit context (Node.js) import { AsyncLocalStorage } from "async_hooks"; const auditContext = new AsyncLocalStorage(); function withAuditContext(context: AuditContext, fn: () => Promise): Promise { return auditContext.run(context, fn); } function getAuditContext(): AuditContext | undefined { return auditContext.getStore(); } ``` ## Audited Operations ### Create with Audit ```typescript async function createSetting( key: string, value: string, category: string ): Promise> { const ctx = getAuditContext(); if (!ctx) throw new Error("Audit context required"); return store.transaction(async (tx) => { // Create the setting const setting = await tx.nodes.Setting.create({ key, value, category, }); // Create audit entry await createAuditEntry(tx, { entityType: "Setting", entityId: setting.id, action: "create", changes: { key: { after: key }, value: { after: value }, category: { after: category }, }, }); return setting; }); } ``` ### Update with Audit ```typescript async function updateSetting( id: string, updates: Partial<{ value: string; description: string }> ): Promise> { const ctx = getAuditContext(); if (!ctx) throw new Error("Audit context required"); return store.transaction(async (tx) => { // Get current state const current = await tx.nodes.Setting.getById(id); if (!current) throw new Error(`Setting not found: ${id}`); // Calculate changes const changes: Record = {}; for (const [key, newValue] of Object.entries(updates)) { const oldValue = current[key as keyof typeof current]; if (oldValue !== newValue) { changes[key] = { before: oldValue, after: newValue }; } } // Skip if no actual changes if (Object.keys(changes).length === 0) { return current; } // Update the setting const updated = await tx.nodes.Setting.update(id, updates); // Create audit entry await createAuditEntry(tx, { entityType: "Setting", entityId: id, action: "update", changes, }); return updated; }); } ``` ### Delete with Audit ```typescript async function deleteSetting(id: string): Promise { const ctx = getAuditContext(); if (!ctx) throw new Error("Audit context required"); await store.transaction(async (tx) => { // Get current state for audit const current = await tx.nodes.Setting.getById(id); if (!current) throw new Error(`Setting not found: ${id}`); // Delete (soft delete) await tx.nodes.Setting.delete(id); // Create audit entry await createAuditEntry(tx, { entityType: "Setting", entityId: id, action: "delete", changes: { key: { before: current.key }, value: { before: current.value }, category: { before: current.category }, }, }); }); } ``` ### Create Audit Entry ```typescript interface AuditEntryInput { entityType: string; entityId: string; action: "create" | "update" | "delete" | "restore"; changes?: Record; } async function createAuditEntry( tx: Transaction, input: AuditEntryInput ): Promise> { const ctx = getAuditContext()!; const entry = await tx.nodes.AuditEntry.create({ entityType: input.entityType, entityId: input.entityId, action: input.action, timestamp: new Date().toISOString(), changes: input.changes, reason: ctx.reason, ipAddress: ctx.ipAddress, userAgent: ctx.userAgent, }); // Link to user const user = await tx.nodes.User.getById(ctx.userId); if (!user) throw new Error(`User not found: ${ctx.userId}`); await tx.edges.performedBy.create(entry, user, {}); // Link to session if present if (ctx.sessionId) { const session = await tx.nodes.Session.getById(ctx.sessionId); if (!session) throw new Error(`Session not found: ${ctx.sessionId}`); await tx.edges.inSession.create(entry, session, {}); } return entry; } ``` ## Querying Audit History ### Get Entity History ```typescript interface HistoryEntry { version: number; timestamp: string; action: string; changes?: Record; user: { name: string; email: string }; reason?: string; } async function getEntityHistory( entityType: string, entityId: string ): Promise { return store .query() .from("AuditEntry", "a") .whereNode("a", (a) => a.entityType.eq(entityType).and(a.entityId.eq(entityId)) ) .traverse("performedBy", "e") .to("User", "u") .orderBy((ctx) => ctx.a.timestamp, "desc") .select((ctx) => ({ version: ctx.a.version, timestamp: ctx.a.timestamp, action: ctx.a.action, changes: ctx.a.changes, user: { name: ctx.u.name, email: ctx.u.email, }, reason: ctx.a.reason, })) .execute(); } ``` ### Get User Activity ```typescript interface UserActivity { timestamp: string; entityType: string; entityId: string; action: string; } async function getUserActivity( userId: string, options: { since?: Date; limit?: number } = {} ): Promise { const { since, limit = 100 } = options; let query = store .query() .from("User", "u") .whereNode("u", (u) => u.id.eq(userId)) .traverse("performedBy", "e", { direction: "in" }) .to("AuditEntry", "a"); if (since) { query = query.whereNode("a", (a) => a.timestamp.gte(since.toISOString())); } return query .orderBy((ctx) => ctx.a.timestamp, "desc") .limit(limit) .select((ctx) => ({ timestamp: ctx.a.timestamp, entityType: ctx.a.entityType, entityId: ctx.a.entityId, action: ctx.a.action, })) .execute(); } ``` ### Changes in Time Range ```typescript interface ChangeReport { entityType: string; entityId: string; changeCount: number; users: string[]; lastChange: string; } async function getChangesInRange( startDate: Date, endDate: Date ): Promise { const entries = await store .query() .from("AuditEntry", "a") .whereNode("a", (a) => a.timestamp.gte(startDate.toISOString()).and( a.timestamp.lte(endDate.toISOString()) ) ) .traverse("performedBy", "e") .to("User", "u") .select((ctx) => ({ entityType: ctx.a.entityType, entityId: ctx.a.entityId, timestamp: ctx.a.timestamp, userName: ctx.u.name, })) .execute(); // Group by entity const grouped = new Map(); for (const entry of entries) { const key = `${entry.entityType}:${entry.entityId}`; const existing = grouped.get(key); if (existing) { existing.changeCount++; if (!existing.users.includes(entry.userName)) { existing.users.push(entry.userName); } if (entry.timestamp > existing.lastChange) { existing.lastChange = entry.timestamp; } } else { grouped.set(key, { entityType: entry.entityType, entityId: entry.entityId, changeCount: 1, users: [entry.userName], lastChange: entry.timestamp, }); } } return Array.from(grouped.values()); } ``` ## Using TypeGraph's Built-in Temporal Features ### View Entity at Point in Time ```typescript async function getSettingAsOf( id: string, timestamp: Date ): Promise { return store .query() .from("Setting", "s") .temporal("asOf", timestamp.toISOString()) .whereNode("s", (s) => s.id.eq(id)) .select((ctx) => ctx.s) .first(); } ``` ### Get All Versions ```typescript interface SettingVersion { props: SettingProps; validFrom: string; validTo: string | undefined; version: number; } async function getSettingVersions(id: string): Promise { return store .query() .from("Setting", "s") .temporal("includeEnded") .whereNode("s", (s) => s.id.eq(id)) .orderBy((ctx) => ctx.s.validFrom, "desc") .select((ctx) => ({ props: ctx.s, validFrom: ctx.s.validFrom, validTo: ctx.s.validTo, version: ctx.s.version, })) .execute(); } ``` ### Compare Versions ```typescript interface VersionDiff { field: string; before: unknown; after: unknown; } async function compareVersions( id: string, version1: number, version2: number ): Promise { const versions = await store .query() .from("Setting", "s") .temporal("includeEnded") .whereNode("s", (s) => s.id.eq(id).and(s.version.in([version1, version2]))) .orderBy((ctx) => ctx.s.version, "asc") .select((ctx) => ctx.s) .execute(); if (versions.length !== 2) { throw new Error("Versions not found"); } const [before, after] = versions; const diffs: VersionDiff[] = []; const allKeys = new Set([...Object.keys(before), ...Object.keys(after)]); for (const key of allKeys) { const beforeVal = before[key as keyof typeof before]; const afterVal = after[key as keyof typeof after]; if (JSON.stringify(beforeVal) !== JSON.stringify(afterVal)) { diffs.push({ field: key, before: beforeVal, after: afterVal }); } } return diffs; } ``` ## Session Management ### Start Session ```typescript async function startSession( userId: string, metadata: { ipAddress?: string; userAgent?: string } ): Promise> { return store.transaction(async (tx) => { const session = await tx.nodes.Session.create({ startedAt: new Date().toISOString(), ipAddress: metadata.ipAddress, userAgent: metadata.userAgent, }); const user = await tx.nodes.User.getById(userId); if (!user) throw new Error(`User not found: ${userId}`); await tx.edges.hasSession.create(user, session, {}); return session; }); } ``` ### End Session ```typescript async function endSession(sessionId: string): Promise { await store.nodes.Session.update(sessionId, { endedAt: new Date().toISOString(), }); } ``` ### Get Session Activity ```typescript async function getSessionActivity( sessionId: string ): Promise> { return store .query() .from("Session", "s") .whereNode("s", (s) => s.id.eq(sessionId)) .traverse("inSession", "e", { direction: "in" }) .to("AuditEntry", "a") .orderBy((ctx) => ctx.a.timestamp, "asc") .select((ctx) => ({ timestamp: ctx.a.timestamp, action: ctx.a.action, entityType: ctx.a.entityType, })) .execute(); } ``` ## Compliance Queries ### Who Changed This? ```typescript async function whoChanged( entityType: string, entityId: string, field: string ): Promise> { const entries = await store .query() .from("AuditEntry", "a") .whereNode("a", (a) => a.entityType.eq(entityType).and(a.entityId.eq(entityId)) ) .traverse("performedBy", "e") .to("User", "u") .orderBy((ctx) => ctx.a.timestamp, "desc") .select((ctx) => ({ changes: ctx.a.changes, user: ctx.u.name, timestamp: ctx.a.timestamp, })) .execute(); return entries .filter((e) => e.changes && field in e.changes) .map((e) => ({ user: e.user, timestamp: e.timestamp, before: e.changes![field].before, after: e.changes![field].after, })); } ``` ### When Was This Value Set? ```typescript async function whenWasValueSet( entityType: string, entityId: string, field: string, value: unknown ): Promise<{ timestamp: string; user: string } | undefined> { const entries = await store .query() .from("AuditEntry", "a") .whereNode("a", (a) => a.entityType.eq(entityType).and(a.entityId.eq(entityId)) ) .traverse("performedBy", "e") .to("User", "u") .orderBy((ctx) => ctx.a.timestamp, "asc") .select((ctx) => ({ changes: ctx.a.changes, user: ctx.u.name, timestamp: ctx.a.timestamp, })) .execute(); const entry = entries.find( (e) => e.changes && field in e.changes && e.changes[field].after === value ); return entry ? { timestamp: entry.timestamp, user: entry.user } : undefined; } ``` ### Search Audit Trail by Reason Compliance questions rarely come with exact match criteria — "find everything related to the data-retention incident in Q3" is more common than "find entries with reason = 'X'". Fulltext over the `reason` field gives reviewers a BM25-ranked list instead of a brittle substring match: ```typescript async function searchAuditByReason( query: string, options: { since?: Date; limit?: number } = {}, ) { const { since, limit = 25 } = options; const hits = await store.search.fulltext("AuditEntry", { query, limit: since ? limit * 4 : limit, includeSnippets: true, }); if (since === undefined) { return hits.map((hit) => ({ entry: hit.node, score: hit.score, snippet: hit.snippet, })); } const cutoff = since.toISOString(); return hits .filter((hit) => hit.node.timestamp >= cutoff) .slice(0, limit) .map((hit) => ({ entry: hit.node, score: hit.score, snippet: hit.snippet, })); } ``` For stricter date-range composition (the fulltext candidate pool is bounded by the `limit` above), use the query-builder path: ```typescript async function searchAuditInRange( query: string, startDate: Date, endDate: Date, limit = 25, ) { return store .query() .from("AuditEntry", "a") .whereNode("a", (a) => a.$fulltext .matches(query, limit * 2) .and(a.timestamp.gte(startDate.toISOString())) .and(a.timestamp.lte(endDate.toISOString())), ) .select((ctx) => ctx.a) .execute(); } ``` ## Export Audit Logs ### Stream to External System ```typescript async function* exportAuditLogs( since: Date, batchSize = 1000 ): AsyncGenerator { const stream = store .query() .from("AuditEntry", "a") .whereNode("a", (a) => a.timestamp.gte(since.toISOString())) .traverse("performedBy", "e") .to("User", "u") .orderBy((ctx) => ctx.a.timestamp, "asc") .select((ctx) => ({ ...ctx.a, performedBy: ctx.u.email, })) .stream({ batchSize }); let batch: AuditEntryProps[] = []; for await (const entry of stream) { batch.push(entry); if (batch.length >= batchSize) { yield batch; batch = []; } } if (batch.length > 0) { yield batch; } } // Usage async function syncToExternalAuditSystem(since: Date): Promise { for await (const batch of exportAuditLogs(since)) { await externalAuditApi.ingestBatch(batch); } } ``` ## Next Steps - [Document Management](/examples/document-management) - CMS with semantic search - [Product Catalog](/examples/product-catalog) - Categories, variants, inventory - [Workflow Engine](/examples/workflow-engine) - State machines with approvals # Bitemporal Time Travel > Valid time plus recorded time for corrections, effective dating, and the bitemporal 2x2 This example shows TypeGraph's built-in temporal history path: - **Valid time**: when a fact is true in the world, controlled by `validFrom` / `validTo` and read with `store.asOf(T)`. - **Recorded time**: when TypeGraph wrote the fact down, enabled with `history: true` and read with `store.asOfRecorded(T)`. Together they let you answer the audit question a single clock cannot express for TypeGraph-managed writes: what TypeGraph captured as true at a recorded commit instant. :::tip[Just want the code?] Full source on GitHub: [`packages/typegraph/examples/20-bitemporal-time-travel.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/20-bitemporal-time-travel.ts) ::: ## What It Demonstrates - Recorded-time reconstruction after a correction: the invoice amount as it was first reported versus the amount known now. - Valid-time effective dating: a promotion active only inside its validity window. - The bitemporal 2x2: the same subscription question answered at two valid-time instants and two recorded-time instants. - `store.recordedNow()` as the stable recorded-time anchor after writes. - `store.asOf(validT).asOfRecorded(recordedT)` for independent valid and recorded axes. ## Run It From the repository root: ```bash pnpm --filter @nicia-ai/typegraph exec tsx examples/20-bitemporal-time-travel.ts ``` Or from `packages/typegraph`: ```bash npx tsx examples/20-bitemporal-time-travel.ts ``` It uses an in-memory SQLite backend with `history: true`, so it needs no Docker or external services. ## Core API ```typescript const [store] = await createStoreWithSchema(graph, backend, { history: true, }); const invoice = await store.nodes.Invoice.create({ vendor: "Acme", amount: 1000, }); const asReported = await store.recordedNow(); if (asReported === undefined) throw new Error("expected recorded history"); await store.nodes.Invoice.update(invoice.id, { vendor: "Acme", amount: 1250, }); const reportedThen = await store .asOfRecorded(asReported) .nodes.Invoice.getById(invoice.id); ``` For independent axes, start from a valid-time view and add the recorded-time pin: ```typescript const capturedOnJul15BeforeCorrection = await store .asOf("2024-07-15T00:00:00.000Z") .asOfRecorded(beforeCorrection) .nodes.Subscription.getById(subscriptionId); ``` Direct `store.asOfRecorded(T)` is diagonal sugar: it pins the recorded axis and the valid-time axis to the same recorded instant. Chaining from `store.asOf(T)` is the form to use when the domain-effective date and the TypeGraph-capture date are different. ## Sample Output ```text [1] Recorded time - a correction Invoice inv_... (vendor: Acme) as reported (asOfRecorded): $1000 as known now (live read): $1250 [3] Both axes - the bitemporal 2x2 valid-time ↓ \ recorded-time -> before correction now valid May 1 ✓ active ✓ active valid Jul 15 ✗ inactive ✓ active ``` The Jul-15 / before-correction cell is the important one: at that valid date, the captured graph state had the subscription already cancelled, even though you now know it was still active. ## When to Use This Pattern Use bitemporal reads when the difference between **truth in the domain** and **captured TypeGraph state** matters: - Financial restatements and audit reports - Policy or contract effective dating - Compliance snapshots generated from later-corrected data - Support investigations where the system's earlier captured state matters See [Temporal queries](/queries/temporal#recorded-time-bitemporal) for the full API contract and limitations. # Breach Forensics > Pin an access graph to the breach instant and traverse the real blast radius This example uses bitemporal graph reconstruction for incident response. The question after a breach is not "what can this account reach now?" It is "what could this account reach at the moment of compromise?" By the time the investigation starts, dangerous grants may have been revoked or hard-deleted. A live graph can understate exposure. With `history: true`, you can pin the access graph to the recorded breach instant and run `reachable()` over the reconstructed graph. :::tip[Just want the code?] Full source on GitHub: [`packages/typegraph/examples/22-breach-forensics.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/22-breach-forensics.ts) ::: ## What It Demonstrates - A role/access graph with `Account -> Role -> Resource` paths. - A dangerous `deployer -> admin` escalation edge that is later hard-deleted. - `store.recordedNow()` as the breach-time recorded anchor. - `store.asOfRecorded(breachTime).reachable(...)` to reconstruct the blast radius at compromise time. - Point reads (`getByIds`) on the recorded view to resolve reached resources. ## Run It From the repository root: ```bash pnpm --filter @nicia-ai/typegraph exec tsx examples/22-breach-forensics.ts ``` Or from `packages/typegraph`: ```bash npx tsx examples/22-breach-forensics.ts ``` ## The Access Graph ```text svc-deploy --assumes--> deployer --grants--> ci-secrets | escalates v admin --grants--> prod-db, customer-pii ``` The `escalates` edge is the dangerous misconfiguration. Incident response removes it, but the recorded graph still knows it existed at the breach instant. ## Core API ```typescript const breachTime = await store.recordedNow(); if (breachTime === undefined) throw new Error("expected recorded history"); await store.edges.escalates.hardDelete(overGrant.id); const reachedAtBreach = await store .asOfRecorded(breachTime) .reachable(account.id, { edges: ["assumes", "escalates", "grants"], maxHops: 10, }); ``` The example wraps this in a small helper that accepts either a live view or a recorded view: ```typescript type AccessView = Pick< RecordedStoreView, "reachable" | "nodes" >; ``` That shape matters: incident-response code can run the same traversal against "now" and "then" without duplicating logic. ## Sample Output ```text Reachable resources on the CURRENT graph: - ci-secrets (medium) Looks contained - but the escalation was deleted. Reachable resources reconstructed AT THE BREACH INSTANT: - ci-secrets (medium) - prod-db (high) <- EXPOSED - customer-pii (critical) <- EXPOSED ``` ## When to Use This Pattern Use bitemporal graph forensics when deleted or corrected relationships affect the answer: - Identity and access blast-radius analysis - Data-sharing and entitlement investigations - Incident timelines where cleanup changed the graph - "Who could reach what?" reports at a historical recorded instant See [Temporal queries](/queries/temporal#recorded-time-bitemporal) for recorded-time constraints and [Graph Algorithms](/graph-algorithms) for `reachable()`. # Document Management System > A complete CMS example with semantic search, versioning, and access control This example builds a document management system with: - **Document hierarchy** (folders, documents, sections) - **Semantic search** with vector embeddings - **Version history** using temporal queries - **Access control** with permission inheritance - **Related documents** discovery ## Schema Definition ```typescript import { z } from "zod"; import { defineNode, defineEdge, defineGraph, embedding, searchable, subClassOf, partOf, hasPart, } from "@nicia-ai/typegraph"; // Base content type (abstract) const Content = defineNode("Content", { schema: z.object({ title: z.string(), createdBy: z.string(), status: z.enum(["draft", "published", "archived"]).default("draft"), }), }); // Folder extends Content const Folder = defineNode("Folder", { schema: z.object({ title: z.string(), createdBy: z.string(), status: z.enum(["draft", "published", "archived"]).default("draft"), path: z.string(), // /engineering/specs }), }); // Document extends Content const Document = defineNode("Document", { schema: z.object({ // Both fields are indexed for BM25 ranked fulltext. Combined with // the embedding below, this unlocks hybrid retrieval: title matches // (proper nouns, acronyms, terms-of-art) via BM25 plus paraphrased // / conceptual matches via the embedding. title: searchable({ language: "english" }), content: searchable({ language: "english" }), createdBy: z.string(), status: z.enum(["draft", "published", "archived"]).default("draft"), contentType: z.enum(["markdown", "html", "plaintext"]).default("markdown"), embedding: embedding(1536).optional(), }), }); // Users and permissions const User = defineNode("User", { schema: z.object({ email: z.string().email(), name: z.string(), role: z.enum(["admin", "editor", "viewer"]).default("viewer"), }), }); const Permission = defineNode("Permission", { schema: z.object({ level: z.enum(["read", "write", "admin"]), }), }); // Edges const contains = defineEdge("contains"); const relatedTo = defineEdge("relatedTo", { schema: z.object({ type: z.enum(["references", "supersedes", "related"]), confidence: z.number().min(0).max(1).optional(), }), }); const hasPermission = defineEdge("hasPermission"); const createdBy = defineEdge("createdBy"); // Graph definition const graph = defineGraph({ id: "document_management", nodes: { Content: { type: Content }, Folder: { type: Folder, unique: [ { name: "folder_path", fields: ["path"], scope: "kind", collation: "binary", }, ], }, Document: { type: Document }, User: { type: User, unique: [ { name: "user_email", fields: ["email"], scope: "kind", collation: "caseInsensitive", }, ], }, Permission: { type: Permission }, }, edges: { contains: { type: contains, from: [Folder], to: [Folder, Document] }, relatedTo: { type: relatedTo, from: [Document], to: [Document] }, hasPermission: { type: hasPermission, from: [User], to: [Content] }, createdBy: { type: createdBy, from: [Content], to: [User] }, }, ontology: [ // Type hierarchy subClassOf(Folder, Content), subClassOf(Document, Content), // Compositional relationships partOf(Document, Folder), hasPart(Folder, Document), ], }); ``` ## Database Setup ```typescript import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local"; import { createStoreWithSchema } from "@nicia-ai/typegraph"; // `createLocalSqliteBackend` loads sqlite-vec (an optional peer dep) and wires // the vector strategy, so `embedding()` fields are persisted to real `vec0` // KNN storage. It also runs the base DDL for you. const { backend } = createLocalSqliteBackend({ path: "documents.db" }); // `searchable()` fields require the durable fulltext-materialization // step that `createStoreWithSchema` performs at boot. Bare // `createStore()` is an attach-only path and would throw // `StoreNotInitializedError` on the first fulltext write. const [store] = await createStoreWithSchema(graph, backend); ``` ## Core Operations ### Creating Folder Structure ```typescript async function createFolderPath(path: string, userId: string): Promise> { const parts = path.split("/").filter(Boolean); let currentPath = ""; let parentFolder: Node | undefined; for (const part of parts) { currentPath += `/${part}`; // The `folder_path` unique constraint makes this atomic: concurrent // callers converge on one folder instead of racing to create duplicates. const result = await store.nodes.Folder.getOrCreateByConstraint( "folder_path", { title: part, path: currentPath, createdBy: userId, status: "published", }, ); if (result.action === "created" && parentFolder) { await store.edges.contains.create(parentFolder, result.node, {}); } parentFolder = result.node; } return parentFolder!; } ``` ### Creating Documents with Embeddings ```typescript import OpenAI from "openai"; const openai = new OpenAI(); async function generateEmbedding(text: string): Promise { const response = await openai.embeddings.create({ model: "text-embedding-ada-002", input: text, }); return response.data[0].embedding; } async function createDocument( folderId: string, title: string, content: string, userId: string ): Promise> { const embedding = await generateEmbedding(`${title}\n\n${content}`); const document = await store.nodes.Document.create({ title, content, createdBy: userId, status: "draft", contentType: "markdown", embedding, }); // Link to folder const folder = await store.nodes.Folder.getById(folderId); if (!folder) throw new Error(`Folder not found: ${folderId}`); await store.edges.contains.create(folder, document, {}); // Link to creator const user = await store.nodes.User.getById(userId); if (!user) throw new Error(`User not found: ${userId}`); await store.edges.createdBy.create(document, user, {}); return document; } ``` ### Updating Documents (Versioned) ```typescript async function updateDocument( documentId: string, updates: { title?: string; content?: string; status?: "draft" | "published" | "archived" } ): Promise> { const current = await store.nodes.Document.getById(documentId); if (!current) throw new Error(`Document not found: ${documentId}`); // If content changed, regenerate embedding let embedding = current.embedding; if (updates.content && updates.content !== current.content) { const text = `${updates.title ?? current.title}\n\n${updates.content}`; embedding = await generateEmbedding(text); } // Update creates a new version automatically return store.nodes.Document.update(documentId, { ...updates, embedding, }); } ``` ## Searching Documents Document search is the canonical hybrid-retrieval use case. Users search for proper nouns, project names, and quoted phrases that embeddings smooth over, plus conceptual questions that keyword search alone can't answer. This example shows both sides. ### Semantic Search Embedding-only search is still useful when the query is a paraphrase or a question rather than a set of keywords: ```typescript async function searchDocumentsSemantically( query: string, options: { folderId?: string; status?: "draft" | "published" | "archived"; limit?: number; minScore?: number; } = {} ): Promise { const { folderId, status = "published", limit = 10, minScore = 0.7 } = options; const queryEmbedding = await generateEmbedding(query); let queryBuilder = store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding .similarTo(queryEmbedding, limit, { metric: "cosine", minScore }) .and(d.status.eq(status)), ); // If folderId specified, filter to folder descendants if (folderId) { queryBuilder = store .query() .from("Folder", "f") .whereNode("f", (f) => f.id.eq(folderId)) .traverse("contains", "e") .recursive() .to("Document", "d") .whereNode("d", (d) => d.embedding .similarTo(queryEmbedding, limit, { metric: "cosine", minScore }) .and(d.status.eq(status)), ); } // Results are already ordered by similarity (most similar first). return queryBuilder.select((ctx) => ctx.d).execute(); } ``` ### Fulltext Search (BM25 with snippets) When the user types a proper noun, filename, or exact phrase, keyword search outperforms embeddings — and snippets give them a preview of where the match occurred: ```typescript async function findDocumentsByKeyword( query: string, options: { limit?: number } = {}, ): Promise> { const hits = await store.search.fulltext("Document", { query, limit: options.limit ?? 10, includeSnippets: true, }); return hits.map((hit) => ({ document: hit.node, score: hit.score, snippet: hit.snippet, })); } ``` ### Hybrid Search (the production-grade path) Most document-search products want both signals. `store.search.hybrid()` runs fulltext + vector in parallel and fuses the rankings with RRF. Each hit carries sub-scores from each half for debugging: ```typescript async function searchDocuments( query: string, options: { status?: "draft" | "published" | "archived"; limit?: number; } = {}, ): Promise> { const { status = "published", limit = 10 } = options; const queryEmbedding = await generateEmbedding(query); const hits = await store.search.hybrid("Document", { limit, vector: { fieldPath: "embedding", queryEmbedding, metric: "cosine", k: limit * 4, }, fulltext: { query, k: limit * 4, includeSnippets: true, }, fusion: { method: "rrf", k: 60 }, }); return hits .filter((hit) => hit.node.status === status) .map((hit) => ({ document: hit.node, score: hit.score, snippet: hit.fulltext?.snippet, })); } ``` ### Folder-Scoped Hybrid Search (query builder path) For tighter composition — "only within this folder subtree, using hybrid retrieval" — `$fulltext.matches()` and `.similarTo()` combine in one query-builder statement. The fusion runs at the SQL layer: ```typescript async function searchInFolder( folderId: string, query: string, limit = 10, ): Promise { const queryEmbedding = await generateEmbedding(query); return store .query() .from("Folder", "f") .whereNode("f", (f) => f.id.eq(folderId)) .traverse("contains", "e") .recursive() .to("Document", "d") .whereNode("d", (d) => d.$fulltext .matches(query, limit * 4) .and(d.embedding.similarTo(queryEmbedding, limit * 4)) .and(d.status.eq("published")), ) .fuseWith({ k: 60 }) .select((ctx) => ctx.d) .limit(limit) .execute(); } ``` ### Find Related Documents ```typescript async function findRelatedDocuments( documentId: string, limit = 5 ): Promise> { // First, get explicit relationships const explicit = await store .query() .from("Document", "d") .whereNode("d", (d) => d.id.eq(documentId)) .traverse("relatedTo", "e") .to("Document", "related") .select((ctx) => ({ document: ctx.related, relationship: ctx.e.type, })) .execute(); // Then, find semantically similar documents const source = await store.nodes.Document.getById(documentId); if (!source) throw new Error(`Document not found: ${documentId}`); if (!source.embedding) { return explicit; } const similar = await store .query() .from("Document", "d") .whereNode("d", (d) => d.embedding .similarTo(source.embedding!, limit * 2, { metric: "cosine", minScore: 0.8 }) .and(d.id.neq(documentId)) ) .select((ctx) => ({ document: ctx.d, relationship: "similar" as const, })) .limit(limit) .execute(); return [...explicit, ...similar].slice(0, limit); } ``` ## Version History ### Get Document History ```typescript interface DocumentVersion { title: string; content: string; status: string; validFrom: string; validTo: string | undefined; version: number; } async function getDocumentHistory(documentId: string): Promise { return store .query() .from("Document", "d") .temporal("includeEnded") .whereNode("d", (d) => d.id.eq(documentId)) .orderBy((ctx) => ctx.d.validFrom, "desc") .select((ctx) => ({ title: ctx.d.title, content: ctx.d.content, status: ctx.d.status, validFrom: ctx.d.validFrom, validTo: ctx.d.validTo, version: ctx.d.version, })) .execute(); } ``` ### View Document at Point in Time ```typescript async function getDocumentAsOf( documentId: string, timestamp: Date ): Promise { return store .query() .from("Document", "d") .temporal("asOf", timestamp.toISOString()) .whereNode("d", (d) => d.id.eq(documentId)) .select((ctx) => ctx.d) .first(); } ``` ### Compare Versions ```typescript async function compareVersions( documentId: string, version1: number, version2: number ): Promise<{ before: DocumentProps; after: DocumentProps } | undefined> { const versions = await store .query() .from("Document", "d") .temporal("includeEnded") .whereNode("d", (d) => d.id.eq(documentId).and(d.version.in([version1, version2])) ) .orderBy((ctx) => ctx.d.version, "asc") .select((ctx) => ctx.d) .execute(); if (versions.length !== 2) return undefined; return { before: versions[0], after: versions[1] }; } ``` ## Access Control ### Check Read Permission ```typescript async function canRead(userId: string, contentId: string): Promise { // Check direct permission const directPermission = await store .query() .from("User", "u") .whereNode("u", (u) => u.id.eq(userId)) .traverse("hasPermission", "p") .to("Content", "c", { includeSubClasses: true }) .whereNode("c", (c) => c.id.eq(contentId)) .first(); if (directPermission) return true; // Check inherited permission (from parent folders) const content = await store.nodes.Content.getById(contentId); if (!content) return false; // Walk up the folder tree checking permissions const parentFolders = await store .query() .from("Folder", "f") .traverse("contains", "e") .recursive() .to("Content", "c", { includeSubClasses: true }) .whereNode("c", (c) => c.id.eq(contentId)) .select((ctx) => ctx.f.id) .execute(); for (const folderId of parentFolders) { const folderPermission = await store .query() .from("User", "u") .whereNode("u", (u) => u.id.eq(userId)) .traverse("hasPermission", "p") .to("Folder", "f") .whereNode("f", (f) => f.id.eq(folderId)) .first(); if (folderPermission) return true; } return false; } ``` ### Grant Permission ```typescript async function grantPermission( userId: string, contentId: string, level: "read" | "write" | "admin" ): Promise { const user = await store.nodes.User.getById(userId); if (!user) throw new Error(`User not found: ${userId}`); const content = await store.nodes.Content.getById(contentId); if (!content) throw new Error(`Content not found: ${contentId}`); // Create permission node const permission = await store.nodes.Permission.create({ level }); // Link user to content via permission await store.edges.hasPermission.create(user, content, {}); } ``` ## Folder Navigation ### Get Folder Contents ```typescript interface FolderContents { folders: Array<{ id: string; title: string; path: string }>; documents: Array<{ id: string; title: string; status: string }>; } async function getFolderContents(folderId: string): Promise { const folders = await store .query() .from("Folder", "parent") .whereNode("parent", (f) => f.id.eq(folderId)) .traverse("contains", "e") .to("Folder", "child") .select((ctx) => ({ id: ctx.child.id, title: ctx.child.title, path: ctx.child.path, })) .execute(); const documents = await store .query() .from("Folder", "parent") .whereNode("parent", (f) => f.id.eq(folderId)) .traverse("contains", "e") .to("Document", "doc") .select((ctx) => ({ id: ctx.doc.id, title: ctx.doc.title, status: ctx.doc.status, })) .execute(); return { folders, documents }; } ``` ### Get Breadcrumb Path `store.algorithms.reachable` walks `contains` edges in reverse to collect every ancestor folder, tagged with its depth from the starting content: ```typescript async function getBreadcrumb( contentId: string ): Promise> { const ancestors = await store.algorithms.reachable(contentId, { edges: ["contains"], direction: "in", excludeSource: true, }); if (ancestors.length === 0) return []; const folderIds = ancestors .filter((node) => node.kind === "Folder") .toSorted((a, b) => b.depth - a.depth) // root first .map((node) => node.id); const folders = await store.nodes.Folder.getByIds(folderIds); return folders .filter((folder): folder is NonNullable => folder !== undefined) .map((folder) => ({ id: folder.id, title: folder.title, path: folder.path, })); } ``` `reachable` returns `{ id, kind, depth }` from a set-based BFS frontier, then a single batched `getByIds` hydrates the folder properties. ## Bulk Operations ### Move Document to Folder ```typescript async function moveDocument(documentId: string, targetFolderId: string): Promise { await store.transaction(async (tx) => { // Remove from current folder const currentEdge = await tx .query() .from("Folder", "f") .traverse("contains", "e") .to("Document", "d") .whereNode("d", (d) => d.id.eq(documentId)) .select((ctx) => ctx.e.id) .first(); if (currentEdge) { await tx.edges.contains.delete(currentEdge); } // Add to new folder const document = await tx.nodes.Document.getById(documentId); if (!document) throw new Error(`Document not found: ${documentId}`); const targetFolder = await tx.nodes.Folder.getById(targetFolderId); if (!targetFolder) throw new Error(`Folder not found: ${targetFolderId}`); await tx.edges.contains.create(targetFolder, document, {}); }); } ``` ### Bulk Archive ```typescript async function archiveFolder(folderId: string): Promise { // Get all documents in folder and subfolders const documents = await store .query() .from("Folder", "f") .whereNode("f", (f) => f.id.eq(folderId)) .traverse("contains", "e") .recursive() .to("Document", "d") .select((ctx) => ctx.d.id) .execute(); // Archive each document await store.transaction(async (tx) => { for (const docId of documents) { await tx.nodes.Document.update(docId, { status: "archived" }); } }); return documents.length; } ``` ## Next Steps - [Product Catalog](/examples/product-catalog) - Categories, variants, inventory - [Workflow Engine](/examples/workflow-engine) - State machines with approvals - [Audit Trail](/examples/audit-trail) - Complete change tracking # FHIR Graph Merge > Reconcile overlapping FHIR-style records from independent branches into one canonical patient care graph. This example shows TypeGraph's graph-merge primitive on a concrete healthcare interoperability problem: two independent ingestion agents extract overlapping FHIR-style records, disagree on patient spelling, and produce separate care context. TypeGraph branches isolate those writes. `merge()` folds them into one base graph, resolves the duplicate patient, repoints each branch's clinical edges, and reports the conflict/provenance trail. :::tip[Just want the code?] Full source on GitHub: [`packages/typegraph/examples/18-fhir-graph-merge.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/18-fhir-graph-merge.ts) ::: :::caution[Synthetic demo data only] This is an interoperability and data-exploration example. It is not clinical decision support and should not be run on protected health information as-is. The bundled patient record is synthetic and intentionally small. ::: ## What It Demonstrates - `branch()` creates isolated working copies over fresh local SQLite backends. - Two patient pairs are reconciled by **two different mechanisms**, so you can see exact identity and fuzzy resolution side by side: - **`Anna Rivera` / `Ana Rivera`** share a unique MRN. The unique constraint forces this merge by **exact identity** — fuzzy scoring is bypassed, so the similarity threshold is irrelevant to their collapse. - **`Mohammed Ali` / `Mohamed Ali`** have **different** MRNs, so nothing forces them. Only the in-memory fulltext name similarity clearing the threshold collapses them — this is the case where similarity is actually decisive. - Care-context edges from both branches are repointed onto each canonical patient. The output proves this by reading context through `forPatient` edges *into* the survivor, not via a blanket `.find()`. - The merge report surfaces every spelling / identifier disagreement (name, MRN, FHIR id) instead of silently choosing a winner. - Report provenance shows which branch contributed the canonical nodes and edges. ## Run It From the repository root: ```bash pnpm --filter @nicia-ai/typegraph exec tsx examples/18-fhir-graph-merge.ts ``` Or from `packages/typegraph`: ```bash npx tsx examples/18-fhir-graph-merge.ts ``` The example uses in-memory SQLite backends, so it does not require Docker, a FHIR server, or external services. ## Sample Output ```text === FHIR Graph Merge === Before merge: base patients: 0 EHR branch patients: 2 claims branch patients: 2 After merge: merged nodes: 9 merged edges: 10 entity resolutions: 2 conflicts: - Patient.fhirId on patient-ana: claims-agent="Patient/claims-ana", ehr-agent="Patient/ehr-anna" - Patient.name on patient-ana: claims-agent="Ana Rivera", ehr-agent="Anna Rivera" - Patient.fhirId on patient-mohamed: claims-agent="Patient/claims-mohamed", ehr-agent="Patient/ehr-mohammed" - Patient.mrn on patient-mohamed: claims-agent="MRN-205", ehr-agent="MRN-204" - Patient.name on patient-mohamed: claims-agent="Mohamed Ali", ehr-agent="Mohammed Ali" provenance: - ehr-agent: 6 node(s), 6 edge(s) - claims-agent: 5 node(s), 4 edge(s) Canonical patients and their repointed care context: Ana Rivera (MRN-001, 1974-03-09) - Encounter: Hypertension follow-up (2026-04-11T09:30:00-07:00) - Encounter: Kidney function review (2026-04-14T10:00:00-07:00) - MedicationRequest: Lisinopril 10 MG Oral Tablet - Take one tablet by mouth daily - Observation: Blood pressure panel = 152/96 mmHg (high) - Observation: Estimated glomerular filtration rate = 54 mL/min/1.73m2 (low) Mohamed Ali (MRN-205, 1990-08-21) - Encounter: Cardiology consult (2026-05-02T13:00:00-07:00) - Observation: LDL cholesterol = 168 mg/dL (high) ``` The `Mohamed Ali` line is the proof that fuzzy resolution and edge repointing both work: the two spellings carried **different** MRNs (so no exact match forced them), yet they collapsed to one patient — and that patient now owns the EHR branch's cardiology encounter *and* the claims branch's lab observation, each repointed from its origin branch. ## Graph Model The example keeps the schema intentionally small: | TypeGraph kind | FHIR-ish source | Purpose | | -------------- | --------------- | ------- | | `Patient` | `Patient` | The canonical person being reconciled | | `Encounter` | `Encounter` | Clinical visit context | | `Observation` | `Observation` | Vitals and lab evidence | | `MedicationRequest` | `MedicationRequest` | Medication order | | `forPatient` | `subject` references | Connects clinical resources to the patient | | `duringEncounter` | `encounter` references | Connects observations/orders to visits | | `reasonFor` | `reasonReference` | Connects an order to supporting evidence | The `Patient` kind declares a unique MRN constraint: ```typescript const careGraph = defineGraph({ nodes: { Patient: { type: Patient, unique: [ { name: "patient_mrn", fields: ["mrn"], scope: "kind", collation: "caseInsensitive", }, ], }, }, }); ``` A shared unique value is a **definitional** identity source: when two staged nodes share all of a constraint's fields, they are forced to merge and fuzzy scoring is bypassed entirely. That is what collapses `Anna Rivera` / `Ana Rivera` (both `MRN-001`) — the similarity threshold never enters into it. The merge still reports their property disagreements (`name`, `fhirId`). ## Merge Configuration The example configures entity resolution only for `Patient`; other kinds merge by ID and keep their branch-specific clinical context. ```typescript const mergeOptions = { resolve: { Patient: { // Block by shared birth date. The unique MRN constraint forces exact // matches in its own bucket, so blocking does not need the MRN — and using // it would split the different-MRN fuzzy pair apart. block: (node) => node.birthDate, similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.78, }, }, onPropertyConflict: "flag", branchOrder: [EHR_BRANCH, CLAIMS_BRANCH], provenance: true, }; ``` The `fulltext` strategy is an in-memory trigram scorer. It does not require embeddings, database fulltext indexes, or a specific backend, which makes it a good default for bounded candidate dedup during merge. It is the **decisive** signal for `Mohammed Ali` / `Mohamed Ali`: those records carry different MRNs, so no unique match forces them — only the name score clearing `threshold` collapses them. ## Why This Matters FHIR imports often arrive from multiple systems: EHR exports, claims feeds, lab feeds, patient-reported questionnaires, and agent-produced extraction passes. Each source can be useful, but naive append-only ingestion leaves duplicates and broken care context. Graph merge gives you a deterministic reconciliation step: 1. Run each extractor in an isolated branch. 2. Keep all writes type-checked by the same TypeGraph schema. 3. Resolve entities by exact identity, blocking keys, and similarity. 4. Preserve branch-specific context by repointing edges. 5. Return a merge report that callers can inspect, persist, or route for human review. See [Graph Merge](/graph-merge) for the API reference and option semantics. # Incremental Merge > Ingest into a live graph in waves — mergeIncremental() folds a new source onto an already-committed graph without creating duplicates, and persists a queryable provenance trail. This example shows TypeGraph's **incremental** merge path. Where the snapshot [`merge()`](/examples/fhir-graph-merge) reconciles branches that all forked from the *current* base, `mergeIncremental()` folds a new source's branch into a `target` that has **already advanced** — re-discovering an entity that was committed in an earlier wave and merging onto it instead of creating a duplicate. It is the primitive for **continuous ingestion**: every new feed, crawl, or agent batch lands on the live graph, deduplicated against what is already there. :::tip[Just want the code?] Full source on GitHub: [`packages/typegraph/examples/19-incremental-merge.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/19-incremental-merge.ts) ::: ## What It Demonstrates - A `target` graph that already holds one committed `Company` ("Acme Corp"). - A new provider branch (forked from an empty fork-point in this example) that re-reports the same company under a different spelling ("ACME Corporation", same `acme.com` domain) and adds a genuinely new one ("Globex"). - `mergeIncremental()` recalls the committed company via its unique `domain` and merges the new spelling **onto** it — the target keeps one Acme, not two. - The name disagreement is flagged in the report instead of silently overwriting the committed value. - Provenance is persisted to a sidecar graph and queried back: "which canonical entities did this provider contribute to?" ## Run It From the repository root: ```bash pnpm --filter @nicia-ai/typegraph exec tsx examples/19-incremental-merge.ts ``` Or from `packages/typegraph`: ```bash npx tsx examples/19-incremental-merge.ts ``` It uses in-memory SQLite backends, so it needs no Docker or external services. ## Sample Output ```text === Incremental Graph Merge === Target before: [ 'Acme Corp (acme.com)' ] Target after: [ 'Acme Corp (acme.com)', 'Globex (globex.io)' ] No duplicate was created: the provider's "ACME Corporation" merged onto the committed "Acme Corp" via the shared domain. Merged nodes: 2 Entity resolutions: 1 Conflict on Company.name @ acme: provider-crunchbase="ACME Corporation" Provenance persisted: 3 row(s) in sidecar "company_kb::merge-provenance" Provenance — canonical entities this provider contributed to: - Company "cb-globex" (from source "cb-globex") - Company "acme" (from source "cb-acme") ``` ## How It Works ```typescript const result = await mergeIncremental({ forkPoint, // the frozen ancestor the provider branch forked from target, // the live committed graph (already holds "Acme Corp") branches: [provider], options: { resolve: { Company: { similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.9, }, }, onPropertyConflict: "flag", onBasePropertyConflict: "flag", // required by mergeIncremental (keep-base) branchOrder: [PROVIDER], persistProvenance: true, }, }); ``` Two ideas do the work: 1. **Unique `domain` as definitional identity.** The `Company` kind declares a unique constraint on `domain`. The new-vs-base recall queries the target for committed companies that share a staged company's domain, so "ACME Corporation" (`acme.com`) is matched to the committed "Acme Corp" regardless of the name spelling — no similarity threshold needs to be cleared. 2. **Keep-base conflict handling.** `onBasePropertyConflict: "flag"` guarantees a stale branch value can never overwrite a newer committed one: the committed `name` is kept and the provider's spelling is recorded as a conflict for review. The genuinely-new "Globex" has no committed match, so it is created. After the commit, `persistProvenance` writes one row per contribution to a sidecar graph on the target's backend, which `readProvenance` reads back later. ## Why This Matters Real ingestion is never one-shot. Feeds arrive continuously, crawls re-run, and agents produce overlapping batches. Append-only ingestion turns that into a pile of duplicates; a full re-merge from scratch does not scale. `mergeIncremental()` gives you a **steady-state ingestion loop**: 1. Take each new source as a branch off a known fork-point. 2. Merge it into the live target; already-known entities are recalled and updated in place, inherited modifications/deletions propagate, and new ones are created. 3. Keep concurrently changed committed data authoritative (`onBasePropertyConflict: "flag"`). 4. Persist provenance so every canonical entity carries the trail of which source contributed it. See [Graph Merge](/graph-merge) for the full API reference, and [FHIR Graph Merge](/examples/fhir-graph-merge) for the snapshot-merge counterpart. # Knowledge Graph for RAG > Enhance retrieval with entity linking, relationship traversal, and multi-hop context This example demonstrates how **graph structure enhances RAG** beyond vector similarity. While [Semantic Search](/semantic-search) covers embedding basics, this guide focuses on what graphs uniquely provide: entity disambiguation, relationship traversal, and structured context that flat retrieval cannot offer. ## What Graphs Add to RAG | Flat RAG | Graph RAG | |----------|-----------| | Returns similar chunks | Traverses to related entities and facts | | Treats "Apple" the same everywhere | Disambiguates Apple Inc. vs. apple fruit | | Context is unstructured text | Context includes structured relationships | | Single-hop retrieval | Multi-hop reasoning across connections | **Example**: For "What companies has Elon Musk founded?", flat RAG returns chunks mentioning him. Graph RAG traverses from the "Elon Musk" entity through "founded" edges to return structured company data—regardless of whether those facts appear in the same chunk. ## Schema ```typescript import { z } from "zod"; import { defineNode, defineEdge, defineGraph, embedding, inverseOf, searchable } from "@nicia-ai/typegraph"; // Source documents const Document = defineNode("Document", { schema: z.object({ title: searchable({ language: "english" }), source: z.string(), }), }); // Text chunks with embeddings + fulltext const Chunk = defineNode("Chunk", { schema: z.object({ // `searchable()` enables `$fulltext.matches()` for BM25 ranking, // complementing the embedding-based semantic search below. text: searchable({ language: "english" }), embedding: embedding(1536), position: z.number().int(), }), }); // Extracted entities const Entity = defineNode("Entity", { schema: z.object({ name: searchable({ language: "english" }), type: z.enum(["person", "organization", "location", "concept", "product", "event"]), description: z.string().optional(), embedding: embedding(1536).optional(), }), }); // Edges const containsChunk = defineEdge("containsChunk"); const nextChunk = defineEdge("nextChunk"); const prevChunk = defineEdge("prevChunk"); const mentions = defineEdge("mentions", { schema: z.object({ confidence: z.number().min(0).max(1).optional(), }), }); const relatesTo = defineEdge("relatesTo", { schema: z.object({ relationship: z.string(), // "founded", "works_at", "located_in" }), }); export const graph = defineGraph({ id: "rag_graph", nodes: { Document: { type: Document }, Chunk: { type: Chunk }, Entity: { type: Entity, unique: [ { name: "entity_name_type", fields: ["name", "type"], scope: "kind", collation: "caseInsensitive", }, ], }, }, edges: { containsChunk: { type: containsChunk, from: [Document], to: [Chunk] }, nextChunk: { type: nextChunk, from: [Chunk], to: [Chunk] }, prevChunk: { type: prevChunk, from: [Chunk], to: [Chunk] }, mentions: { type: mentions, from: [Chunk], to: [Entity] }, relatesTo: { type: relatesTo, from: [Entity], to: [Entity] }, }, ontology: [inverseOf(nextChunk, prevChunk)], }); ``` ## Embedding Setup Using [Vercel AI SDK](https://ai-sdk.dev/docs/ai-sdk-core/embeddings): ```typescript import { embed, embedMany } from "ai"; import { openai } from "@ai-sdk/openai"; const embeddingModel = openai.embeddingModel("text-embedding-3-small"); async function generateEmbedding(text: string): Promise { const { embedding } = await embed({ model: embeddingModel, value: text }); return embedding; } async function generateEmbeddings(texts: string[]): Promise { const { embeddings } = await embedMany({ model: embeddingModel, values: texts }); return embeddings; } ``` ## Ingestion with Entity Linking The key graph RAG capability: linking chunks to disambiguated entities. ```typescript import type { Node } from "@nicia-ai/typegraph"; interface ChunkData { text: string; entities: Array<{ name: string; type: "person" | "organization" | "location" | "concept" | "product" | "event"; }>; } async function ingestDocument( title: string, source: string, chunks: ChunkData[] ): Promise { await store.transaction(async (tx) => { const doc = await tx.nodes.Document.create({ title, source }); // Batch embed all chunks const chunkEmbeddings = await generateEmbeddings(chunks.map((c) => c.text)); let prevChunk: Node | undefined; for (const [i, chunkData] of chunks.entries()) { const chunk = await tx.nodes.Chunk.create({ text: chunkData.text, embedding: chunkEmbeddings[i], position: i, }); await tx.edges.containsChunk.create(doc, chunk, {}); if (prevChunk) { await tx.edges.nextChunk.create(prevChunk, chunk, {}); } // Link to entities (dedupe by unique constraint) for (const entityData of chunkData.entities) { const entityResult = await tx.nodes.Entity.getOrCreateByConstraint( "entity_name_type", { name: entityData.name, type: entityData.type, } ); // Compute expensive derived fields only for newly created entities if (entityResult.action === "created") { await tx.nodes.Entity.update(entityResult.node.id, { embedding: await generateEmbedding(entityData.name), }); } await tx.edges.mentions.getOrCreateByEndpoints( chunk, entityResult.node, {}, { ifExists: "return" } ); } prevChunk = chunk; } }); } ``` ## Graph-Specific Query Patterns These patterns demonstrate capabilities that require graph structure—they cannot be replicated with flat vector search. ### Entity-Based Retrieval Find all chunks mentioning a specific entity, regardless of how it's phrased: ```typescript async function findChunksByEntity(entityName: string) { return store .query() .from("Entity", "e") .whereNode("e", (e) => e.name.eq(entityName)) .traverse("mentions", "m", { direction: "in" }) .to("Chunk", "c") .select((ctx) => ctx.c.text) .execute(); } ``` ### Multi-Hop Entity Traversal Find entities connected through relationships: ```typescript async function findRelatedEntities(entityName: string, maxHops = 2) { const rows = await store .query() .from("Entity", "e") .whereNode("e", (e) => e.name.eq(entityName)) .traverse("relatesTo", "r") .recursive({ maxHops, depth: "depth" }) .to("Entity", "related") .select((ctx) => ({ from: ctx.e.name, to: ctx.related.name, toId: ctx.related.id, depth: ctx.depth, })) .execute(); // distinct paths can reach the same target; dedupe by target const seen = new Set(); return rows .filter((row) => { if (seen.has(row.toId)) return false; seen.add(row.toId); return true; }) .map((row) => ({ from: row.from, to: row.to, depth: row.depth, })); } ``` ### Context Window Expansion Get surrounding chunks for a match: ```typescript async function getChunkWithContext(chunkId: string, windowSize = 1) { const [before, after] = await Promise.all([ store .query() .from("Chunk", "c") .whereNode("c", (c) => c.id.eq(chunkId)) .traverse("prevChunk", "e") .recursive({ maxHops: windowSize }) .to("Chunk", "prev") .orderBy("prev", "position", "desc") .select((ctx) => ctx.prev.text) .execute(), store .query() .from("Chunk", "c") .whereNode("c", (c) => c.id.eq(chunkId)) .traverse("nextChunk", "e") .recursive({ maxHops: windowSize }) .to("Chunk", "next") .orderBy("next", "position", "asc") .select((ctx) => ctx.next.text) .execute(), ]); const chunk = await store.nodes.Chunk.getById(chunkId); return { before: before.toReversed(), chunk: chunk?.text ?? "", after, }; } ``` ## Hybrid Retrieval: Vector + Fulltext Vector search handles semantic similarity; fulltext search (`$fulltext.matches()`) nails exact-match terms that embeddings blur — proper nouns, SKUs, technical jargon. Combining both with Reciprocal Rank Fusion is the gold-standard RAG retrieval pattern. ### Query-builder fusion (single SQL query) Use `$fulltext.matches()` and `.similarTo()` in the same `whereNode()` and TypeGraph fuses the two ranked lists with RRF at the SQL layer: ```typescript async function hybridSearch(query: string, limit = 10) { const queryEmbedding = await generateEmbedding(query); return store .query() .from("Chunk", "c") .whereNode("c", (c) => c.$fulltext .matches(query, limit * 4) .and(c.embedding.similarTo(queryEmbedding, limit * 4)) ) .select((ctx) => ({ chunkId: ctx.c.id, text: ctx.c.text, })) .limit(limit) .execute(); } ``` Each source retrieves `limit * 4` candidates, RRF blends the rankings, and the outer `LIMIT` trims to the final top-k. The two CTEs join by `node_id` and the outer ORDER BY uses `1/(60 + rank_vec) + 1/(60 + rank_ft)`. ### Store API (tunable weights) When you need to tune the fusion parameters (e.g. weighting fulltext higher for entity-heavy queries), use `store.search.hybrid()`: ```typescript async function tunedHybrid(query: string, limit = 10) { const queryEmbedding = await generateEmbedding(query); const hits = await store.search.hybrid("Chunk", { limit, vector: { fieldPath: "embedding", queryEmbedding, metric: "cosine", k: 50, }, fulltext: { query, k: 50, includeSnippets: true, }, fusion: { method: "rrf", k: 60, weights: { vector: 1.0, fulltext: 1.5 }, }, }); return hits.map((h) => ({ chunkId: h.node.id, text: h.node.text, score: h.score, vectorRank: h.vector?.rank, fulltextRank: h.fulltext?.rank, snippet: h.fulltext?.snippet, })); } ``` See the [Fulltext Search guide](/fulltext-search) for tuning advice, query modes, and the full hybrid retrieval playbook. ## Hybrid Retrieval: Vector + Graph Combine vector similarity with graph traversal in a single query using the `from` option: ```typescript async function hybridRetrieval(query: string, limit = 10) { const queryEmbedding = await generateEmbedding(query); // Single query: vector search + fan-out to entities AND document const results = await store .query() .from("Chunk", "c") .whereNode("c", (c) => c.embedding.similarTo(queryEmbedding, limit, { metric: "cosine", minScore: 0.7 }) ) .traverse("mentions", "m") .to("Entity", "e") .traverse("containsChunk", "d_edge", { direction: "in", from: "c" }) // Fan-out from chunk .to("Document", "d") // Results are already ordered by similarity (most similar first). // When you need explicit scores, use `store.search.hybrid()` instead — // it returns hits with `.score`, `.vector.score`, and `.fulltext.score`. .select((ctx) => ({ chunkId: ctx.c.id, text: ctx.c.text, source: ctx.d.title, entityName: ctx.e.name, entityType: ctx.e.type, })) .execute(); // Group by chunk (one row per chunk-entity pair) const byChunk = new Map }>(); for (const row of results) { const existing = byChunk.get(row.chunkId); if (existing) { existing.entities.push({ name: row.entityName, type: row.entityType }); } else { byChunk.set(row.chunkId, { ...row, entities: [{ name: row.entityName, type: row.entityType }], }); } } return [...byChunk.values()]; } ``` The `from` option enables **fan-out patterns** where you traverse multiple relationships from the same node. Without `from`, traversals chain sequentially (A→B→C). With `from`, you can branch: traverse from chunk to entities, AND from chunk to document. ## Building Structured Context Format graph-enriched context for an LLM: ```typescript async function buildGraphContext(query: string, extractedEntities: string[]) { const queryEmbedding = await generateEmbedding(query); // Get relevant chunks with sources const chunks = await store .query() .from("Chunk", "c") .whereNode("c", (c) => c.embedding.similarTo(queryEmbedding, 5, { metric: "cosine", minScore: 0.7 }) ) .traverse("containsChunk", "e", { direction: "in" }) .to("Document", "d") .select((ctx) => ({ text: ctx.c.text, source: ctx.d.title })) .execute(); // Get entity relationships from graph const entityFacts = await Promise.all( extractedEntities.map(async (name) => { const relations = await store .query() .from("Entity", "e") .whereNode("e", (e) => e.name.eq(name)) .traverse("relatesTo", "r") .to("Entity", "target") .select((ctx) => ctx.target.name) .execute(); return relations.length > 0 ? { name, relatedTo: relations } : undefined; }) ); return { chunks, entityFacts: entityFacts.filter(Boolean) }; } function formatForPrompt(context: Awaited>): string { let prompt = "## Relevant Passages\n\n"; for (const chunk of context.chunks) { prompt += `**${chunk.source}**: ${chunk.text}\n\n`; } if (context.entityFacts.length > 0) { prompt += "## Entity Relationships\n\n"; for (const entity of context.entityFacts) { if (entity) { prompt += `**${entity.name}** → ${entity.relatedTo.join(", ")}\n`; } } } return prompt; } ``` ## When to Use Graph RAG **Use graph RAG when:** - Queries require connecting facts across documents ("Who founded X and what else did they start?") - Entity disambiguation matters (distinguishing "Apple" the company from "apple" the fruit) - Relationship traversal provides value ("Find all companies in the same industry as X") - You need structured facts alongside unstructured text **Flat vector RAG may suffice when:** - Simple "find similar content" queries - No entity relationships to exploit - Single-document question answering ## Next Steps - [Semantic Search](/semantic-search) — Vector embedding fundamentals - [Traversals](/queries/traverse) — Graph traversal patterns - [Document Management](/examples/document-management) — Versioning and access control # Multi-Tenant SaaS > Complete multi-tenancy patterns with isolation, data partitioning, and tenant management This example shows how to build a multi-tenant SaaS application with: - **Three isolation strategies** (shared tables, schema per tenant, database per tenant) - **Tenant-aware queries** that automatically filter data - **Tenant provisioning** and lifecycle management - **Cross-tenant analytics** for platform operators - **Tenant migration** between isolation levels ## Choosing an Isolation Strategy | Strategy | Isolation | Complexity | Cost | Best For | |----------|-----------|------------|------|----------| | Shared tables | Low | Low | Lowest | Many small tenants, B2C SaaS | | Schema per tenant | Medium | Medium | Low | SMB customers, PostgreSQL only | | Database per tenant | High | High | Highest | Enterprise, compliance requirements | ## Strategy 1: Shared Tables with Row-Level Isolation All tenants share the same database tables, filtered by `tenantId`. ### Schema Definition ```typescript import { z } from "zod"; import { defineNode, defineEdge, defineGraph, searchable } from "@nicia-ai/typegraph"; // Tenant metadata const Tenant = defineNode("Tenant", { schema: z.object({ slug: z.string(), name: z.string(), plan: z.enum(["free", "starter", "pro", "enterprise"]), status: z.enum(["active", "suspended", "cancelled"]).default("active"), createdAt: z.string().datetime(), settings: z.record(z.unknown()).optional(), }), }); // All entities include tenantId const Project = defineNode("Project", { schema: z.object({ tenantId: z.string(), // Searchable fields so tenants can run BM25 search against their // own projects without paying for an external search service. // Fulltext filtering composes with the tenantId predicate — every // query is authoritatively tenant-scoped. name: searchable({ language: "english" }), description: searchable({ language: "english" }).optional(), status: z.enum(["active", "archived"]).default("active"), }), }); const Task = defineNode("Task", { schema: z.object({ tenantId: z.string(), title: searchable({ language: "english" }), status: z.enum(["todo", "in_progress", "done"]).default("todo"), priority: z.enum(["low", "medium", "high"]).default("medium"), }), }); const User = defineNode("User", { schema: z.object({ tenantId: z.string(), email: z.string().email(), name: z.string(), role: z.enum(["owner", "admin", "member", "guest"]).default("member"), }), }); // Edges const hasProject = defineEdge("hasProject"); const hasTask = defineEdge("hasTask"); const assignedTo = defineEdge("assignedTo"); const memberOf = defineEdge("memberOf"); const graph = defineGraph({ id: "multi_tenant", nodes: { Tenant: { type: Tenant, unique: [ { name: "tenant_slug", fields: ["slug"], scope: "kind", collation: "binary", }, ], }, Project: { type: Project }, Task: { type: Task }, User: { type: User, unique: [ // Emails are scoped per tenant in the shared-tables strategy: the // same address can be a member of more than one tenant, but not // twice in the same one. { name: "user_tenant_email", fields: ["tenantId", "email"], scope: "kind", collation: "caseInsensitive", }, ], }, }, edges: { hasProject: { type: hasProject, from: [Tenant], to: [Project] }, hasTask: { type: hasTask, from: [Project], to: [Task] }, assignedTo: { type: assignedTo, from: [Task], to: [User] }, memberOf: { type: memberOf, from: [User], to: [Tenant] }, }, }); ``` ### Tenant-Scoped Store Create a wrapper that automatically filters by tenant: ```typescript interface TenantContext { tenantId: string; userId: string; role: "owner" | "admin" | "member" | "guest"; } function createTenantStore(store: Store, ctx: TenantContext) { const projects = { async list(options: { status?: string } = {}) { let query = store .query() .from("Project", "p") .whereNode("p", (p) => p.tenantId.eq(ctx.tenantId)); if (options.status) { query = query.whereNode("p", (p) => p.status.eq(options.status)); } return query.select((q) => q.p).execute(); }, async create(data: { name: string; description?: string }) { const project = await store.nodes.Project.create({ ...data, tenantId: ctx.tenantId, }); const tenant = await store.nodes.Tenant.getById(ctx.tenantId); if (!tenant) throw new Error(`Tenant not found: ${ctx.tenantId}`); await store.edges.hasProject.create(tenant, project, {}); return project; }, async get(projectId: string) { const project = await store.nodes.Project.getById(projectId); if (!project || project.tenantId !== ctx.tenantId) { throw new Error("Not found"); } return project; }, async update(projectId: string, updates: Partial) { await projects.get(projectId); // Verify access return store.nodes.Project.update(projectId, updates); }, async delete(projectId: string) { await projects.get(projectId); // Verify access await store.nodes.Project.delete(projectId); }, }; const tasks = { async list(projectId: string) { await projects.get(projectId); // Verify access return store .query() .from("Project", "p") .whereNode("p", (p) => p.id.eq(projectId)) .traverse("hasTask", "e") .to("Task", "t") .select((q) => q.t) .execute(); }, async create(projectId: string, data: { title: string; priority?: string }) { const project = await projects.get(projectId); // Verify access const task = await store.nodes.Task.create({ ...data, tenantId: ctx.tenantId, }); await store.edges.hasTask.create(project, task, {}); return task; }, }; const users = { async list() { return store .query() .from("User", "u") .whereNode("u", (u) => u.tenantId.eq(ctx.tenantId)) .select((q) => q.u) .execute(); }, async invite(email: string, name: string, role: string) { if (ctx.role !== "owner" && ctx.role !== "admin") { throw new Error("Insufficient permissions"); } const user = await store.nodes.User.create({ tenantId: ctx.tenantId, email, name, role, }); const tenant = await store.nodes.Tenant.getById(ctx.tenantId); if (!tenant) throw new Error(`Tenant not found: ${ctx.tenantId}`); await store.edges.memberOf.create(user, tenant, {}); return user; }, }; return { projects, tasks, users }; } // Usage in API handler async function handleRequest(req: Request) { const session = await getSession(req); const tenantStore = createTenantStore(store, { tenantId: session.tenantId, userId: session.userId, role: session.role, }); // All queries are automatically tenant-scoped const projects = await tenantStore.projects.list(); } ``` ### Tenant Provisioning ```typescript async function provisionTenant( slug: string, name: string, ownerEmail: string, ownerName: string, plan: "free" | "starter" | "pro" | "enterprise" = "free" ): Promise<{ tenant: Node; owner: Node }> { return store.transaction(async (tx) => { // Atomic uniqueness check — the `tenant_slug` constraint guarantees // concurrent callers can't both succeed. const tenantResult = await tx.nodes.Tenant.getOrCreateByConstraint( "tenant_slug", { slug, name, plan, status: "active", createdAt: new Date().toISOString(), }, ); if (tenantResult.action !== "created") { throw new Error("Tenant slug already exists"); } const owner = await tx.nodes.User.create({ tenantId: tenantResult.node.id, email: ownerEmail, name: ownerName, role: "owner", }); await tx.edges.memberOf.create(owner, tenantResult.node, {}); return { tenant: tenantResult.node, owner }; }); } ``` ## Strategy 2: Schema Per Tenant (PostgreSQL) Each tenant gets their own PostgreSQL schema within the same database. ### Setup ```typescript import { Pool } from "pg"; import { drizzle } from "drizzle-orm/node-postgres"; import { sql } from "drizzle-orm"; import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres"; const pool = new Pool({ connectionString: process.env.DATABASE_URL }); async function createTenantSchema(tenantId: string): Promise { const schemaName = `tenant_${tenantId}`; // Create schema await pool.query(`CREATE SCHEMA IF NOT EXISTS ${schemaName}`); // Run TypeGraph migrations in the tenant schema await pool.query(`SET search_path TO ${schemaName}`); await pool.query(generatePostgresMigrationSQL()); await pool.query(`SET search_path TO public`); } async function getTenantStore(tenantId: string): Promise { const schemaName = `tenant_${tenantId}`; // Create connection with schema const client = await pool.connect(); await client.query(`SET search_path TO ${schemaName}`); const db = drizzle(client); const backend = createPostgresBackend(db); // `searchable()` fields require the durable fulltext-materialization // step `createStoreWithSchema` performs at boot; bare `createStore()` // would throw `StoreNotInitializedError` on the first fulltext op. const [store] = await createStoreWithSchema(graph, backend); return store; } ``` ### Tenant Store Cache ```typescript class TenantStoreManager { private stores = new Map(); private maxCached = 100; async getStore(tenantId: string): Promise { const cached = this.stores.get(tenantId); if (cached) { cached.lastUsed = new Date(); return cached.store; } // Evict oldest if at capacity if (this.stores.size >= this.maxCached) { this.evictOldest(); } const store = await getTenantStore(tenantId); this.stores.set(tenantId, { store, lastUsed: new Date() }); return store; } private evictOldest(): void { let oldest: { id: string; date: Date } | undefined; for (const [id, { lastUsed }] of this.stores) { if (!oldest || lastUsed < oldest.date) { oldest = { id, date: lastUsed }; } } if (oldest) { this.stores.delete(oldest.id); } } } const tenantManager = new TenantStoreManager(); ``` ### Provisioning with Schema ```typescript async function provisionTenantWithSchema( slug: string, name: string, ownerEmail: string ): Promise<{ tenantId: string }> { const tenantId = generateUUID(); // Create schema and tables await createTenantSchema(tenantId); // Get tenant-specific store const tenantStore = await tenantManager.getStore(tenantId); // Create initial data await tenantStore.nodes.User.create({ email: ownerEmail, name: name, role: "owner", }); // Store tenant metadata in public schema const publicDb = drizzle(pool); await publicDb.insert(tenants).values({ id: tenantId, slug, name, createdAt: new Date(), }); return { tenantId }; } ``` ## Strategy 3: Database Per Tenant Each tenant gets their own database for maximum isolation. ### Tenant Database Manager ```typescript interface TenantConfig { id: string; slug: string; databaseUrl: string; status: "active" | "suspended"; } class TenantDatabaseManager { private connections = new Map(); private maxConnections = 50; async getStore(tenantId: string): Promise { const cached = this.connections.get(tenantId); if (cached) return cached.store; // Get tenant config from central registry const config = await this.getTenantConfig(tenantId); if (config.status !== "active") { throw new Error("Tenant is not active"); } // Evict if at capacity if (this.connections.size >= this.maxConnections) { await this.evictLeastUsed(); } // Create new connection const pool = new Pool({ connectionString: config.databaseUrl, max: 5 }); const db = drizzle(pool); const backend = createPostgresBackend(db); const [store] = await createStoreWithSchema(graph, backend); this.connections.set(tenantId, { pool, store }); return store; } async closeConnection(tenantId: string): Promise { const conn = this.connections.get(tenantId); if (conn) { await conn.pool.end(); this.connections.delete(tenantId); } } private async getTenantConfig(tenantId: string): Promise { // Fetch from central tenant registry const result = await centralDb .select() .from(tenantConfigs) .where(eq(tenantConfigs.id, tenantId)) .get(); if (!result) throw new Error("Tenant not found"); return result; } private async evictLeastUsed(): Promise { // Simple LRU eviction const first = this.connections.keys().next().value; if (first) { await this.closeConnection(first); } } } const dbManager = new TenantDatabaseManager(); ``` ### Provisioning New Database ```typescript async function provisionTenantDatabase( slug: string, name: string, ownerEmail: string ): Promise<{ tenantId: string; databaseUrl: string }> { const tenantId = generateUUID(); const dbName = `tenant_${tenantId.replace(/-/g, "_")}`; // Create database (using admin connection) const adminPool = new Pool({ connectionString: process.env.ADMIN_DATABASE_URL }); await adminPool.query(`CREATE DATABASE ${dbName}`); await adminPool.end(); // Build connection URL const baseUrl = new URL(process.env.DATABASE_BASE_URL!); baseUrl.pathname = `/${dbName}`; const databaseUrl = baseUrl.toString(); // Initialize TypeGraph tables const tenantPool = new Pool({ connectionString: databaseUrl }); await tenantPool.query(generatePostgresMigrationSQL()); // Create initial data const db = drizzle(tenantPool); const backend = createPostgresBackend(db); const [store] = await createStoreWithSchema(graph, backend); await store.nodes.User.create({ email: ownerEmail, name: name, role: "owner", }); await tenantPool.end(); // Register in central tenant registry await centralDb.insert(tenantConfigs).values({ id: tenantId, slug, name, databaseUrl, status: "active", createdAt: new Date(), }); return { tenantId, databaseUrl }; } ``` ## Cross-Tenant Operations For platform administrators who need to query across tenants. ### Aggregated Metrics (Shared Tables) ```typescript import { count, field } from "@nicia-ai/typegraph"; async function getTenantMetrics(): Promise< Array<{ tenantId: string; projectCount: number; taskCount: number; userCount: number }> > { // Projects by tenant const projectCounts = await store .query() .from("Project", "p") .groupBy("p", "tenantId") .aggregate({ tenantId: field("p", "tenantId"), projectCount: count("p"), }) .execute(); // Tasks by tenant const taskCounts = await store .query() .from("Task", "t") .groupBy("t", "tenantId") .aggregate({ tenantId: field("t", "tenantId"), taskCount: count("t"), }) .execute(); // Users by tenant const userCounts = await store .query() .from("User", "u") .groupBy("u", "tenantId") .aggregate({ tenantId: field("u", "tenantId"), userCount: count("u"), }) .execute(); // Merge results const metrics = new Map(); for (const p of projectCounts) { metrics.set(p.tenantId, { projectCount: p.projectCount, taskCount: 0, userCount: 0 }); } for (const t of taskCounts) { const existing = metrics.get(t.tenantId) || { projectCount: 0, taskCount: 0, userCount: 0 }; existing.taskCount = t.taskCount; metrics.set(t.tenantId, existing); } for (const u of userCounts) { const existing = metrics.get(u.tenantId) || { projectCount: 0, taskCount: 0, userCount: 0 }; existing.userCount = u.userCount; metrics.set(u.tenantId, existing); } return Array.from(metrics.entries()).map(([tenantId, counts]) => ({ tenantId, ...counts, })); } ``` ### Cross-Tenant Search (Database Per Tenant) Fulltext search composes with tenant scope: the fulltext predicate narrows the candidate pool by relevance, and the per-tenant store supplies the isolation. Nothing from tenant A can ever appear in tenant B's results because each store wraps a different database: ```typescript async function searchAcrossTenants( query: string, tenantIds: string[] ): Promise }>> { const results = await Promise.all( tenantIds.map(async (tenantId) => { try { const store = await dbManager.getStore(tenantId); const hits = await store.search.fulltext("Project", { query, limit: 10, includeSnippets: true, }); return { tenantId, results: hits.map((hit) => ({ ...hit.node, score: hit.score, snippet: hit.snippet, })), }; } catch (error) { console.error(`Failed to search tenant ${tenantId}:`, error); return { tenantId, results: [] }; } }) ); return results; } ``` ### Tenant-Scoped Fulltext (Shared Tables) The composition most multi-tenant teams need: BM25 ranking *plus* a tenantId filter in a single query. The fulltext predicate scales with the tenant's data, not the whole table: ```typescript async function searchTenantProjects( tenantId: string, query: string, limit = 10, ) { return store .query() .from("Project", "p") .whereNode("p", (p) => // Tenant filter first so the fulltext match is always scoped. p.tenantId.eq(tenantId).and(p.$fulltext.matches(query, limit)), ) .select((ctx) => ctx.p) .execute(); } ``` The tenant filter and fulltext match compose in one SQL statement; no post-filter in JS and no risk of leaking another tenant's data even if the caller forgets to verify. ## Tenant Lifecycle ### Suspend Tenant ```typescript async function suspendTenant(tenantId: string, reason: string): Promise { const current = await store.nodes.Tenant.getById(tenantId); if (!current) throw new Error(`Tenant not found: ${tenantId}`); await store.nodes.Tenant.update(tenantId, { status: "suspended", settings: { ...(current.settings || {}), suspendedAt: new Date().toISOString(), suspendReason: reason, }, }); } ``` ### Delete Tenant (Shared Tables) ```typescript async function deleteTenant(tenantId: string): Promise { await store.transaction(async (tx) => { // Delete all tasks const tasks = await tx .query() .from("Task", "t") .whereNode("t", (t) => t.tenantId.eq(tenantId)) .select((ctx) => ctx.t.id) .execute(); for (const taskId of tasks) { await tx.nodes.Task.delete(taskId); } // Delete all projects const projects = await tx .query() .from("Project", "p") .whereNode("p", (p) => p.tenantId.eq(tenantId)) .select((ctx) => ctx.p.id) .execute(); for (const projectId of projects) { await tx.nodes.Project.delete(projectId); } // Delete all users const users = await tx .query() .from("User", "u") .whereNode("u", (u) => u.tenantId.eq(tenantId)) .select((ctx) => ctx.u.id) .execute(); for (const userId of users) { await tx.nodes.User.delete(userId); } // Delete tenant await tx.nodes.Tenant.delete(tenantId); }); } ``` ### Delete Tenant (Database Per Tenant) ```typescript async function deleteTenantDatabase(tenantId: string): Promise { // Close active connection await dbManager.closeConnection(tenantId); // Get database name const config = await getTenantConfig(tenantId); const dbUrl = new URL(config.databaseUrl); const dbName = dbUrl.pathname.slice(1); // Drop database const adminPool = new Pool({ connectionString: process.env.ADMIN_DATABASE_URL }); await adminPool.query(`DROP DATABASE IF EXISTS ${dbName}`); await adminPool.end(); // Remove from registry await centralDb.delete(tenantConfigs).where(eq(tenantConfigs.id, tenantId)); } ``` ## Tenant Migration Move tenant between isolation strategies: ```typescript async function migrateTenantToSeparateDatabase(tenantId: string): Promise { // 1. Create new database const { databaseUrl } = await provisionTenantDatabase( `migrated_${tenantId}`, "Migrated Tenant", "placeholder@example.com" ); // 2. Get tenant data from shared tables const sharedStore = store; const projects = await sharedStore .query() .from("Project", "p") .whereNode("p", (p) => p.tenantId.eq(tenantId)) .select((ctx) => ctx.p) .execute(); const tasks = await sharedStore .query() .from("Task", "t") .whereNode("t", (t) => t.tenantId.eq(tenantId)) .select((ctx) => ctx.t) .execute(); const users = await sharedStore .query() .from("User", "u") .whereNode("u", (u) => u.tenantId.eq(tenantId)) .select((ctx) => ctx.u) .execute(); // 3. Insert into new database const newStore = await dbManager.getStore(tenantId); await newStore.transaction(async (tx) => { for (const project of projects) { await tx.nodes.Project.create(project); } for (const task of tasks) { await tx.nodes.Task.create(task); } for (const user of users) { await tx.nodes.User.create(user); } }); // 4. Delete from shared tables await deleteTenant(tenantId); return databaseUrl; } ``` ## Next Steps - [Document Management](/examples/document-management) - CMS with semantic search - [Product Catalog](/examples/product-catalog) - Categories, variants, inventory - [Integration Patterns](/integration) - More deployment strategies # Product Catalog > E-commerce catalog with categories, variants, and inventory tracking This example builds a product catalog system with: - **Category hierarchy** with inheritance - **Product variants** (size, color, etc.) - **Inventory tracking** across warehouses - **Product relationships** (bundles, accessories, alternatives) - **Price history** using temporal queries ## Schema Definition ```typescript import { z } from "zod"; import { defineNode, defineEdge, defineGraph, embedding, searchable, } from "@nicia-ai/typegraph"; // Category hierarchy const Category = defineNode("Category", { schema: z.object({ name: searchable({ language: "english" }), slug: z.string(), description: searchable({ language: "english" }).optional(), imageUrl: z.string().url().optional(), displayOrder: z.number().default(0), isActive: z.boolean().default(true), }), }); // Products const Product = defineNode("Product", { schema: z.object({ sku: z.string(), // `searchable()` enables BM25 fulltext matching on name + description. // Combined with the `embedding` field below this supports hybrid // retrieval — SKU-style exact matches that embeddings miss, plus // conceptual matches that keyword search alone miss. name: searchable({ language: "english" }), description: searchable({ language: "english" }), basePrice: z.number().positive(), currency: z.string().default("USD"), status: z.enum(["draft", "active", "discontinued"]).default("draft"), embedding: embedding(1536).optional(), }), }); // Product variants (specific size/color combinations) const Variant = defineNode("Variant", { schema: z.object({ sku: z.string(), // "Large / Blue" — variant names contain the exact tokens shoppers // type ("blue", "xl"), so indexing them enables keyword retrieval // that complements the product-level embedding. name: searchable({ language: "english" }), priceModifier: z.number().default(0), // Added to base price attributes: z.record(z.string()), // { size: "L", color: "blue" } isDefault: z.boolean().default(false), }), }); // Inventory const Warehouse = defineNode("Warehouse", { schema: z.object({ code: z.string(), name: z.string(), location: z.string(), isActive: z.boolean().default(true), }), }); const Inventory = defineNode("Inventory", { schema: z.object({ quantity: z.number().int().min(0), reservedQuantity: z.number().int().min(0).default(0), reorderPoint: z.number().int().min(0).default(10), lastCountedAt: z.string().datetime().optional(), }), }); // Edges const parentCategory = defineEdge("parentCategory"); const inCategory = defineEdge("inCategory", { schema: z.object({ isPrimary: z.boolean().default(false) }), }); const hasVariant = defineEdge("hasVariant"); const inventoryFor = defineEdge("inventoryFor"); const atWarehouse = defineEdge("atWarehouse"); const relatedProduct = defineEdge("relatedProduct", { schema: z.object({ type: z.enum(["accessory", "alternative", "bundled", "upsell"]), sortOrder: z.number().default(0), }), }); // Graph const graph = defineGraph({ id: "product_catalog", nodes: { Category: { type: Category, unique: [ { name: "category_slug", fields: ["slug"], scope: "kind", collation: "binary", }, ], }, Product: { type: Product, unique: [ { name: "product_sku", fields: ["sku"], scope: "kind", collation: "binary", }, ], }, Variant: { type: Variant, unique: [ { name: "variant_sku", fields: ["sku"], scope: "kind", collation: "binary", }, ], }, Warehouse: { type: Warehouse, unique: [ { name: "warehouse_code", fields: ["code"], scope: "kind", collation: "binary", }, ], }, Inventory: { type: Inventory }, }, edges: { parentCategory: { type: parentCategory, from: [Category], to: [Category] }, inCategory: { type: inCategory, from: [Product], to: [Category] }, hasVariant: { type: hasVariant, from: [Product], to: [Variant] }, inventoryFor: { type: inventoryFor, from: [Inventory], to: [Variant] }, atWarehouse: { type: atWarehouse, from: [Inventory], to: [Warehouse] }, relatedProduct: { type: relatedProduct, from: [Product], to: [Product] }, }, ontology: [ // Category hierarchy is modeled via the parentCategory edge, not ontology. // Use ontology for type-level constraints, e.g.: // disjointWith(Product, Category), ], }); ``` ## Category Management ### Create Category Tree ```typescript async function createCategory( name: string, slug: string, parentSlug?: string ): Promise> { const result = await store.nodes.Category.getOrCreateByConstraint( "category_slug", { name, slug, isActive: true }, ); if (result.action === "created" && parentSlug) { const parent = await store.nodes.Category.findByConstraint( "category_slug", { slug: parentSlug }, ); if (parent) { await store.edges.parentCategory.create(result.node, parent, {}); } } return result.node; } // Build initial category structure await createCategory("Electronics", "electronics"); await createCategory("Phones", "phones", "electronics"); await createCategory("Accessories", "accessories", "electronics"); await createCategory("Cases", "cases", "accessories"); await createCategory("Chargers", "chargers", "accessories"); ``` ### Get Category with Ancestors ```typescript interface CategoryWithPath { id: string; name: string; slug: string; path: Array<{ name: string; slug: string }>; } async function getCategoryWithPath(slug: string): Promise { const category = await store.nodes.Category.findByConstraint( "category_slug", { slug }, ); if (!category) return undefined; // Walk `parentCategory` edges up to the root. `reachable` returns each // ancestor with its depth from the starting node — sort by depth desc // so the root comes first. const ancestorIds = ( await store.algorithms.reachable(category.id, { edges: ["parentCategory"], excludeSource: true, }) ) .toSorted((a, b) => b.depth - a.depth) .map((node) => node.id); const ancestors = await store.nodes.Category.getByIds(ancestorIds); return { id: category.id, name: category.name, slug: category.slug, path: ancestors .filter((c): c is NonNullable => c !== undefined) .map((c) => ({ name: c.name, slug: c.slug })), }; } ``` ### Get Subcategories ```typescript async function getSubcategories( parentSlug: string, includeNested = false ): Promise> { const parent = await store.nodes.Category.findByConstraint( "category_slug", { slug: parentSlug }, ); if (!parent) return []; // `reachable` returns descendants tagged with their depth. Cap at 1 for // immediate children only, or let it run to the configured default // (10 hops) for the full subtree. const descendants = await store.algorithms.reachable(parent.id, { edges: ["parentCategory"], direction: "in", excludeSource: true, maxHops: includeNested ? undefined : 1, }); const children = (await store.nodes.Category.getByIds( descendants.map((node) => node.id), )).filter( (category): category is NonNullable => category !== undefined && category.isActive, ); const depthById = new Map(descendants.map((row) => [row.id, row.depth])); return children .map((category) => ({ id: category.id, name: category.name, slug: category.slug, depth: depthById.get(category.id) ?? 1, })) .toSorted((a, b) => a.depth - b.depth); } ``` ## Product Management ### Create Product with Variants ```typescript interface ProductInput { sku: string; name: string; description: string; basePrice: number; categorySlug: string; variants: Array<{ sku: string; name: string; priceModifier?: number; attributes: Record; isDefault?: boolean; }>; } async function createProduct(input: ProductInput): Promise> { return store.transaction(async (tx) => { // Generate embedding for semantic search const embedding = await generateEmbedding(`${input.name} ${input.description}`); // Create product const product = await tx.nodes.Product.create({ sku: input.sku, name: input.name, description: input.description, basePrice: input.basePrice, status: "draft", embedding, }); // Link to category const category = await tx .query() .from("Category", "c") .whereNode("c", (c) => c.slug.eq(input.categorySlug)) .select((ctx) => ctx.c) .first(); if (category) { await tx.edges.inCategory.create(product, category, { isPrimary: true }); } // Create variants for (const v of input.variants) { const variant = await tx.nodes.Variant.create({ sku: v.sku, name: v.name, priceModifier: v.priceModifier ?? 0, attributes: v.attributes, isDefault: v.isDefault ?? false, }); await tx.edges.hasVariant.create(product, variant, {}); } return product; }); } ``` ### Get Product Details ```typescript interface ProductDetails { id: string; sku: string; name: string; description: string; basePrice: number; status: string; categories: Array<{ name: string; slug: string; isPrimary: boolean }>; variants: Array<{ id: string; sku: string; name: string; price: number; attributes: Record; inventory: number; }>; related: Array<{ id: string; name: string; type: string }>; } async function getProductDetails(sku: string): Promise { const product = await store.nodes.Product.findByConstraint( "product_sku", { sku }, ); if (!product) return undefined; // `store.batch()` runs all three queries in sequence, one at a time — // still three statements, and read-committed isolation means a write can // land between them. It is the pool-pressure win, not a snapshot. const [categories, variants, related] = await store.batch( store .query() .from("Product", "p") .whereNode("p", (p) => p.id.eq(product.id)) .traverse("inCategory", "e") .to("Category", "c") .select((ctx) => ({ name: ctx.c.name, slug: ctx.c.slug, isPrimary: ctx.e.isPrimary, })), store .query() .from("Product", "p") .whereNode("p", (p) => p.id.eq(product.id)) .traverse("hasVariant", "e") .to("Variant", "v") .optionalTraverse("inventoryFor", "inv", { direction: "in" }) .to("Inventory", "i") .select((ctx) => ({ id: ctx.v.id, sku: ctx.v.sku, name: ctx.v.name, priceModifier: ctx.v.priceModifier, attributes: ctx.v.attributes, quantity: ctx.i?.quantity ?? 0, reservedQuantity: ctx.i?.reservedQuantity ?? 0, })), store .query() .from("Product", "p") .whereNode("p", (p) => p.id.eq(product.id)) .traverse("relatedProduct", "e") .to("Product", "r") .orderBy("e", "sortOrder", "asc") .select((ctx) => ({ id: ctx.r.id, name: ctx.r.name, type: ctx.e.type, })), ); return { id: product.id, sku: product.sku, name: product.name, description: product.description, basePrice: product.basePrice, status: product.status, categories, variants: variants.map((v) => ({ ...v, price: product.basePrice + v.priceModifier, inventory: v.quantity - v.reservedQuantity, })), related, }; } ``` ## Inventory Management ### Update Inventory ```typescript async function updateInventory( variantSku: string, warehouseCode: string, quantity: number ): Promise { const variant = await store .query() .from("Variant", "v") .whereNode("v", (v) => v.sku.eq(variantSku)) .select((ctx) => ctx.v) .first(); const warehouse = await store .query() .from("Warehouse", "w") .whereNode("w", (w) => w.code.eq(warehouseCode)) .select((ctx) => ctx.w) .first(); if (!variant || !warehouse) { throw new Error("Variant or warehouse not found"); } // Find existing inventory record const existingInventory = await store .query() .from("Inventory", "i") .traverse("inventoryFor", "e1") .to("Variant", "v") .whereNode("v", (v) => v.id.eq(variant.id)) .traverse("atWarehouse", "e2", { direction: "in" }) .to("Warehouse", "w") .whereNode("w", (w) => w.id.eq(warehouse.id)) .select((ctx) => ctx.i) .first(); if (existingInventory) { await store.nodes.Inventory.update(existingInventory.id, { quantity, lastCountedAt: new Date().toISOString(), }); } else { const inventory = await store.nodes.Inventory.create({ quantity, reservedQuantity: 0, lastCountedAt: new Date().toISOString(), }); await store.edges.inventoryFor.create(inventory, variant, {}); await store.edges.atWarehouse.create(inventory, warehouse, {}); } } ``` ### Reserve Inventory ```typescript async function reserveInventory( variantSku: string, quantity: number ): Promise<{ success: boolean; warehouseCode?: string }> { const inventories = await store .query() .from("Variant", "v") .whereNode("v", (v) => v.sku.eq(variantSku)) .traverse("inventoryFor", "e", { direction: "in" }) .to("Inventory", "i") .traverse("atWarehouse", "e2") .to("Warehouse", "w") .whereNode("w", (w) => w.isActive.eq(true)) .select((ctx) => ({ inventoryId: ctx.i.id, warehouseCode: ctx.w.code, available: ctx.i.quantity - ctx.i.reservedQuantity, reservedQuantity: ctx.i.reservedQuantity, })) .execute(); // Find warehouse with enough inventory const available = inventories.find((i) => i.available >= quantity); if (!available) { return { success: false }; } await store.nodes.Inventory.update(available.inventoryId, { reservedQuantity: available.reservedQuantity + quantity, }); return { success: true, warehouseCode: available.warehouseCode }; } ``` ### Low Stock Report ```typescript import { field, sum, havingLt } from "@nicia-ai/typegraph"; interface LowStockItem { productName: string; variantSku: string; variantName: string; totalQuantity: number; reorderPoint: number; } async function getLowStockItems(): Promise { return store .query() .from("Product", "p") .traverse("hasVariant", "e1") .to("Variant", "v") .traverse("inventoryFor", "e2", { direction: "in" }) .to("Inventory", "i") .groupByNode("v") .having(havingLt(sum("i", "quantity"), field("i", "reorderPoint"))) .aggregate({ productName: field("p", "name"), variantSku: field("v", "sku"), variantName: field("v", "name"), totalQuantity: sum("i", "quantity"), reorderPoint: field("i", "reorderPoint"), }) .execute(); } ``` ## Search and Discovery Product search is a textbook hybrid-search problem: users type SKUs, brand names, and category jargon that embeddings blur together, and conceptual queries ("warm winter jacket") that exact keyword matching misses. TypeGraph supports both in one query. ### Fulltext-only search (SKU / keyword hits) `store.search.fulltext()` runs a ranked BM25 query across every `searchable()` field on the node. Good for search-box autocomplete and SKU lookups where the query is a bag of keywords rather than a description: ```typescript const hits = await store.search.fulltext("Product", { query: "waterproof jacket", limit: 20, includeSnippets: true, }); for (const hit of hits) { console.log(hit.node.sku, hit.node.name, hit.score, hit.snippet); } ``` ### Hybrid product search (fulltext + semantic, fused) For a production product-search feature, fuse fulltext and vector retrieval with Reciprocal Rank Fusion. RRF is rank-based, so it handles the score-scale mismatch between BM25 and cosine automatically: ```typescript async function searchProducts( query: string, options: { categorySlug?: string; minPrice?: number; maxPrice?: number; limit?: number; } = {} ): Promise> { const { categorySlug, minPrice, maxPrice, limit = 20 } = options; const queryEmbedding = await generateEmbedding(query); // Limit category scope (if requested) to a set of IDs we can pass in // as an equality filter. Applying this ahead of RRF shrinks the // candidate pool before fusion — better latency and ranking quality. let categoryIds: readonly string[] | undefined; if (categorySlug) { const root = await store.nodes.Category.findByConstraint( "category_slug", { slug: categorySlug }, ); if (!root) return []; const subtree = await store.algorithms.reachable(root.id, { edges: ["parentCategory"], direction: "in", }); categoryIds = subtree.map((node) => node.id); } const hits = await store.search.hybrid("Product", { limit, vector: { fieldPath: "embedding", queryEmbedding, metric: "cosine", k: limit * 4, }, fulltext: { query, k: limit * 4, includeSnippets: true, }, // Exact-name matches matter in commerce search — boost fulltext. fusion: { method: "rrf", k: 60, weights: { vector: 1, fulltext: 1.5 } }, }); // Post-filter by status / price / category. The hybrid API does not // compose with the query builder's predicates, so apply these after // fusion. For heavier filtering, switch to the query-builder path // below and filter inside the same SQL statement. const filtered = hits.filter((hit) => { if (hit.node.status !== "active") return false; if (minPrice !== undefined && hit.node.basePrice < minPrice) return false; if (maxPrice !== undefined && hit.node.basePrice > maxPrice) return false; return true; }); if (categoryIds === undefined) { return filtered.map((hit) => ({ product: hit.node, score: hit.score, snippet: hit.fulltext?.snippet, })); } // Category membership is a graph edge — check with a single batch query. const categoryIdSet = new Set(categoryIds); const productToCategories = new Map>(); const memberships = await store .query() .from("Product", "p") .whereNode("p", (p) => p.id.in(filtered.map((hit) => hit.node.id))) .traverse("inCategory", "e") .to("Category", "c") .select((ctx) => ({ productId: ctx.p.id, categoryId: ctx.c.id })) .execute(); for (const row of memberships) { const cats = productToCategories.get(row.productId) ?? new Set(); cats.add(row.categoryId); productToCategories.set(row.productId, cats); } return filtered .filter((hit) => { const cats = productToCategories.get(hit.node.id) ?? new Set(); return [...cats].some((id) => categoryIdSet.has(id)); }) .map((hit) => ({ product: hit.node, score: hit.score, snippet: hit.fulltext?.snippet, })); } ``` ### Hybrid search composed with graph traversal (query builder) When you need tighter composition with predicates and traversals — for example, "only products in these categories, active, in stock" — use the query builder. `$fulltext.matches()` and `.similarTo()` in the same `whereNode()` compile to a two-CTE SQL statement with RRF at the ORDER BY: ```typescript const hits = await store .query() .from("Product", "p") .whereNode("p", (p) => p.$fulltext .matches(query, limit * 4) .and(p.embedding.similarTo(queryEmbedding, limit * 4)) .and(p.status.eq("active")), ) .traverse("inCategory", "e") .to("Category", "c") .whereNode("c", (c) => c.id.in([...categoryIdSet])) .fuseWith({ k: 60, weights: { vector: 1, fulltext: 1.5 } }) .select((ctx) => ctx.p) .limit(limit) .execute(); ``` Results come back already ranked by the fused RRF score. The traversal filter is applied inside the same SQL statement, before the final `LIMIT`, so recall is not sacrificed for composition. ### Get Products in Category ```typescript async function getProductsInCategory( categorySlug: string, options: { includeSubcategories?: boolean; page?: number; pageSize?: number; sortBy?: "name" | "price" | "newest"; } = {} ): Promise<{ products: ProductProps[]; total: number }> { const { includeSubcategories = true, page = 1, pageSize = 20, sortBy = "name" } = options; const root = await store.nodes.Category.findByConstraint( "category_slug", { slug: categorySlug }, ); if (!root) return { products: [], total: 0 }; const categoryIds = includeSubcategories ? ( await store.algorithms.reachable(root.id, { edges: ["parentCategory"], direction: "in", }) ).map((node) => node.id) : [root.id]; const query = store .query() .from("Product", "p") .whereNode("p", (p) => p.status.eq("active")) .traverse("inCategory", "e") .to("Category", "c") .whereNode("c", (c) => c.id.in(categoryIds)) .select((ctx) => ctx.p); // Apply sorting const sortedQuery = sortBy === "price" ? query.orderBy((ctx) => ctx.p.basePrice, "asc") : sortBy === "newest" ? query.orderBy((ctx) => ctx.p.createdAt, "desc") : query.orderBy((ctx) => ctx.p.name, "asc"); const products = await sortedQuery .limit(pageSize) .offset((page - 1) * pageSize) .execute(); const total = await store .query() .from("Product", "p") .whereNode("p", (p) => p.status.eq("active")) .traverse("inCategory", "e") .to("Category", "c") .whereNode("c", (c) => c.id.in(categoryIds)) .count(); return { products, total }; } ``` ## Price History ### Track Price Changes TypeGraph's temporal model automatically tracks all changes: ```typescript async function getPriceHistory( sku: string ): Promise> { return store .query() .from("Product", "p") .temporal("includeEnded") .whereNode("p", (p) => p.sku.eq(sku)) .orderBy((ctx) => ctx.p.validFrom, "desc") .select((ctx) => ({ price: ctx.p.basePrice, validFrom: ctx.p.validFrom, validTo: ctx.p.validTo, })) .execute(); } ``` ### Price at Point in Time ```typescript async function getPriceAsOf(sku: string, date: Date): Promise { const product = await store .query() .from("Product", "p") .temporal("asOf", date.toISOString()) .whereNode("p", (p) => p.sku.eq(sku)) .select((ctx) => ctx.p.basePrice) .first(); return product; } ``` ## Next Steps - [Document Management](/examples/document-management) - CMS with semantic search - [Workflow Engine](/examples/workflow-engine) - State machines with approvals - [Audit Trail](/examples/audit-trail) - Complete change tracking # Provenance Retraction > Retract bad sources, keep facts with alternate support, and replay belief changes with recorded time This example shows provenance retraction on a derived knowledge graph. Distinct source kinds support fact kinds through explicit justification nodes. When a source is retracted, TypeGraph recomputes the well-founded support set: - facts with alternate support stay current - terminal facts can be derived without being valid premises themselves - facts with no remaining support become non-current - recorded-time reads can replay what the graph believed before and after :::tip[Just want the code?] Full source on GitHub: [`packages/typegraph/examples/23-provenance-retraction.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/23-provenance-retraction.ts) ::: ## What It Demonstrates - `@nicia-ai/typegraph/provenance` as a first-class package subpath. - Multiple source node kinds via `source: { kinds: [...] }`. - A terminal `DeployDecision` fact kind that is derived but is not allowed in `premiseOf.from`. - `createRetractionCapability()` over a store created with `{ history: true }`. - `holding()` for current well-founded believed facts. - `retract()` reports where facts died or survived through alternate justifications. - `store.asOfRecorded(T)` replays the fact currency before and after the retraction. - `unRetract()` clears a source's retraction flag and reopens supported facts. ## Run It From the repository root: ```bash pnpm --filter @nicia-ai/typegraph exec tsx examples/23-provenance-retraction.ts ``` Or from `packages/typegraph`: ```bash npx tsx examples/23-provenance-retraction.ts ``` The example uses an in-memory SQLite backend with `history: true`, so it does not require Docker or external services. ## Graph Shape The example models a security advisory workflow: ```text scanner-source -> justification-scanner-finding -> vulnerability-libvector vendor-source -> justification-vendor-advisory -> vulnerability-libvector vulnerability-libvector -> justification-block-deploy -> decision-block-deploy ``` `vulnerability-libvector` has two independent source-level supports. `decision-block-deploy` depends on `vulnerability-libvector`, so it should survive as long as the vulnerability fact still has at least one grounded support path. The graph is ordinary TypeGraph schema: ```typescript const ScannerSource = defineNode("ScannerSource", { schema: z.object({ title: z.string(), retracted: z.boolean().default(false), }), }); const VendorSource = defineNode("VendorSource", { schema: z.object({ title: z.string(), retracted: z.boolean().default(false), }), }); const Vulnerability = defineNode("Vulnerability", { schema: z.object({ cve: z.string(), packageName: z.string() }), }); const DeployDecision = defineNode("DeployDecision", { schema: z.object({ action: z.string() }), }); ``` `DeployDecision` is terminal: it is a fact kind and a valid derivation target, but it is not a premise kind. ```typescript const graph = defineGraph({ edges: { premiseOf: { type: premiseOf, from: [ScannerSource, VendorSource, Vulnerability], to: [Justification], }, derives: { type: derives, from: [Justification], to: [Vulnerability, DeployDecision], }, }, }); ``` Then the roles are mapped into provenance retraction: ```typescript const provenance = createRetractionCapability(store, { source: { kinds: ["ScannerSource", "VendorSource"] }, justification: { kind: "Justification" }, fact: { kinds: ["Vulnerability", "DeployDecision"] }, premiseOf: { kind: "premiseOf" }, derives: { kind: "derives" }, }); ``` ## Retraction Retracting the scanner source does not kill the facts, because the vendor advisory still supports `vulnerability-libvector`: ```typescript const scannerReport = await provenance.retract(scanner); console.log(scannerReport.survivedVia); ``` Retracting the vendor source too removes the last grounded support. The vulnerability and deploy decision become non-current, but the recorded relation still knows when they were believed: ```typescript const beforeVendorRetraction = await store.recordedNow(); if (beforeVendorRetraction === undefined) throw new Error("expected history"); await provenance.retract(vendor); const afterVendorRetraction = await store.recordedNow(); if (afterVendorRetraction === undefined) throw new Error("expected history"); const before = await store.asOfRecorded(beforeVendorRetraction).nodes.DeployDecision.getById(blockDeploy.id); const after = await store.asOfRecorded(afterVendorRetraction).nodes.DeployDecision.getById(blockDeploy.id); ``` ## Sample Output ```text Initial derived beliefs: holding(): decision-block-deploy, vulnerability-libvector Retract the unverified scanner source. died: (none) survived: decision-block-deploy via justification-block-deploy; vulnerability-libvector via justification-vendor-advisory unaffected: (none) holding(): decision-block-deploy, vulnerability-libvector Retract the vendor advisory too. died: decision-block-deploy, vulnerability-libvector survived: (none) unaffected: (none) holding(): (none) Recorded-time replay of the deploy-block fact: before vendor retraction: Block the production deploy after vendor retraction: not current ``` ## When to Use This Pattern Use provenance retraction when derived facts must respond to source quality: - AI memory and RAG citations where a source document is later invalidated - security or compliance knowledge graphs with advisory retractions - ingestion pipelines where downstream facts depend on upstream source trust - audit workflows that need both current belief state and past belief replay See [Provenance and Retraction](/provenance) for the API guide and [Temporal queries](/queries/temporal#recorded-time-bitemporal) for the recorded-time rules. # Research Copilot > Semantic search, ontology expansion, and graph algorithms combined into an explainable literature-review digest over a citation graph A single runnable example that exercises nearly every TypeGraph capability — typed schema, [ontology](/ontology), vector embeddings, recursive traversals, and the [graph algorithms](/graph-algorithms) — against a corpus of landmark ML papers. It produces an explainable literature-review digest in one run against a single SQLite file, with zero external services. :::tip[Just want the code?] Full source on GitHub: [`packages/typegraph/examples/14-research-copilot.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/14-research-copilot.ts) ::: ## What You Get A natural-language query comes in. The copilot returns a ranked, chronological reading list with citation counts, authors, and topics — all computed against a single in-memory SQLite database: ```text Query: "contrastive self-supervised representation learning for vision" Recommended reading order (chronological among top-ranked): 2012 ImageNet Classification with Deep Convolutional Neural Networks [3 citations] Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton topics: CNN, ComputerVision, DeepLearning why: semantic 0.449 · topic match: DeepLearning · 3 incoming citations 2014 Adam: A Method for Stochastic Optimization [1 citation] Diederik Kingma, Jimmy Ba topics: Optimization why: semantic 0.429 · 1 incoming citation 2019 Momentum Contrast for Unsupervised Visual Representation Learning [1 citation] Kaiming He, Haoqi Fan, Yuxin Wu, et al. topics: Contrastive, SelfSupervised, ComputerVision why: semantic 0.523 · topic match: SelfSupervised, Contrastive · 1 incoming citation 2020 A Simple Framework for Contrastive Learning of Visual Representations [1 citation] Ting Chen, Simon Kornblith, Mohammad Norouzi, et al. topics: Contrastive, SelfSupervised, ComputerVision why: semantic 0.436 · topic match: SelfSupervised, Contrastive · 1 incoming citation 2020 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale [1 citation] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. topics: Transformer, ComputerVision, DeepLearning why: semantic 0.350 · topic match: DeepLearning · 1 incoming citation ``` ## Architecture Each moving part maps to a single TypeGraph primitive: | Feature | TypeGraph capability | | ------------------------------------ | ----------------------------------------------------------- | | Semantic paper retrieval | `embedding()` fields + cosine similarity | | Topic hierarchy expansion | [Ontology](/ontology) + `store.algorithms.reachable()` | | Citation-authority ranking | `store.algorithms.degree()` over `cites` | | Explainable paper lineage | `store.algorithms.shortestPath()` over `cites` | | "Does this trace back to X?" | `store.algorithms.canReach()` | | Co-author discovery (2-hop) | `store.algorithms.neighbors()` | | Reading-list assembly | [Query builder](/queries/overview) with typed traversals | ## Schema Three node kinds and four edges model the citation graph plus a topic hierarchy that supports query expansion: ```typescript const Paper = defineNode("Paper", { schema: z.object({ // Title + abstract are `searchable()` so BM25 ranks papers by // keyword hits (rare technical terms, author surnames, dataset // names) — exactly the queries where embeddings are least // discriminative. title: searchable({ language: "english" }), year: z.number().int(), abstract: searchable({ language: "english" }), embedding: embedding(128), }), }); const Author = defineNode("Author", { schema: z.object({ name: z.string() }), }); const Topic = defineNode("Topic", { schema: z.object({ name: z.string() }), }); const cites = defineEdge("cites", { schema: z.object({}) }); const authoredBy = defineEdge("authored_by", { schema: z.object({}) }); const coversTopic = defineEdge("covers_topic", { schema: z.object({}) }); // Topic hierarchy: `CNN broader_than DL` reads "CNN is a more specific // concept than DL". Recursive traversal expands narrow query terms into // their ancestor concepts for higher recall. const broaderThan = defineEdge("broader_than", { schema: z.object({}) }); const graph = defineGraph({ id: "research_copilot", nodes: { Paper: { type: Paper }, Author: { type: Author }, Topic: { type: Topic } }, edges: { cites: { type: cites, from: [Paper], to: [Paper] }, authored_by: { type: authoredBy, from: [Paper], to: [Author] }, covers_topic: { type: coversTopic, from: [Paper], to: [Topic] }, broader_than: { type: broaderThan, from: [Topic], to: [Topic] }, }, }); ``` ## Scene by Scene The example walks through five capabilities end-to-end. Each produces real console output against the seeded corpus of 18 landmark papers. ### 1. Semantic retrieval Every paper has a 128-dimensional embedding. Rank the corpus against a query embedding and take the top hits: ```typescript const queryEmbedding = mockEmbedding(query); const allPapers = await store.nodes.Paper.find(); const ranked = allPapers .map((paper) => ({ paper, similarity: cosine(queryEmbedding, paper.embedding), })) .sort((a, b) => b.similarity - a.similarity); ``` In production, swap the in-JS ranking for `p.embedding.similarTo(queryEmbedding, k)` in a [query builder](/queries/overview) predicate — backed by pgvector or sqlite-vec — to do the scoring in SQL. See [Semantic Search](/semantic-search). `title` and `abstract` are declared `searchable()`, so the same corpus is also indexed for BM25 via SQLite's FTS5. The example runs a rare-token query against the fulltext index to show where BM25 wins — dataset names, method acronyms, proper nouns — exactly the queries embeddings smooth out: ```typescript const fulltextHits = await store.search.fulltext("Paper", { query: "Dropout", limit: 3, includeSnippets: true, }); ``` ```text ─── Fulltext retrieval (BM25 via FTS5) for: "Dropout" ─── 2.619 Dropout: A Simple Way to Prevent Neural Networks from Overfitting Dropout: A Simple Way to Prevent Neural Networks from Overfitting Randomly zeroing unit activations during training prevents co-adaptation and… ``` In production you'd fuse the two via `store.search.hybrid()`, which runs both retrievers and blends them with Reciprocal Rank Fusion at the SQL layer: ```typescript const hits = await store.search.hybrid("Paper", { limit: 10, vector: { fieldPath: "embedding", queryEmbedding, metric: "cosine" }, fulltext: { query, includeSnippets: true }, // Weight fulltext slightly higher for the entity-heavy queries // typical of literature search. fusion: { method: "rrf", k: 60, weights: { vector: 1, fulltext: 1.25 } }, }); ``` See the [Fulltext Search guide](/fulltext-search) for tuning and [Example 15](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/15-fulltext-hybrid-search.ts) for an end-to-end hybrid walkthrough. ### 2. Ontology-expanded topic matching A query for the narrow topic `Contrastive` should also return papers tagged with its ancestors (`SelfSupervised`, `DeepLearning`). `reachable()` walks the `broader_than` edge recursively and returns every ancestor topic: ```typescript const topicAncestors = await store.algorithms.reachable(contrastiveTopic, { edges: ["broader_than"], maxHops: 10, excludeSource: true, }); ``` Then filter papers whose `covers_topic` edge lands in the expanded set: ```typescript const topicMatches = await store .query() .from("Paper", "p") .traverse("covers_topic", "e") .to("Topic", "t") .whereNode("t", (t) => t.id.in([...expandedTopicIds])) .select((ctx) => ({ id: ctx.p.id, title: ctx.p.title, topic: ctx.t.name })) .execute(); ``` Output: ```text Expanded set: {Contrastive, SelfSupervised, DeepLearning} ``` ### 3. Citation-authority re-ranking Pure vector similarity is noisy. Fuse it with in-degree on the `cites` edge so highly-cited papers bubble up: ```typescript const citationCount = await store.algorithms.degree(paperId, { edges: ["cites"], direction: "in", }); const score = similarity + topicBonus + Math.log(citationCount + 1) / 10; ``` Output: ```text score = similarity + 0.05 * topicMatches + log(1 + citations) / 10 rank score sim topic cites title ─────────────────────────────────────────────────────────────────── 1 0.692 0.523 2 1 Momentum Contrast for Unsupervised Visual Representation Learning 2 0.638 0.449 1 3 ImageNet Classification with Deep Convolutional Neural Networks 3 0.606 0.436 2 1 A Simple Framework for Contrastive Learning of Visual Representations ``` ### 4. Explainable lineage "You've read AlexNet — how does SimCLR trace back to it?" `shortestPath` returns an ordered list of nodes, which the example formats as a tree: ```typescript const lineage = await store.algorithms.shortestPath(simclr.id, alex.id, { edges: ["cites"], maxHops: 6, }); ``` ```text 2-hop citation lineage: A Simple Framework for Contrastive Learning of Visual Representations └─▶ Deep Residual Learning for Image Recognition └─▶ ImageNet Classification with Deep Convolutional Neural Networks ``` ### 5. Heritage check `canReach` is the boolean sibling of `shortestPath` — useful when you don't need the path, just the answer. Here: "which of these modern papers still trace back to Rumelhart's 1986 backprop paper?" ```typescript const reaches = await store.algorithms.canReach(paper.id, backprop.id, { edges: ["cites"], maxHops: 10, }); ``` ```text ✓ "LLaMA: Open and Efficient Foundation Language Models" traces to Rumelhart 1986 ✓ "Learning Transferable Visual Models From Natural Language" traces to Rumelhart 1986 ✓ "Chain-of-Thought Prompting Elicits Reasoning in Large LMs" traces to Rumelhart 1986 ✓ "A Simple Framework for Contrastive Learning of Visual Reps" traces to Rumelhart 1986 ``` ### 6. Collaborator discovery `neighbors` returns the direct neighborhood of a node. Compose it — authors of CLIP → their other papers → co-authors on those papers — to rank natural collaborators by shared-paper count: ```typescript const clipAuthors = await store.algorithms.neighbors(clip.id, { edges: ["authored_by"], depth: 1, }); // For each CLIP author: walk authored_by backwards to all their papers, // then forwards to all their co-authors. const perAuthorPapers = await Promise.all( clipAuthors.map((author) => store.algorithms.neighbors(author.id, { edges: ["authored_by"], direction: "in", depth: 1, }), ), ); ``` Issuing each level in parallel keeps the fan-out at `O(depth)` round-trips instead of `O(authors × papers)`. ```text Seed paper authors: Ilya Sutskever, Jong Wook Kim, Aditya Ramesh, Alec Radford, Chris Hallacy Nearby collaborators beyond the original CLIP paper: 2× shared papers with CLIP authors Alex Krizhevsky 2× shared papers with CLIP authors Geoffrey Hinton 2× shared papers with CLIP authors Rewon Child 2× shared papers with CLIP authors Jeffrey Wu 1× shared papers with CLIP authors Nitish Srivastava ``` ## Run It The full source lives at [`packages/typegraph/examples/14-research-copilot.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/14-research-copilot.ts). From a checkout of the repository: ```bash pnpm install npx tsx packages/typegraph/examples/14-research-copilot.ts ``` The example builds the graph, runs every scene, and tears down — all against an in-memory SQLite database. To persist it, point `createExampleBackend()` at a file path. To run it on Postgres, swap the import to `createPostgresBackend` — see [Backend Setup](/backend-setup). ## Next Steps - [Graph Algorithms](/graph-algorithms) — the full API for `shortestPath`, `reachable`, `canReach`, `neighbors`, and `degree` - [Knowledge Graph for RAG](/examples/knowledge-graph-rag) — entity linking, chunk traversal, and hybrid vector + fulltext retrieval - [Ontology & Reasoning](/ontology) — inverse edges, subclass hierarchies, and other ontology primitives beyond `broader_than` - [Semantic Search](/semantic-search) — production vector search with pgvector, sqlite-vec, and native libSQL / Turso vectors # Workflow Engine > State machines with approvals, assignments, and escalations This example builds a workflow engine with: - **State machine definitions** as graph schemas - **Approval chains** with multiple approvers - **Task assignment** and delegation - **Escalation rules** based on time - **Audit trail** of all state changes ## Schema Definition ```typescript import { z } from "zod"; import { defineNode, defineEdge, defineGraph, implies } from "@nicia-ai/typegraph"; // Workflow definition (template) const WorkflowDefinition = defineNode("WorkflowDefinition", { schema: z.object({ name: z.string(), description: z.string().optional(), version: z.number().int().positive(), isActive: z.boolean().default(true), }), }); // States within a workflow const State = defineNode("State", { schema: z.object({ name: z.string(), type: z.enum(["initial", "intermediate", "terminal", "approval"]), config: z.record(z.unknown()).optional(), // State-specific config }), }); // Transitions between states const Transition = defineNode("Transition", { schema: z.object({ name: z.string(), condition: z.string().optional(), // Expression to evaluate requiredRole: z.string().optional(), }), }); // Workflow instances const WorkflowInstance = defineNode("WorkflowInstance", { schema: z.object({ referenceId: z.string(), // ID of the entity being processed referenceType: z.string(), // Type of entity (e.g., "PurchaseOrder") status: z.enum(["active", "completed", "cancelled", "failed"]).default("active"), data: z.record(z.unknown()).optional(), // Instance-specific data createdAt: z.string().datetime(), completedAt: z.string().datetime().optional(), }), }); // Tasks assigned to users const Task = defineNode("Task", { schema: z.object({ title: z.string(), description: z.string().optional(), type: z.enum(["action", "approval", "review", "notification"]), status: z.enum(["pending", "in_progress", "completed", "rejected", "escalated"]).default("pending"), dueDate: z.string().datetime().optional(), priority: z.enum(["low", "medium", "high", "urgent"]).default("medium"), result: z.record(z.unknown()).optional(), completedAt: z.string().datetime().optional(), }), }); // Users const User = defineNode("User", { schema: z.object({ email: z.string().email(), name: z.string(), role: z.string(), department: z.string().optional(), }), }); // Comments on tasks const Comment = defineNode("Comment", { schema: z.object({ content: z.string(), createdAt: z.string().datetime(), }), }); // Edges const hasState = defineEdge("hasState"); const hasTransition = defineEdge("hasTransition"); const fromState = defineEdge("fromState"); const toState = defineEdge("toState"); const usesDefinition = defineEdge("usesDefinition"); const currentState = defineEdge("currentState"); const hasTask = defineEdge("hasTask"); const assignedTo = defineEdge("assignedTo"); const createdBy = defineEdge("createdBy"); const hasComment = defineEdge("hasComment"); const reportsTo = defineEdge("reportsTo"); // For escalation chain // Graph const graph = defineGraph({ id: "workflow_engine", nodes: { WorkflowDefinition: { type: WorkflowDefinition, unique: [ { name: "workflow_name_version", fields: ["name", "version"], scope: "kind", collation: "binary", }, ], }, State: { type: State }, Transition: { type: Transition }, WorkflowInstance: { type: WorkflowInstance }, Task: { type: Task }, User: { type: User, unique: [ { name: "user_email", fields: ["email"], scope: "kind", collation: "caseInsensitive", }, ], }, Comment: { type: Comment }, }, edges: { hasState: { type: hasState, from: [WorkflowDefinition], to: [State] }, hasTransition: { type: hasTransition, from: [WorkflowDefinition], to: [Transition] }, fromState: { type: fromState, from: [Transition], to: [State] }, toState: { type: toState, from: [Transition], to: [State] }, usesDefinition: { type: usesDefinition, from: [WorkflowInstance], to: [WorkflowDefinition] }, currentState: { type: currentState, from: [WorkflowInstance], to: [State] }, hasTask: { type: hasTask, from: [WorkflowInstance], to: [Task] }, assignedTo: { type: assignedTo, from: [Task], to: [User] }, createdBy: { type: createdBy, from: [Task, Comment, WorkflowInstance], to: [User] }, hasComment: { type: hasComment, from: [Task], to: [Comment] }, reportsTo: { type: reportsTo, from: [User], to: [User] }, }, ontology: [ // Escalation implies assignment implies(reportsTo, assignedTo), ], }); ``` ## Workflow Definition ### Create Approval Workflow ```typescript async function createApprovalWorkflow(): Promise> { return store.transaction(async (tx) => { // Create workflow definition const workflow = await tx.nodes.WorkflowDefinition.create({ name: "Purchase Order Approval", description: "Multi-level approval for purchase orders", version: 1, isActive: true, }); // Create states const states = { draft: await tx.nodes.State.create({ name: "Draft", type: "initial", }), pendingManagerApproval: await tx.nodes.State.create({ name: "Pending Manager Approval", type: "approval", config: { approverRole: "manager", timeout: "48h" }, }), pendingFinanceApproval: await tx.nodes.State.create({ name: "Pending Finance Approval", type: "approval", config: { approverRole: "finance", timeout: "24h" }, }), approved: await tx.nodes.State.create({ name: "Approved", type: "terminal", }), rejected: await tx.nodes.State.create({ name: "Rejected", type: "terminal", }), }; // Link states to workflow for (const state of Object.values(states)) { await tx.edges.hasState.create(workflow, state, {}); } // Create transitions const transitions = [ { from: states.draft, to: states.pendingManagerApproval, name: "Submit", requiredRole: "requester", }, { from: states.pendingManagerApproval, to: states.pendingFinanceApproval, name: "Approve", requiredRole: "manager", condition: "amount > 1000", }, { from: states.pendingManagerApproval, to: states.approved, name: "Approve", requiredRole: "manager", condition: "amount <= 1000", }, { from: states.pendingManagerApproval, to: states.rejected, name: "Reject", requiredRole: "manager", }, { from: states.pendingFinanceApproval, to: states.approved, name: "Approve", requiredRole: "finance", }, { from: states.pendingFinanceApproval, to: states.rejected, name: "Reject", requiredRole: "finance", }, ]; for (const t of transitions) { const transition = await tx.nodes.Transition.create({ name: t.name, requiredRole: t.requiredRole, condition: t.condition, }); await tx.edges.hasTransition.create(workflow, transition, {}); await tx.edges.fromState.create(transition, t.from, {}); await tx.edges.toState.create(transition, t.to, {}); } return workflow; }); } ``` ## Workflow Instances ### Start Workflow ```typescript interface StartWorkflowInput { workflowName: string; referenceId: string; referenceType: string; data?: Record; createdByUserId: string; } async function startWorkflow(input: StartWorkflowInput): Promise> { return store.transaction(async (tx) => { // Find workflow definition const workflow = await tx .query() .from("WorkflowDefinition", "w") .whereNode("w", (w) => w.name.eq(input.workflowName).and(w.isActive.eq(true))) .select((ctx) => ctx.w) .first(); if (!workflow) { throw new Error(`Workflow '${input.workflowName}' not found`); } // Find initial state const initialState = await tx .query() .from("WorkflowDefinition", "w") .whereNode("w", (w) => w.id.eq(workflow.id)) .traverse("hasState", "e") .to("State", "s") .whereNode("s", (s) => s.type.eq("initial")) .select((ctx) => ctx.s) .first(); if (!initialState) { throw new Error("Workflow has no initial state"); } // Create instance const instance = await tx.nodes.WorkflowInstance.create({ referenceId: input.referenceId, referenceType: input.referenceType, status: "active", data: input.data, createdAt: new Date().toISOString(), }); // Link to definition and state await tx.edges.usesDefinition.create(instance, workflow, {}); await tx.edges.currentState.create(instance, initialState, {}); // Link to creator const creator = await tx.nodes.User.getById(input.createdByUserId); if (!creator) throw new Error(`User not found: ${input.createdByUserId}`); await tx.edges.createdBy.create(instance, creator, {}); return instance; }); } ``` ### Get Available Transitions ```typescript interface AvailableTransition { id: string; name: string; targetState: string; requiredRole?: string; condition?: string; } async function getAvailableTransitions( instanceId: string, userId: string ): Promise { // Get user's role const user = await store.nodes.User.getById(userId); if (!user) throw new Error(`User not found: ${userId}`); const userRole = user.role; // Get current state const currentState = await store .query() .from("WorkflowInstance", "i") .whereNode("i", (i) => i.id.eq(instanceId)) .traverse("currentState", "e") .to("State", "s") .select((ctx) => ctx.s) .first(); if (!currentState) { throw new Error("Instance has no current state"); } // Get transitions from current state const transitions = await store .query() .from("State", "s") .whereNode("s", (s) => s.id.eq(currentState.id)) .traverse("fromState", "e1", { direction: "in" }) .to("Transition", "t") .traverse("toState", "e2") .to("State", "target") .select((ctx) => ({ id: ctx.t.id, name: ctx.t.name, targetState: ctx.target.name, requiredRole: ctx.t.requiredRole, condition: ctx.t.condition, })) .execute(); // Filter by role return transitions.filter( (t) => !t.requiredRole || t.requiredRole === userRole || userRole === "admin" ); } ``` ### Execute Transition ```typescript async function executeTransition( instanceId: string, transitionId: string, userId: string, result?: Record ): Promise { await store.transaction(async (tx) => { const instance = await tx.nodes.WorkflowInstance.getById(instanceId); if (!instance) throw new Error(`WorkflowInstance not found: ${instanceId}`); if (instance.status !== "active") { throw new Error("Workflow is not active"); } // Verify transition is valid const available = await getAvailableTransitions(instanceId, userId); const transition = available.find((t) => t.id === transitionId); if (!transition) { throw new Error("Transition not available"); } // Get target state const targetState = await tx .query() .from("Transition", "t") .whereNode("t", (t) => t.id.eq(transitionId)) .traverse("toState", "e") .to("State", "s") .select((ctx) => ctx.s) .first(); // Remove current state edge const currentStateEdge = await tx .query() .from("WorkflowInstance", "i") .whereNode("i", (i) => i.id.eq(instanceId)) .traverse("currentState", "e") .to("State", "s") .select((ctx) => ctx.e.id) .first(); if (currentStateEdge) { await tx.edges.currentState.delete(currentStateEdge); } // Add new state edge await tx.edges.currentState.create(instance, targetState!, {}); // Update instance data const updatedData = { ...instance.data, lastTransition: transition.name, ...result }; const updates: Partial = { data: updatedData }; // Check if terminal state if (targetState!.type === "terminal") { updates.status = "completed"; updates.completedAt = new Date().toISOString(); } await tx.nodes.WorkflowInstance.update(instanceId, updates); // Complete any pending tasks const pendingTasks = await tx .query() .from("WorkflowInstance", "i") .whereNode("i", (i) => i.id.eq(instanceId)) .traverse("hasTask", "e") .to("Task", "t") .whereNode("t", (t) => t.status.in(["pending", "in_progress"])) .select((ctx) => ctx.t.id) .execute(); for (const taskId of pendingTasks) { await tx.nodes.Task.update(taskId, { status: "completed", completedAt: new Date().toISOString(), }); } // Create tasks for new state if needed if (targetState!.type === "approval") { await createApprovalTask(tx, instanceId, targetState!, userId); } }); } ``` ## Task Management ### Create Approval Task ```typescript async function createApprovalTask( tx: Transaction, instanceId: string, state: Node, requesterId: string ): Promise { const config = state.config as { approverRole: string; timeout: string } | undefined; if (!config) return; // Find approver (first user with matching role, or requester's manager) let approver = await tx .query() .from("User", "u") .whereNode("u", (u) => u.role.eq(config.approverRole)) .select((ctx) => ctx.u) .first(); // If no direct match, walk the reporting chain upward until we hit someone // with the approver role. The recursive traversal tags each hop with its // depth so we can pick the nearest matching manager. if (!approver) { const candidates = await tx .query() .from("User", "requester") .whereNode("requester", (u) => u.id.eq(requesterId)) .traverse("reportsTo", "e") .recursive({ depth: "depth" }) .to("User", "manager") .whereNode("manager", (u) => u.role.eq(config.approverRole)) .orderBy("depth", "asc") .select((ctx) => ctx.manager) .execute(); approver = candidates[0]; } if (!approver) { throw new Error(`No approver found with role '${config.approverRole}'`); } // Calculate due date const dueDate = calculateDueDate(config.timeout); // Create task const task = await tx.nodes.Task.create({ title: `Approval Required: ${state.name}`, description: `Please review and approve or reject.`, type: "approval", status: "pending", priority: "medium", dueDate: dueDate.toISOString(), }); // Link task to instance and approver const instance = await tx.nodes.WorkflowInstance.getById(instanceId); if (!instance) throw new Error(`WorkflowInstance not found: ${instanceId}`); await tx.edges.hasTask.create(instance, task, {}); await tx.edges.assignedTo.create(task, approver, {}); } function calculateDueDate(timeout: string): Date { const now = new Date(); const match = timeout.match(/^(\d+)(h|d)$/); if (!match) return new Date(now.getTime() + 24 * 60 * 60 * 1000); // Default 24h const value = parseInt(match[1], 10); const unit = match[2]; if (unit === "h") { return new Date(now.getTime() + value * 60 * 60 * 1000); } else { return new Date(now.getTime() + value * 24 * 60 * 60 * 1000); } } ``` ### Get User's Tasks ```typescript interface TaskWithContext { id: string; title: string; type: string; status: string; priority: string; dueDate?: string; workflowName: string; referenceId: string; referenceType: string; } async function getUserTasks( userId: string, status?: "pending" | "in_progress" ): Promise { let query = store .query() .from("User", "u") .whereNode("u", (u) => u.id.eq(userId)) .traverse("assignedTo", "e", { direction: "in" }) .to("Task", "t"); if (status) { query = query.whereNode("t", (t) => t.status.eq(status)); } else { query = query.whereNode("t", (t) => t.status.in(["pending", "in_progress"])); } return query .traverse("hasTask", "e2", { direction: "in" }) .to("WorkflowInstance", "i") .traverse("usesDefinition", "e3") .to("WorkflowDefinition", "w") .select((ctx) => ({ id: ctx.t.id, title: ctx.t.title, type: ctx.t.type, status: ctx.t.status, priority: ctx.t.priority, dueDate: ctx.t.dueDate, workflowName: ctx.w.name, referenceId: ctx.i.referenceId, referenceType: ctx.i.referenceType, })) .orderBy((ctx) => ctx.t.dueDate, "asc") .execute(); } ``` ### Complete Task ```typescript async function completeTask( taskId: string, userId: string, decision: "approve" | "reject", comment?: string ): Promise { await store.transaction(async (tx) => { const task = await tx.nodes.Task.getById(taskId); if (!task) throw new Error(`Task not found: ${taskId}`); // Verify user is assigned const assignee = await tx .query() .from("Task", "t") .whereNode("t", (t) => t.id.eq(taskId)) .traverse("assignedTo", "e") .to("User", "u") .select((ctx) => ctx.u.id) .first(); if (assignee !== userId) { throw new Error("User is not assigned to this task"); } // Update task await tx.nodes.Task.update(taskId, { status: decision === "approve" ? "completed" : "rejected", completedAt: new Date().toISOString(), result: { decision }, }); // Add comment if provided if (comment) { const commentNode = await tx.nodes.Comment.create({ content: comment, createdAt: new Date().toISOString(), }); await tx.edges.hasComment.create(task, commentNode, {}); const user = await tx.nodes.User.getById(userId); if (!user) throw new Error(`User not found: ${userId}`); await tx.edges.createdBy.create(commentNode, user, {}); } // Get workflow instance const instance = await tx .query() .from("Task", "t") .whereNode("t", (t) => t.id.eq(taskId)) .traverse("hasTask", "e", { direction: "in" }) .to("WorkflowInstance", "i") .select((ctx) => ctx.i) .first(); // Find and execute the appropriate transition const transitions = await getAvailableTransitions(instance!.id, userId); const transition = transitions.find((t) => decision === "approve" ? t.name === "Approve" : t.name === "Reject" ); if (transition) { await executeTransition(instance!.id, transition.id, userId, { decision }); } }); } ``` ## Escalation ### Check Overdue Tasks ```typescript async function getOverdueTasks(): Promise> { const now = new Date().toISOString(); return store .query() .from("Task", "t") .whereNode("t", (t) => t.status .in(["pending", "in_progress"]) .and(t.dueDate.isNotNull()) .and(t.dueDate.lt(now)) ) .traverse("assignedTo", "e") .to("User", "u") .select((ctx) => ({ task: ctx.t, assignee: ctx.u, })) .execute(); } ``` ### Escalate Task ```typescript async function escalateTask(taskId: string): Promise { await store.transaction(async (tx) => { // Get current assignee const currentAssignment = await tx .query() .from("Task", "t") .whereNode("t", (t) => t.id.eq(taskId)) .traverse("assignedTo", "e") .to("User", "u") .select((ctx) => ({ edgeId: ctx.e.id, user: ctx.u })) .first(); if (!currentAssignment) { throw new Error("Task has no assignee"); } // Single-hop traversal to the direct manager const manager = await tx .query() .from("User", "u") .whereNode("u", (u) => u.id.eq(currentAssignment.user.id)) .traverse("reportsTo", "e") .to("User", "manager") .select((ctx) => ctx.manager) .first(); if (!manager) { throw new Error("No manager found for escalation"); } // Update task await tx.nodes.Task.update(taskId, { status: "escalated", priority: "urgent", }); // Reassign to manager await tx.edges.assignedTo.delete(currentAssignment.edgeId); const task = await tx.nodes.Task.getById(taskId); if (!task) throw new Error(`Task not found: ${taskId}`); await tx.edges.assignedTo.create(task, manager, {}); // Add escalation comment const comment = await tx.nodes.Comment.create({ content: `Task escalated from ${currentAssignment.user.name} due to timeout`, createdAt: new Date().toISOString(), }); await tx.edges.hasComment.create(task, comment, {}); }); } ``` ### Run Escalation Job ```typescript async function runEscalationJob(): Promise<{ escalated: number }> { const overdueTasks = await getOverdueTasks(); let escalated = 0; for (const { task } of overdueTasks) { try { await escalateTask(task.id); escalated++; } catch (error) { console.error(`Failed to escalate task ${task.id}:`, error); } } return { escalated }; } ``` ## Workflow History ### Get Instance Timeline ```typescript interface TimelineEvent { timestamp: string; type: "state_change" | "task_created" | "task_completed" | "comment"; description: string; actor?: string; } async function getInstanceTimeline(instanceId: string): Promise { const events: TimelineEvent[] = []; // Run both history reads in sequence. Two statements, not one, and // not a snapshot — under read-committed isolation the task view can observe // a write the state-change view did not. const [stateHistory, tasks] = await store.batch( store .query() .from("WorkflowInstance", "i") .temporal("includeEnded") .whereNode("i", (i) => i.id.eq(instanceId)) .traverse("currentState", "e") .to("State", "s") .orderBy("e", "validFrom", "asc") .select((ctx) => ({ stateName: ctx.s.name, timestamp: ctx.e.validFrom, })), store .query() .from("WorkflowInstance", "i") .whereNode("i", (i) => i.id.eq(instanceId)) .traverse("hasTask", "e") .to("Task", "t") .optionalTraverse("assignedTo", "a") .to("User", "u") .select((ctx) => ({ title: ctx.t.title, status: ctx.t.status, createdAt: ctx.t.createdAt, completedAt: ctx.t.completedAt, assignee: ctx.u?.name, })), ); for (const state of stateHistory) { events.push({ timestamp: state.timestamp, type: "state_change", description: `Entered state: ${state.stateName}`, }); } for (const task of tasks) { events.push({ timestamp: task.createdAt.toISOString(), type: "task_created", description: `Task created: ${task.title}`, actor: task.assignee, }); if (task.completedAt) { events.push({ timestamp: task.completedAt, type: "task_completed", description: `Task ${task.status}: ${task.title}`, actor: task.assignee, }); } } // Sort by timestamp return events.sort((a, b) => a.timestamp.localeCompare(b.timestamp)); } ``` ## Next Steps - [Document Management](/examples/document-management) - CMS with semantic search - [Product Catalog](/examples/product-catalog) - Categories, variants, inventory - [Audit Trail](/examples/audit-trail) - Complete change tracking