This is the abridged developer documentation for TypeGraph
# What is TypeGraph?
> A TypeScript-first embedded knowledge graph library
TypeGraph is a **TypeScript-first, embedded knowledge graph library** that brings property graph semantics and
ontological reasoning to applications using standard relational databases. Rather than introducing a separate graph
database, TypeGraph lives inside your application as a library, storing graph data in your existing SQLite or
PostgreSQL database.
## Architecture

## Core Capabilities
### 1. Type-Driven Schema Definition
Zod schemas are the single source of truth. From one schema definition, TypeGraph derives:
- Runtime validation rules
- TypeScript types (inferred, not duplicated)
- Database storage requirements
- Query builder type constraints
```typescript
const Person = defineNode("Person", {
schema: z.object({
fullName: z.string().min(1),
email: z.string().email().optional(),
dateOfBirth: z.date().optional(),
}),
});
```
### 2. Semantic Layer with Ontological Reasoning
Type-level relationships enable sophisticated inference:
| Relationship | Meaning | Use Case |
| -------------- | ----------------------------------------- | ---------------------- |
| `subClassOf` | Instance inheritance (Podcast IS-A Media) | Query expansion |
| `broader` | Hierarchical concept (ML broader than DL) | Topic navigation |
| `equivalentTo` | Same concept, different name | Cross-system mapping |
| `disjointWith` | Cannot be both (Person ≠ Organization) | Constraint validation |
| `implies` | Edge entailment (marriedTo implies knows) | Relationship inference |
| `inverseOf` | Edge pairs (manages/managedBy) | Bidirectional queries |
### 3. Self-Describing Schema (Homoiconic)
The schema and ontology are stored in the database as data, enabling:
- Runtime schema introspection
- Versioned schema history
- Self-describing exports and backups
- Migration tooling
### 4. Type-Safe Query Compilation
Queries compile to an AST before targeting SQL:
- Consistent semantics across SQLite and PostgreSQL
- Type-checked at compile time
- Query results have inferred types
### 5. Temporal and Bitemporal History
Every node and edge has a valid-time window (`validFrom` / `validTo`), so you
can ask what was true at a domain instant with `.temporal("asOf", T)` or
`store.asOf(T)`. Stores created with `{ history: true }` also capture recorded
time for TypeGraph-managed writes, the system-time axis that remembers when the
graph wrote each fact down. `store.asOfRecorded(T)` reconstructs what the graph
captured at a recorded instant, and
`store.asOf(validT).asOfRecorded(recordedT)` pins both axes independently.
Use it for audit trails, agent decision replay, effective-dated policies, and
breach forensics. See [Temporal queries](/queries/temporal) and the
[Bitemporal Time Travel](/examples/bitemporal-time-travel) example.
## Design Philosophy
### Embedded, Not External
TypeGraph is a library dependency, not a networked service. TypeGraph initializes with your application, uses your
database connection, and requires no separate deployment.
### Schema-First, Type-Driven
Define your schemas once with Zod, and TypeGraph handles validation, type inference, and storage.
No duplicate type definitions or manual synchronization.
### Explicit Over Implicit
TypeGraph favors explicit declarations:
- Relationships are declared, not inferred from foreign keys
- Semantic relationships are explicit in the ontology
- Cascade behavior is configured, not assumed
### Portable Abstractions
The query builder generates portable ASTs that can target different SQL dialects.
The same query code works with SQLite and PostgreSQL.
## What TypeGraph Is Not
TypeGraph deliberately excludes:
- **Broad graph analytics suites**: Focused PageRank, connectivity, and
deterministic label-propagation primitives are built in; modularity
optimization and most centrality measures are not
- **Distributed storage**: Single-database deployment only
These exclusions keep TypeGraph focused and maintainable.
Note: TypeGraph **does support** semantic search via native database vector
engines: pgvector for PostgreSQL, sqlite-vec for the local (better-sqlite3)
SQLite backend, and libSQL's built-in vectors for the libSQL / Turso backend.
See [Semantic Search](/semantic-search) for details.
Note: TypeGraph **does support** fulltext search — native BM25 on SQLite (FTS5)
and `tsvector` + GIN on PostgreSQL, with a query-builder
`n.$fulltext.matches()` predicate that composes with any other predicate.
Combine with semantic search for hybrid RAG retrieval. See
[Fulltext Search](/fulltext-search) for details.
Note: TypeGraph does support **variable-length paths** via `.recursive()` with
configurable depth limits, optional path/depth projection, and explicit cycle
policy. Cycle prevention is the default.
See [Recursive Traversals](/queries/recursive) for details.
Note: TypeGraph ships **Tier 1 graph algorithms** (shortest path, reachability,
neighborhoods, and degree) on `store.algorithms.*`. Traversal calls use a
set-based BFS frontier, while degree uses a single count query. See
[Graph Algorithms](/graph-algorithms) for details.
Note: TypeGraph supports **runtime schema induction** via graph
extensions. An LLM or ingestion agent can propose a typed schema as a
JSON-serializable document, an operator approves it, and `store.evolve()`
atomically commits a new schema version — no redeploy, full Zod
validation, restart parity. See [Graph Extensions](/graph-extensions)
for the agent-driven workflow.
Note: TypeGraph ships **graph merge** — fork a store into isolated working
copies, let many writers (parallel agents, importers, reviewers) edit
independently, then reconcile them into one canonical graph with deterministic
entity resolution (exact / blocking / fulltext / vector / hybrid), edge
repointing, conflict reporting, and provenance. `mergeIncremental()` folds new
sources into a *live* graph without creating duplicates — the primitive for
multi-agent knowledge-graph construction and continuous ingestion. See
[Graph Merge](/graph-merge) for the full guide.
Note: TypeGraph supports **bitemporal graph reads**. Valid time answers "when
was this fact true in the domain?" Recorded time answers "when did the graph
record it?" Together they reconstruct prior captured state after corrections, replay
agent decisions against the graph they actually saw, and traverse access graphs
at a breach instant for TypeGraph-managed writes. See
[Temporal queries](/queries/temporal) and the
[Agent Decision Replay](/examples/agent-decision-replay) example.
## Why TypeGraph?
### Compared to Graph Databases (Neo4j, Amazon Neptune)
Graph databases are powerful but come with operational overhead:
| Aspect | Graph Database | TypeGraph |
|--------|---------------|-----------|
| **Deployment** | Separate service to manage, scale, and monitor | Library in your app, uses existing database |
| **Network** | Additional latency for every query | In-process, no network hop |
| **Transactions** | Separate transaction scope from your SQL data | Same ACID transaction as your other data |
| **Learning curve** | New query language (Cypher, Gremlin) | TypeScript you already know |
| **Graph algorithms** | Broad suites (PageRank, shortest path, community detection) | Focused algorithms (shortest path, reachability, neighborhoods, degree, WCC, label propagation, PageRank/PPR) |
| **Scale** | Optimized for billions of nodes | Best for thousands to millions |
**Choose TypeGraph** when your graph is part of your application domain (knowledge bases, org
charts, content relationships) rather than a standalone analytical system.
### Compared to ORMs (Prisma, Drizzle, TypeORM)
ORMs model relations through foreign keys, which works well for simple associations but lacks graph semantics:
| Aspect | Traditional ORM | TypeGraph |
|--------|----------------|-----------|
| **Relationships** | Foreign keys, eager/lazy loading | First-class edges with properties |
| **Traversals** | Manual joins or N+1 queries | Fluent traversal API, compiled to efficient SQL |
| **Inheritance** | Table-per-class or single-table | Semantic `subClassOf` with query expansion |
| **Constraints** | Foreign key constraints | Disjointness, cardinality, implications |
| **Schema** | Migrations alter tables | Schema versioning, JSON properties |
**Choose TypeGraph** when you need to traverse relationships, model type hierarchies, or enforce
semantic constraints beyond what foreign keys provide.
### Compared to Triple Stores (RDF, SPARQL)
Triple stores and RDF provide rich ontological modeling but have practical challenges:
| Aspect | Triple Store | TypeGraph |
|--------|-------------|-----------|
| **Type safety** | Runtime validation, stringly-typed | Full TypeScript inference |
| **Query language** | SPARQL (powerful but verbose) | TypeScript fluent API |
| **Schema** | OWL/RDFS (complex specification) | Zod schemas (familiar, composable) |
| **Integration** | Separate system, data sync required | Embedded in your app |
| **Inference** | Full reasoning engines available | Precomputed closures, practical subset |
**Choose TypeGraph** when you want ontological concepts (subclass, disjoint, implies) without the
complexity of full semantic web stack.
### The TypeGraph Sweet Spot
TypeGraph is designed for applications where:
1. **The graph is your domain model** — not a separate analytical system
2. **You already use SQL** — and don't want another database to manage
3. **Type safety matters** — you want compile-time checking, not runtime surprises
4. **Semantic relationships help** — inheritance, implications, constraints add value
5. **Scale is moderate** — thousands to millions of nodes, not billions
## When to Use TypeGraph
TypeGraph is ideal for:
- **Knowledge bases** with typed entities and relationships
- **Organizational structures** with hierarchies and roles
- **Content graphs** with topics, articles, and references
- **Domain models** requiring semantic constraints
- **RAG applications** combining graph traversal with vector search
- **Multi-source ingestion & entity resolution** — reconcile parallel agent or
importer outputs into one canonical graph with [graph merge](/graph-merge)
- **Auditable AI systems and forensics** — reconstruct the graph an agent or
investigator saw at a recorded instant with [bitemporal reads](/queries/temporal#recorded-time-bitemporal)
TypeGraph is not ideal for:
- Large-scale graph analytics requiring distributed processing
- Social networks with billions of edges
- Real-time streaming graph data
- Applications requiring a broad graph-data-science suite such as community
detection or betweenness centrality (use Neo4j or a graph library; focused
algorithms—including PageRank and weighted shortest path—ship on
`store.algorithms.*`)
# Quick Start
> Set up TypeGraph and build your first knowledge graph
Get TypeGraph running in your project with this minimal example.
## 1. Install
```bash
npm install @nicia-ai/typegraph zod drizzle-orm better-sqlite3
npm install -D @types/better-sqlite3
```
> **Edge environments or libsql:** Skip `better-sqlite3` and use
> `@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql` with `@libsql/client`, or
> `@nicia-ai/typegraph/adapters/drizzle/sqlite` with your edge-compatible driver (D1, bun:sqlite).
> See [Backend Setup](/backend-setup#libsql--turso) and [Edge and Serverless](/integration#edge-and-serverless).
## 2. Create Your First Graph
```typescript
import { z } from "zod";
import { defineNode, defineEdge, defineGraph } from "@nicia-ai/typegraph";
import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local";
// Define your schema
const Person = defineNode("Person", {
schema: z.object({ name: z.string(), role: z.string().optional() }),
});
const Project = defineNode("Project", {
schema: z.object({ name: z.string(), status: z.enum(["active", "done"]) }),
});
const worksOn = defineEdge("worksOn");
const graph = defineGraph({
id: "my_app",
nodes: { Person: { type: Person }, Project: { type: Project } },
edges: { worksOn: { type: worksOn, from: [Person], to: [Project] } },
});
// Provision an in-memory database and create the store
const store = await createLocalSqliteStore(graph);
// Use it!
const alice = await store.nodes.Person.create({ name: "Alice", role: "Engineer" });
const project = await store.nodes.Project.create({ name: "Website", status: "active" });
await store.edges.worksOn.create(alice, project, {});
// Query with full type safety
const results = await store
.query()
.from("Person", "p")
.traverse("worksOn", "e")
.to("Project", "proj")
.select((ctx) => ({ person: ctx.p.name, project: ctx.proj.name }))
.execute();
console.log(results); // [{ person: "Alice", project: "Website" }]
```
That's it! You have a working knowledge graph. Read on for the complete setup guide.
This managed entrypoint returns the complete typed `Store` while keeping its
public declaration surface independent of Drizzle. Use
`@nicia-ai/typegraph/postgres/pglite` for the same setup with in-process
PostgreSQL. If your application owns the database connection or needs direct
driver access, use the adapter entrypoints described below instead.
---
## Complete Setup Guide
This section covers production setup with SQLite and PostgreSQL in detail.
### Installation
```bash
npm install @nicia-ai/typegraph zod drizzle-orm better-sqlite3
npm install -D @types/better-sqlite3
```
> `better-sqlite3` is optional. For libsql/Turso, use `@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql`.
> For D1 or bun:sqlite, use `@nicia-ai/typegraph/adapters/drizzle/sqlite` with the matching Drizzle driver.
### SQLite Setup
TypeGraph provides two ways to set up SQLite:
#### Managed Store (Recommended)
Use the managed Store when TypeGraph should own the connection and provision
its schema:
```typescript
import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local";
const store = await createLocalSqliteStore(graph, { path: "./my-app.db" });
// The Store owns the connection.
await store.close();
```
The return value is a `Store`: typed node and edge collections, queries,
algorithms, graph-owned transactions, schema evolution, and schema-derived
property types are all available. Adapter-native handles and caller-owned
transaction adoption are absent by design; opt into `AdapterStore` through a
Drizzle adapter entrypoint when application tables must share a transaction.
#### Quick Setup (Recommended for Development)
Use the backend wrapper when you also need the underlying Drizzle database or
want to choose how the Store is created.
> **Note:** `createLocalSqliteBackend` requires `better-sqlite3` and only works in Node.js.
> For edge environments, see [Manual Setup](#manual-setup-full-control) with
> `/adapters/drizzle/sqlite`.
```typescript
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
// In-memory database (data lost on restart)
const { backend } = createLocalSqliteBackend();
// File-based database (persistent)
const { backend, db } = createLocalSqliteBackend({ path: "./my-app.db" });
```
The function returns both the `backend` (for use with `createStore`) and `db`
(the underlying Drizzle instance for direct SQL access if needed).
#### Manual Setup (Full Control)
For production deployments or when you need full control over the database configuration:
```typescript
import Database from "better-sqlite3";
import { drizzle } from "drizzle-orm/better-sqlite3";
import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// Create database connection
const sqlite = new Database("my-app.db");
// Run TypeGraph migrations (creates required tables)
sqlite.exec(generateSqliteMigrationSQL());
// Create Drizzle instance
const db = drizzle(sqlite);
// Create the backend
const backend = createSqliteBackend(db);
```
#### libsql / Turso Setup
For Turso, embedded replicas, or sharing a libsql connection with other libraries:
```typescript
import { createClient } from "@libsql/client";
import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql";
const client = createClient({ url: "file:app.db" });
const { backend } = await createLibsqlBackend(client);
```
`createLibsqlBackend` handles DDL automatically. The caller owns the client and
is responsible for closing it. See [Backend Setup](/backend-setup#libsql--turso) for
remote Turso URLs and caveats.
#### Edge-Compatible Setup (D1, bun:sqlite)
For Cloudflare Workers or Bun, use the driver-agnostic backend:
```typescript
import { drizzle } from "drizzle-orm/d1"; // or bun-sqlite
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// D1 example
const db = drizzle(env.DB);
const backend = createSqliteBackend(db);
```
Use [drizzle-kit managed migrations](/integration#drizzle-kit-managed-migrations-recommended)
to set up the schema.
#### Drizzle-Kit Managed Migrations
If you already use `drizzle-kit` for migrations, see [Drizzle-Kit Managed Migrations](/integration#drizzle-kit-managed-migrations-recommended)
for how to import TypeGraph's schema into your `schema.ts` file.
## Defining Your Schema
### Step 1: Define Node Types
Nodes represent entities in your graph. Each node type has a name and a Zod schema:
```typescript
import { z } from "zod";
import { defineNode } from "@nicia-ai/typegraph";
const Person = defineNode("Person", {
schema: z.object({
name: z.string().min(1),
email: z.string().email().optional(),
bio: z.string().optional(),
}),
});
const Project = defineNode("Project", {
schema: z.object({
name: z.string(),
description: z.string().optional(),
status: z.enum(["planning", "active", "completed"]),
}),
});
const Task = defineNode("Task", {
schema: z.object({
title: z.string(),
priority: z.enum(["low", "medium", "high"]),
completed: z.boolean().default(false),
}),
});
```
### Step 2: Define Edge Types
Edges represent relationships between nodes:
```typescript
import { defineEdge } from "@nicia-ai/typegraph";
const worksOn = defineEdge("worksOn", {
schema: z.object({
role: z.string().optional(),
since: z.string().optional(),
}),
});
const hasTask = defineEdge("hasTask", {
schema: z.object({}),
});
const assignedTo = defineEdge("assignedTo", {
schema: z.object({
assignedAt: z.string().optional(),
}),
});
// Unconstrained edge — connects any node to any node
const related = defineEdge("related");
```
### Step 3: Create the Graph Definition
Combine nodes, edges, and ontology into a graph:
```typescript
import { defineGraph, disjointWith } from "@nicia-ai/typegraph";
const graph = defineGraph({
id: "project_management",
nodes: {
Person: { type: Person },
Project: { type: Project },
Task: { type: Task },
},
edges: {
worksOn: { type: worksOn, from: [Person], to: [Project] },
hasTask: { type: hasTask, from: [Project], to: [Task] },
assignedTo: { type: assignedTo, from: [Task], to: [Person] },
related, // any→any
},
ontology: [
// A Person cannot be a Project or Task
disjointWith(Person, Project),
disjointWith(Person, Task),
disjointWith(Project, Task),
],
});
```
### Step 4: Create the Store
The store connects your graph definition to the database:
```typescript
import { createStore } from "@nicia-ai/typegraph";
const store = createStore(graph, backend);
```
#### Store Creation: Which Function to Use
| Function | Schema Handling | Use Case |
| -------------------------------- | ---------------------------------------------- | -------------------------------------------------- |
| `createLocalSqliteBackend` | Automatic | Quick start, development, tests (Node.js) |
| `createLibsqlBackend` | Automatic | libsql/Turso (Node.js, Workers, browser) |
| `createLocalPgliteBackend` | Automatic | In-process Postgres, embedded apps, pgvector tests |
| `createStore` + manual migration | None | When you manage migrations externally |
| `createStoreWithSchema` | Auto-creates tables, validates & auto-migrates | **Recommended for production** |
:::caution[Fulltext requires `createStoreWithSchema`]
If your graph has any `searchable()` fields, you must boot through
`createStoreWithSchema` once at startup. It durably materializes the
fulltext storage; bare `createStore()` is an attach-only path and throws
`StoreNotInitializedError` on the first fulltext operation. Graphs
without `searchable()` fields are unaffected.
:::
For production, use `createStoreWithSchema` to validate and auto-apply safe schema changes:
```typescript
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const [store, result] = await createStoreWithSchema(graph, backend);
if (result.status === "initialized") {
console.log("Schema initialized at version", result.version);
} else if (result.status === "migrated") {
console.log(`Migrated from v${result.fromVersion} to v${result.toVersion}`);
}
// Other statuses: "unchanged", "pending", "breaking"
// See Schema Migrations for full details
```
#### Graph ID
Every graph has a unique `id` that scopes its data:
```typescript
const graph = defineGraph({
id: "my_app", // Scopes all nodes/edges to this graph
// ...
});
```
**Key behaviors:**
- All nodes and edges are stored with this `graph_id` in the database
- Multiple graphs can share the same database tables (isolated by `graph_id`)
- Changing the ID creates a new, empty graph (existing data is orphaned)
See [Multiple Graphs](/multiple-graphs) for multi-graph deployments.
## Working with Data
### Creating Nodes
```typescript
const alice = await store.nodes.Person.create({
name: "Alice Smith",
email: "alice@example.com",
});
const project = await store.nodes.Project.create({
name: "Website Redesign",
status: "active",
});
const task = await store.nodes.Task.create({
title: "Design mockups",
priority: "high",
});
```
### Creating Edges
Pass node objects directly to create edges:
```typescript
await store.edges.worksOn.create(alice, project, { role: "Lead Designer" });
await store.edges.hasTask.create(project, task, {});
await store.edges.assignedTo.create(task, alice, { assignedAt: new Date().toISOString() });
```
### Retrieving Nodes
```typescript
const person = await store.nodes.Person.getById(alice.id);
console.log(person?.name); // "Alice Smith"
```
### Updating Nodes
```typescript
const updated = await store.nodes.Task.update(task.id, { completed: true });
```
### Deleting Nodes
```typescript
await store.nodes.Task.delete(task.id);
```
## Querying Data
TypeGraph provides a fluent query builder:
```typescript
// Find all active projects
const activeProjects = await store
.query()
.from("Project", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ctx.p)
.execute();
// Find people working on a project
const teamMembers = await store
.query()
.from("Project", "p")
.traverse("worksOn", "e", { direction: "in" })
.to("Person", "person")
.select((ctx) => ({
project: ctx.p.name,
person: ctx.person.name,
}))
.execute();
// Multi-hop traversal: find tasks for a person
const myTasks = await store
.query()
.from("Person", "person")
.whereNode("person", (p) => p.name.eq("Alice Smith"))
.traverse("worksOn", "e1")
.to("Project", "project")
.traverse("hasTask", "e2")
.to("Task", "task")
.select((ctx) => ({
project: ctx.project.name,
task: ctx.task.title,
priority: ctx.task.priority,
}))
.execute();
```
`whereNode()` and `whereEdge()` constrain graph matches while traversal is
built. Use `.where((ctx) => ...)` when a condition should filter completed
rows, including optional or recursive results. Each successful traversal
combination is one row, so fanout can repeat a source entity; project an
identity and call relation `.distinct()` when the intended result is one row
per entity.
Traversal continues from the latest target by default. Reusable branching
fragments should state their source explicitly with `{ from: "alias" }` so
their behavior does not depend on which traversal preceded them. Direction,
ontology expansion, and temporal coordinates retain their ordinary query
defaults.
## Transactions
Group operations in transactions for atomicity:
```typescript
await store.transaction(async (tx) => {
const project = await tx.nodes.Project.create({
name: "New Feature",
status: "planning",
});
const task1 = await tx.nodes.Task.create({
title: "Research",
priority: "high",
});
const task2 = await tx.nodes.Task.create({
title: "Implementation",
priority: "medium",
});
await tx.edges.hasTask.create(project, task1, {});
await tx.edges.hasTask.create(project, task2, {});
});
```
## Error Handling
TypeGraph provides specific error types:
```typescript
import { ValidationError, NodeNotFoundError, DisjointError, RestrictedDeleteError } from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create({ name: "" }); // Invalid: empty name
} catch (error) {
if (error instanceof ValidationError) {
console.log("Validation failed:", error.message);
}
}
try {
await store.nodes.Project.delete(project.id);
} catch (error) {
if (error instanceof RestrictedDeleteError) {
console.log("Cannot delete: edges exist");
}
}
```
## PostgreSQL Setup
TypeGraph also supports PostgreSQL for production deployments with better concurrency and JSON support.
For in-process Postgres during local development or tests, see
[PGlite in Backend Setup](/backend-setup#pglite-postgres-in-wasm).
### Installation
```bash
npm install @nicia-ai/typegraph zod drizzle-orm pg
npm install -D @types/pg
```
### Database Setup
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Create connection pool
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Connection pool size
});
// Run TypeGraph migrations
await pool.query(generatePostgresMigrationSQL());
// Create Drizzle instance and backend
const db = drizzle(pool);
const backend = createPostgresBackend(db);
```
If you use `drizzle-kit` for migrations, see [Drizzle-Kit Managed Migrations](/integration#drizzle-kit-managed-migrations-recommended).
### PostgreSQL Advantages
- **JSONB**: Native JSON type with efficient indexing
- **Connection pooling**: Better concurrency handling
- **Partial indexes**: More efficient uniqueness constraints
- **Full transactions**: ACID guarantees across operations
### Using with Connection Pools
For production, always use connection pooling:
```typescript
import { Pool } from "pg";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20,
idleTimeoutMillis: 30000,
connectionTimeoutMillis: 2000,
});
// Graceful shutdown
process.on("SIGTERM", async () => {
await pool.end();
});
```
## Next Steps
- [Project Structure](/project-structure) - Organize your graph definitions as your project grows
- [Schemas & Types](/core-concepts) - Deep dive into nodes, edges, and schemas
- [Ontology](/ontology) - Learn about semantic relationships
- [Query Builder](/queries/overview) - Query patterns and traversals
- [Schemas & Stores](/schemas-stores) - Complete API documentation
# Schemas & Types
> Defining nodes, edges, and leveraging TypeScript inference
TypeGraph's power comes from its type system. Define your schema once with Zod, and get:
- **Runtime validation** on every create and update
- **TypeScript types** inferred automatically (no duplication)
- **Query builder constraints** that prevent invalid queries at compile time
## Contents
- [Nodes](#nodes) — Entities with properties and metadata
- [Defining Node Types](#defining-node-types)
- [Schema Features](#schema-features)
- [Node Operations](#node-operations)
- [Edges](#edges) — Relationships between nodes
- [Defining Edge Types](#defining-edge-types) (domain/range constraints)
- [Edge Constraints](#edge-constraints) (cardinality)
- [Edge Operations](#edge-operations)
- [Graph Definition](#graph-definition) — Combining nodes, edges, and ontology
- [Delete Behaviors](#delete-behaviors) — Restrict, cascade, disconnect
- [Uniqueness Constraints](#uniqueness-constraints) — Enforcing unique values
- [Type Inference](#type-inference) — Extracting TypeScript types from schemas
## Nodes
Nodes represent entities in your graph. Each node has:
- **Type**: The type of node (e.g., "Person", "Company")
- **ID**: A unique identifier within the graph
- **Props**: Properties defined by a Zod schema
- **Metadata**: Version, timestamps, and soft-delete state
### Defining Node Types
```typescript
import { z } from "zod";
import { defineNode } from "@nicia-ai/typegraph";
const Person = defineNode("Person", {
schema: z.object({
fullName: z.string().min(1),
email: z.string().email().optional(),
dateOfBirth: z.string().optional(),
tags: z.array(z.string()).default([]),
}),
description: "A person in the system", // Optional
});
```
### Schema Features
TypeGraph supports all Zod validation features:
```typescript
const Product = defineNode("Product", {
schema: z.object({
// Required string
name: z.string().min(1).max(200),
// Optional with default
status: z.enum(["draft", "active", "archived"]).default("draft"),
// Number with constraints
price: z.number().positive(),
// Array with items validation
categories: z.array(z.string()).min(1),
// Regex pattern
sku: z.string().regex(/^[A-Z]{2,4}-\d{4,8}$/),
// Nullable field
description: z.string().nullable(),
// Transform on validation
slug: z.string().transform((s) => s.toLowerCase().replace(/\s+/g, "-")),
}),
});
```
### Node Operations
```typescript
// Create with auto-generated ID
const node = await store.nodes.Person.create({ fullName: "Alice Smith" });
// Create with specific ID
const node = await store.nodes.Person.create({ fullName: "Alice Smith" }, { id: "person-alice" });
// Retrieve
const person = await store.nodes.Person.getById("person-alice");
// Update (partial)
const updated = await store.nodes.Person.update("person-alice", {
email: "alice@example.com",
});
// Delete (soft delete by default)
await store.nodes.Person.delete("person-alice");
// Hard delete (permanent removal) - use carefully!
await store.nodes.Person.hardDelete("person-alice");
```
### Node Object Shape
A node returned from the store has this structure:
```typescript
const alice = await store.nodes.Person.create({ name: "Alice", email: "a@example.com" });
// alice = {
// id: "01HX...", // Generated ULID (or your custom ID)
// kind: "Person", // The node type name
// name: "Alice", // Schema property (flattened to top level)
// email: "a@example.com", // Schema property
// meta: {
// version: 1,
// createdAt: "2024-01-15T10:30:00.000Z",
// updatedAt: "2024-01-15T10:30:00.000Z",
// deletedAt: undefined,
// validFrom: "2024-01-15T10:30:00.000Z", // defaults to createdAt when omitted,
// // unless a stated past validTo makes
// // the row "born already ended" (undefined)
// validTo: undefined,
// }
// }
```
Schema properties are flattened to the top level for ergonomic access (`alice.name` instead of
`alice.props.name`). System metadata lives under `meta`.
### Soft Delete vs Hard Delete
By default, `delete()` performs a **soft delete**—it sets the `deletedAt` timestamp but preserves the record:
```typescript
await store.nodes.Person.delete(alice.id); // Sets deletedAt, keeps the record
```
For permanent removal, use `hardDelete()`:
```typescript
await store.nodes.Person.hardDelete(alice.id); // Removes from database
```
**When to use each:**
| Method | Use Case |
|--------|----------|
| `delete()` | Standard deletions, audit trails, undo capability |
| `hardDelete()` | GDPR erasure, storage cleanup, removing test data |
**Warning:** `hardDelete()` is irreversible. It also removes associated uniqueness entries and
embeddings. Consider using soft delete for most use cases.
## Edges
Edges represent relationships between nodes. Each edge has:
- **Type**: The type of relationship (e.g., "worksAt", "knows")
- **ID**: A unique identifier
- **From**: Source node (type + ID)
- **To**: Target node (type + ID)
- **Props**: Properties defined by a Zod schema
### Defining Edge Types
```typescript
import { defineEdge } from "@nicia-ai/typegraph";
// Edge with properties
const worksAt = defineEdge("worksAt", {
schema: z.object({
role: z.string(),
startDate: z.string().optional(),
isPrimary: z.boolean().default(true),
}),
});
// Edge without properties
const knows = defineEdge("knows");
// Equivalent to: defineEdge("knows", { schema: z.object({}) })
```
#### Unconstrained Edges
Edges defined without `from` and `to` are **unconstrained** — they can connect any
node type to any node type. When used directly in `defineGraph`, they are automatically
allowed for all node types in the graph:
```typescript
const sameAs = defineEdge("sameAs");
const related = defineEdge("related", {
schema: z.object({ reason: z.string() }),
});
const graph = defineGraph({
id: "my_graph",
nodes: {
Person: { type: Person },
Company: { type: Company },
},
edges: {
sameAs, // any→any (Person↔Person, Person↔Company, Company↔Company)
related, // any→any, with properties
worksAt: { type: worksAt, from: [Person], to: [Company] }, // constrained
},
});
// All of these work:
await store.edges.sameAs.create(alice, bob, {}); // Person→Person
await store.edges.sameAs.create(alice, acme, {}); // Person→Company
await store.edges.sameAs.create(acme, alice, {}); // Company→Person
```
This is useful for semantic relationships like `sameAs`, `seeAlso`, `related`, or
`tagged` that apply broadly across node types.
#### Domain and Range Constraints
Edges can include built-in domain (source types) and range (target types) constraints
directly in their definition. This makes edge definitions self-contained and reusable:
```typescript
// Edge with built-in domain/range constraints
const worksAt = defineEdge("worksAt", {
schema: z.object({
role: z.string(),
startDate: z.string().optional(),
}),
from: [Person], // Domain: only Person can be the source
to: [Company], // Range: only Company can be the target
});
// Edge connecting multiple types
const mentions = defineEdge("mentions", {
from: [Article, Comment],
to: [Person, Company, Topic],
});
```
Any edge type can be used directly in `defineGraph` without an `EdgeRegistration`
wrapper. Constrained edges use their built-in `from`/`to`; unconstrained edges
allow all node types:
```typescript
const graph = defineGraph({
nodes: { Person: { type: Person }, Company: { type: Company } },
edges: {
worksAt, // Constrained - uses built-in from/to
sameAs, // Unconstrained - connects any node to any node
},
});
```
You can still use `EdgeRegistration` to narrow (but not widen) the constraints:
```typescript
const worksAt = defineEdge("worksAt", {
from: [Person],
to: [Company, Subsidiary], // Allows both Company and Subsidiary
});
const graph = defineGraph({
edges: {
// Narrow to only Subsidiary targets in this graph
worksAt: { type: worksAt, from: [Person], to: [Subsidiary] },
},
});
```
Attempting to widen beyond the edge's built-in constraints throws a `ConfigurationError`:
```typescript
const worksAt = defineEdge("worksAt", {
from: [Person],
to: [Company],
});
// This throws ConfigurationError - OtherEntity is not in the edge's range
defineGraph({
edges: {
worksAt: { type: worksAt, from: [Person], to: [OtherEntity] },
},
});
```
#### Source-Dependent Targets
An array-valued `to` allows every combination of the source and target types.
When the valid target depends on the source, use a map instead:
```typescript
const Employee = defineNode("Employee", { schema: z.object({ name: z.string() }) });
const Student = defineNode("Student", { schema: z.object({ name: z.string() }) });
const Department = defineNode("Department", { schema: z.object({ name: z.string() }) });
const Course = defineNode("Course", { schema: z.object({ name: z.string() }) });
const assignedTo = defineEdge("assignedTo", {
from: [Employee, Student],
to: {
Employee: [Department],
Student: [Course],
},
});
const graph = defineGraph({
id: "assignments",
nodes: {
Employee: { type: Employee },
Student: { type: Student },
Department: { type: Department },
Course: { type: Course },
},
edges: { assignedTo },
});
```
This permits `Employee → Department` and `Student → Course`. It rejects
`Employee → Course` and `Student → Department`. Using
`to: [Department, Course]` would permit all four combinations.
Map keys are the literal node kind names (`Employee.kind`), not aliases used to
register nodes in a graph. Every kind in `from` must have a map entry, no other
keys are allowed, and each target array must be nonempty. You can also use
computed keys such as `[Employee.kind]: [Department]`.
The map syntax works in an explicit graph registration too:
```typescript
edges: {
assignedTo: {
type: assignedTo,
from: [Employee],
to: { Employee: [Department] },
},
}
```
A registration may narrow the built-in allowed pairs, but it cannot introduce
new pairs. Replacing a map with arrays is valid only when every resulting
combination is already allowed by the edge definition.
At runtime, both endpoints must match the **same** declared pair, including
`subClassOf` assignability. A source matching several source entries can use the
targets allowed by any of those entries. An undeclared pair fails with
[`EndpointPairError`](/errors#endpointpairerror); an invalid source kind still
fails with `EndpointError`. Malformed declarations fail with `ConfigurationError`.
Typed collection writes preserve the source/target relationship; dynamic writes
and imports enforce it at runtime. Bulk writes reject invalid pairs atomically.
Import pair validation remains active even when reference validation is disabled;
imports retain their own documented error-handling and partial-success behavior.
See [collection types](/types#typededgecollectionr) for inference limits,
[graph extensions](/graph-extensions#edges) for runtime declarations, and
[schema management](/schema-management#endpoint-pair-changes) for schema changes.
### Edge Constraints
#### Cardinality
Control how many edges can exist:
```typescript
const graph = defineGraph({
edges: {
// Default: no limit
knows: { type: knows, from: [Person], to: [Person], cardinality: "many" },
// At most one edge of this type from any source node
currentEmployer: {
type: currentEmployer,
from: [Person],
to: [Company],
cardinality: "one",
},
// At most one edge between any (source, target) pair
rated: { type: rated, from: [Person], to: [Product], cardinality: "unique" },
// At most one active edge (valid_to IS NULL) from any source
currentRole: {
type: currentRole,
from: [Person],
to: [Company],
cardinality: "oneActive",
},
},
});
```
| Cardinality | Description |
|-------------|-------------|
| `"many"` | No limit (default) |
| `"one"` | At most one edge of this type from any source node |
| `"unique"` | At most one edge between any (source, target) pair |
| `"oneActive"` | At most one edge with `valid_to IS NULL` from any source |
#### Enforcement Timing
Cardinality constraints are checked at edge **creation time**, before the insert:
```typescript
// With cardinality: "one" on currentEmployer:
await store.edges.currentEmployer.create(alice, acme, {}); // OK
await store.edges.currentEmployer.create(alice, other, {}); // Throws CardinalityError
```
The check queries existing edges and throws `CardinalityError` if violated.
For `oneActive`, only edges with `validTo` unset count toward the limit.
### Edge Operations
```typescript
// Create edge - pass nodes directly
const edge = await store.edges.worksAt.create(alice, acme, { role: "Engineer" });
// Retrieve edge
const e = await store.edges.worksAt.getById(edge.id);
// Delete edge
await store.edges.worksAt.delete(edge.id);
```
## Graph Definition
The graph definition combines all components:
```typescript
import { defineGraph } from "@nicia-ai/typegraph";
const graph = defineGraph({
// Unique identifier for this graph
id: "my_application",
// Node registrations
nodes: {
Person: {
type: Person,
onDelete: "restrict", // Default behavior
},
Company: {
type: Company,
onDelete: "cascade",
},
Employment: {
type: Employment,
onDelete: "disconnect",
},
},
// Edge registrations
edges: {
worksAt: {
type: worksAt,
from: [Person],
to: [Company],
cardinality: "many",
},
employedAt: {
type: employedAt,
from: [Company],
to: [Employment],
cardinality: "many",
},
},
// Semantic relationships
ontology: [subClassOf(Company, Organization), disjointWith(Person, Company)],
});
```
## Delete Behaviors
Control what happens when nodes are deleted:
### Restrict (Default)
Blocks deletion if any edges are connected:
```typescript
nodes: {
Author: { type: Author }, // onDelete defaults to "restrict"
}
// This throws RestrictedDeleteError if Author has edges
await store.nodes.Author.delete(authorId);
```
### Cascade
Automatically deletes all connected edges:
```typescript
nodes: {
Book: { type: Book, onDelete: "cascade" },
}
// Deletes the book and all edges connected to it
await store.nodes.Book.delete(bookId);
```
### Disconnect
Soft-deletes edges (preserves history):
```typescript
nodes: {
Review: { type: Review, onDelete: "disconnect" },
}
// Marks connected edges as deleted (deleted_at is set)
await store.nodes.Review.delete(reviewId);
```
## Uniqueness Constraints
Ensure unique values within node types:
```typescript
const graph = defineGraph({
nodes: {
Person: {
type: Person,
unique: [
{
name: "person_email",
fields: ["email"],
where: (props) => props.email.isNotNull(),
scope: "kind",
collation: "caseInsensitive",
},
],
},
Company: {
type: Company,
unique: [
{
name: "company_ticker",
fields: ["ticker"],
scope: "kind",
collation: "binary",
},
],
},
},
});
```
### Scope Options
- `"kind"`: Unique within this exact type only
- `"kindWithSubClasses"`: Unique across this type and all subclasses
### Collation Options
- `"binary"`: Case-sensitive comparison
- `"caseInsensitive"`: Case-insensitive comparison
## Type Inference
TypeGraph infers TypeScript types from Zod schemas—you never duplicate type definitions.
### Extracting Types from Definitions
```typescript
import { z } from "zod";
import { defineNode, type Node, type NodeProps, type NodeId } from "@nicia-ai/typegraph";
const Person = defineNode("Person", {
schema: z.object({
name: z.string(),
email: z.string().email().optional(),
age: z.number().optional(),
}),
});
// For functions that work with full nodes (id, kind, metadata, props):
type PersonNode = Node;
// { id: NodeId; kind: "Person"; name: string; email?: string; version: number; createdAt: Date; ... }
// For functions that only need the property data:
type PersonProps = NodeProps;
// { name: string; email?: string; age?: number }
// For type-safe node IDs (prevents mixing IDs from different node types):
type PersonId = NodeId;
// string & { readonly [__nodeId]: typeof Person }
```
Use `Node` when your function needs the full node with metadata.
Use `NodeProps` when you only care about the schema properties (e.g., for form validation or API payloads).
### Typed Store Operations
```typescript
// Create returns a fully typed Node
const alice: Node = await store.nodes.Person.create({
name: "Alice",
email: "alice@example.com",
});
// TypeScript knows the structure
alice.id; // NodeId - branded string
alice.name; // string
alice.email; // string | undefined
alice.age; // number | undefined
alice.version; // number
alice.createdAt; // Date
// Type errors caught at compile time
await store.nodes.Person.create({
name: 123, // Error: Type 'number' is not assignable to type 'string'
invalid: "field", // Error: Object literal may only specify known properties
});
```
### Typed Query Results
```typescript
// Result type is inferred from your select projection
const results = await store
.query()
.from("Person", "p")
.select((ctx) => ({
name: ctx.p.name, // TypeScript knows: string
email: ctx.p.email, // TypeScript knows: string | undefined
id: ctx.p.id, // TypeScript knows: NodeId
}))
.execute();
// results: Array<{ name: string; email: string | undefined; id: NodeId }>
// Invalid property access is caught
.select((ctx) => ({
invalid: ctx.p.nonexistent, // TypeScript error!
}))
```
### Typed Edge Operations
Edge endpoints are constrained to valid node types:
```typescript
// Edge definition: worksAt goes from Person → Company
const graph = defineGraph({
// ...
edges: {
worksAt: { type: worksAt, from: [Person], to: [Company] },
},
});
// TypeScript enforces valid endpoints
await store.edges.worksAt.create(alice, acmeCorp, { role: "Engineer" }); // OK
await store.edges.worksAt.create(acmeCorp, alice, { role: "Engineer" });
// Error: Argument of type 'Node' is not assignable to parameter of type 'Node'
```
# Backend Setup
> Configure SQLite and PostgreSQL backends for TypeGraph
TypeGraph stores graph data in your existing relational database using Drizzle ORM adapters.
This guide covers setting up SQLite, PostgreSQL, and PGlite backends.
:::note[Custom indexes]
TypeGraph migrations create the core tables and built-in indexes. For application-specific indexes
on JSON properties (and Drizzle/drizzle-kit integration), see [Indexes](/performance/indexes).
:::
## SQLite
SQLite is ideal for development, testing, single-server deployments, and embedded applications.
### Quick Setup
For development and testing, use the convenience function that owns the
connection and provisions TypeGraph's base tables:
```typescript
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
import { createStore } from "@nicia-ai/typegraph";
// In-memory database (resets on restart)
const { backend } = createLocalSqliteBackend();
const store = createStore(graph, backend);
// File-based database (persisted)
const { backend, db } = createLocalSqliteBackend({ path: "./app.db" });
const store = createStore(graph, backend);
```
The local backend owns its connection, so it applies performance pragmas at
open: `journal_mode=WAL`, `synchronous=NORMAL`, and a 5s `busy_timeout`. On
file databases this makes single-operation writes roughly 5× faster than the
driver defaults (rollback journal, `synchronous=FULL`). Override individual
values or opt out entirely:
```typescript
// Override one value, keep the other defaults
createLocalSqliteBackend({ path: "./app.db", pragmas: { busyTimeoutMs: 10_000 } });
// Keep better-sqlite3's driver defaults untouched
createLocalSqliteBackend({ path: "./app.db", pragmas: false });
```
:::caution[Fulltext and embeddings require `createStoreWithSchema`]
`createLocalSqliteBackend` creates the base tables but does not durably
materialize strategy-owned storage. If your graph has `searchable()` or
`embedding()` fields, boot with
`const [store] = await createStoreWithSchema(graph, backend);` instead of
bare `createStore()` — otherwise the first fulltext or embedding operation
throws `StoreNotInitializedError`.
:::
### Manual Setup
For full control over the database connection:
```typescript
import Database from "better-sqlite3";
import { drizzle } from "drizzle-orm/better-sqlite3";
import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { createStoreWithSchema } from "@nicia-ai/typegraph";
// Create and configure the database
const sqlite = new Database("app.db");
sqlite.pragma("journal_mode = WAL"); // Recommended for performance
sqlite.pragma("foreign_keys = ON");
// Create Drizzle instance and backend
const db = drizzle(sqlite);
const backend = createSqliteBackend(db);
// createStoreWithSchema auto-creates tables on first run
const [store] = await createStoreWithSchema(graph, backend);
// Clean up when done
process.on("exit", () => sqlite.close());
```
For a fresh database whose DDL is managed externally, use
`generateSqliteMigrationSQL()` with `createStore()` instead:
```typescript
sqlite.exec(generateSqliteMigrationSQL());
const store = createStore(graph, backend);
```
The generated script is complete installation DDL and stamps the current
deployment-wide base-schema marker last; it is not an incremental upgrade
planner. Existing databases attached only through the zero-DDL runtime
factories must apply release-specific additive migrations through their
migration tool. See
[Upgrading deployment-wide base storage](#upgrading-deployment-wide-base-storage)
for the exact SQLite and PostgreSQL statements. A privileged
`createStoreWithSchema()` open adopts missing release storage once, then stamps
a deployment-wide base-schema marker. Warm opens read that marker and issue no
base-adoption DDL.
### SQLite with Vector Search
For semantic search, use the sqlite-vec extension. `createLocalSqliteBackend()` wires the
`sqliteVecStrategy` automatically when the extension loads. For a bring-your-own connection, load the
extension and pass the strategy explicitly:
```typescript
import Database from "better-sqlite3";
import { drizzle } from "drizzle-orm/better-sqlite3";
import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { sqliteVecStrategy } from "@nicia-ai/typegraph";
const sqlite = new Database("app.db");
// Load sqlite-vec extension
sqlite.loadExtension("vec0");
// Run migrations (core tables)
sqlite.exec(generateSqliteMigrationSQL());
const db = drizzle(sqlite);
const backend = createSqliteBackend(db, { vector: sqliteVecStrategy });
```
sqlite-vec stores embeddings in `vec0` virtual tables and supports the `cosine` and `l2` metrics. Per-field
vector tables are provisioned by `createStoreWithSchema` at boot (not by the generated migration SQL), and the
runtime asserts a durable marker rather than issuing DDL on first write — see
[Database roles & least privilege](#database-roles--least-privilege).
See [Semantic Search](/semantic-search) for query examples.
### libsql / Turso
For edge deployments, shared-driver setups, or Turso cloud databases, use the first-class
libsql backend:
```bash
npm install @libsql/client
```
```typescript
import { createClient } from "@libsql/client";
import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql";
import { createStore } from "@nicia-ai/typegraph";
// Local file
const client = createClient({ url: "file:app.db" });
// Or remote Turso database
// const client = createClient({ url: "libsql://my-db.turso.io", authToken: "..." });
const { backend, db } = await createLibsqlBackend(client);
const store = createStore(graph, backend);
```
`createLibsqlBackend` handles DDL execution and configures the correct async
execution profile automatically. It returns both the `backend` and the underlying
Drizzle `db` instance for direct SQL access. The caller retains ownership of the
client and is responsible for closing it when done — this allows sharing a single
client across TypeGraph and other libraries. Its installation is complete: the
factory publishes the deployment-wide base-schema marker, and when it encounters
a pre-0.52 edge table it applies the focused match-identity storage adoption
before retrying the idempotent installation script. The local SQLite factory has
the same behavior.
The libsql backend has native vector and hybrid search, wired automatically via `libsqlVectorStrategy` — no
extension to load. It uses libSQL's built-in engine (`F32_BLOB(N)` storage, `vector_distance_cos` /
`vector_distance_l2`, and DiskANN approximate nearest neighbor via `libsql_vector_idx` + `vector_top_k`) and
supports the `cosine` and `l2` metrics. See [Semantic Search](/semantic-search) for query examples.
:::caution[In-memory databases and transactions]
libsql's `file::memory:` creates a separate database per connection. Since transactions
open a new connection, the original database is destroyed after a transaction completes
([tursodatabase/libsql-client-ts#229](https://github.com/tursodatabase/libsql-client-ts/issues/229)).
Use a file-based database (`file:path.db`) or remote URL when transactions are needed.
:::
### API Reference
#### `createLocalSqliteBackend(options?)`
Creates a SQLite backend with automatic database and schema setup.
```typescript
function createLocalSqliteBackend(options?: {
path?: string; // Database path, defaults to ":memory:"
tables?: SqliteTables;
/**
* Override the fulltext strategy. Defaults to `fts5Strategy` (SQLite's
* built-in FTS5 virtual table). Pass `false` to disable fulltext support
* entirely — the backend then advertises no `capabilities.fulltext` and
* omits the fulltext CRUD/search methods, and the managed installation
* never creates the fulltext table. Forwarded to both the installation
* DDL and `createSqliteBackend`.
*/
fulltext?: FulltextStrategy | false;
}): { backend: GraphBackend; db: BetterSQLite3Database };
```
#### `createSqliteBackend(db, options?)`
Creates a SQLite backend from an existing Drizzle database instance. Pass `vector` to enable vector search
(for example `sqliteVecStrategy` after loading the sqlite-vec extension).
```typescript
function createSqliteBackend(
db: BetterSQLite3Database,
options?: {
tables?: SqliteTables;
/**
* Override the fulltext strategy. Defaults to `fts5Strategy` (SQLite's
* built-in FTS5 virtual table). Pass `false` to disable fulltext
* support entirely — the backend then advertises no
* `capabilities.fulltext` and omits the fulltext CRUD/search methods,
* mirroring `vector` left unset. Required for a SQLite build without
* FTS5 compiled in.
*/
fulltext?: FulltextStrategy | false;
vector?: VectorStrategy;
capabilities?: BundledBackendCapabilityOverrides;
},
): GraphBackend;
```
Pass `{ fulltext: false }` on a SQLite build without FTS5 compiled in, or
whenever the graph has no `searchable()` fields and you would rather skip
the virtual table than carry it unused:
```typescript
const backend = createSqliteBackend(db, { fulltext: false });
```
#### `generateSqliteMigrationSQL()`
Returns complete fresh-installation SQL for creating TypeGraph tables and
stamping the current deployment-wide base-schema marker in SQLite.
```typescript
function generateSqliteMigrationSQL(
tables?: SqliteTables,
fulltextStrategy?: FulltextStrategy | false,
): string;
```
`generateSqliteDDL()` is the lower-level table/index statement array used by
backend bootstrap. It deliberately omits the deployment-wide marker row and is
therefore not a complete installation script. Use `generateSqliteMigrationSQL()`
when the resulting database will be opened through `createVerifiedStore()` or
the DML-only graph-template APIs.
#### `createLibsqlBackend(client, options?)`
Creates a SQLite backend from a `@libsql/client` instance. Runs DDL automatically.
The caller retains ownership of the client and is responsible for closing it.
```typescript
async function createLibsqlBackend(client: Client, options?: { tables?: SqliteTables }): Promise<{ backend: GraphBackend; db: LibSQLDatabase }>;
```
## PostgreSQL
PostgreSQL is recommended for production deployments with concurrent access, large datasets,
or when you need advanced features like pgvector.
`createPostgresBackend` is driver-agnostic. Pick the Drizzle adapter that matches your
runtime, and TypeGraph works the same way against each.
### Choosing a PostgreSQL driver
| Runtime | Recommended driver | Drizzle adapter |
| --------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | -------------------------------------------------------- |
| Long-lived Node server (Fly, Render, Cloud Run, containers) | `pg` (node-postgres) or `postgres` (postgres-js) | `drizzle-orm/node-postgres` or `drizzle-orm/postgres-js` |
| Node serverless (Vercel Functions, AWS Lambda, Netlify Functions) | `postgres` (postgres-js) — faster cold start, lower per-query overhead | `drizzle-orm/postgres-js` |
| Bun server | `postgres` (postgres-js) or Bun's built-in SQL | `drizzle-orm/postgres-js` or `drizzle-orm/bun-sql` |
| Edge runtime (Cloudflare Workers, Vercel Edge, Netlify Edge) — needs transactions | `@neondatabase/serverless` Pool over WebSockets | `drizzle-orm/neon-serverless` |
| Edge runtime — single-statement reads/writes only | `@neondatabase/serverless` `neon(url)` over HTTP | `drizzle-orm/neon-http` |
| Cloudflare Hyperdrive | `pg` or `postgres` (through the Hyperdrive pooler) | `drizzle-orm/node-postgres` or `drizzle-orm/postgres-js` |
| Embedded apps, local development, Postgres dialect tests | `@electric-sql/pglite` | `drizzle-orm/pglite` |
:::note[Neon HTTP vs WebSocket]
Both Neon drivers work with TypeGraph. They have different tradeoffs:
- **`drizzle-orm/neon-http`** uses HTTP per statement. Lowest cold-start cost; survives Workers'
per-request isolation. **Cannot hold a session across statements**, so multi-statement transactions
are unavailable — TypeGraph auto-detects this driver and sets `capabilities.execution.interactiveTransactions = false`,
so `store.transaction(...)` refuses rather than pretending to provide rollback. Eligible
atomic-batch operations remain available when the transport is certified for them.
A schema-managed Store's write fuses its schema fence into the write's own statement when the
write fuses, and fails closed otherwise — see
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for which
writes fuse and the reasons a write that cannot refuses with.
- **`drizzle-orm/neon-serverless`** uses a WebSocket Pool. Holds a session, supports full transactional
semantics, but the WebSocket connection lifecycle needs care in serverless / per-request contexts
(you typically want a fresh Pool per request).
Pick HTTP for stateless reads and for the fused schema-managed writes. Pick WebSockets for schema
migrations, and for any write outside that fused envelope.
:::
### node-postgres (pg)
The default choice for long-lived Node servers. Widest ecosystem and most deployment
documentation.
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20,
});
const db = drizzle(pool);
const backend = createPostgresBackend(db);
const [store] = await createStoreWithSchema(graph, backend);
```
For a fresh database managed externally, use `generatePostgresMigrationSQL()` with `createStore()`:
```typescript
import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
await pool.query(generatePostgresMigrationSQL());
const store = createStore(graph, backend);
```
As with SQLite, this is complete installation DDL rather than an incremental
upgrade plan. Apply the
[base-schema upgrade](#upgrading-deployment-wide-base-storage)
to an existing database, or let a privileged `createStoreWithSchema()`
preparation adopt the storage before runtime workers use `createStore()`.
### postgres-js
A leaner Postgres client with lower per-query overhead and smaller bundle size. Good
default for Node serverless platforms and Bun. Fully tested against TypeGraph's adapter
and integration suites.
```bash
npm install postgres drizzle-orm
```
```typescript
import postgres from "postgres";
import { drizzle } from "drizzle-orm/postgres-js";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const sql = postgres(process.env.DATABASE_URL, {
max: 10,
idle_timeout: 30,
});
const db = drizzle(sql);
const backend = createPostgresBackend(db);
const [store] = await createStoreWithSchema(graph, backend);
```
Transactions go through `sql.begin(fn)`; TypeGraph handles this automatically via
Drizzle's `db.transaction()`. Isolation levels are honored the same way as with
node-postgres.
### Neon serverless (WebSockets)
For edge runtimes like Cloudflare Workers, Vercel Edge, and Netlify Edge — anywhere
native TCP sockets aren't available. Neon's `@neondatabase/serverless` driver speaks
the Postgres wire protocol over WebSockets and exposes a pg-Pool-compatible API.
```bash
npm install @neondatabase/serverless drizzle-orm
```
```typescript
import { Pool } from "@neondatabase/serverless";
import { drizzle } from "drizzle-orm/neon-serverless";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const pool = new Pool({ connectionString: env.NEON_DATABASE_URL });
const db = drizzle(pool);
const backend = createPostgresBackend(db);
const [store] = await createStoreWithSchema(graph, backend);
```
When running under Node.js (for local testing), install `ws` and configure it once
before connecting:
```typescript
import { neonConfig } from "@neondatabase/serverless";
import ws from "ws";
neonConfig.webSocketConstructor = ws;
```
Edge runtimes expose `WebSocket` globally and need no extra setup.
### Neon HTTP
For stateless edge workloads where you don't need transactional writes. The HTTP
driver issues one request per query — lowest cold-start cost, no session lifecycle
to manage. TypeGraph auto-detects this driver and sets
`capabilities.execution.interactiveTransactions` to `false` and
`capabilities.execution.unitOfWork` to `"batch"`. On a raw Store,
`store.transaction(...)` refuses rather than silently falling through to
sequential execution.
A schema-managed or verified Store's first write does not universally fail
closed here — it depends on whether the write fuses. A singleton node
create, update, `upsertById`, or delete fuses on a kind with no declared
unique constraint (a create takes a generated or a caller-supplied id) —
except a node delete, which fuses even when the kind DOES carry a declared
unique constraint, because the atomic delete program releases that claim in
the same statement. A singleton edge create fuses when the kind's
cardinality is `"many"`, and edge update and delete fuse the same way
(`EdgeCollection` has no `upsertById`). So do
`bulkInsert`/`bulkCreate`/`bulkDelete`/`bulkReplaceById`/`bulkUpsertById`,
and a constrained write inside an atomic program's claim envelope. Each of
these asserts the active schema version inside the statements neon-http
submits together, and `transaction(queries)` commits or rejects that
submission as a whole. A write that cannot fuse either fails closed with
`BATCH_WRITE_UNSUPPORTED` naming a proven reason (an interactive callback, a
probe-then-write constraint check, Operational Identity, history, or a
schema commit), or — for a write that simply doesn't fit the fused shape,
such as a singleton create, update, or `upsertById` on a uniquely-constrained
kind, or a supplied-id tombstone resurrection — fails closed with the plain
`SCHEMA_WRITE_FENCE_UNSUPPORTED` limitation and no named reason. See
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares)
for the shared guard and the full reason table.
Schema commits stay refused regardless: `commitSchemaVersion` and
`setActiveVersion` require holding one transaction across their
compare-and-swap read and activating write to eliminate the orphan-row crash
window they exist to fix, so they refuse with a typed `ConfigurationError` on
non-transactional backends. Run schema migrations from a process with a
transactional driver (`drizzle-orm/neon-serverless`, regular `pg`, etc.); the
edge worker can keep using neon-http for reads and for the fused writes
above. A raw `createStore()` remains available for writes outside that
envelope when the application explicitly accepts they are not fenced against
schema changes.
```bash
npm install @neondatabase/serverless drizzle-orm
```
```typescript
import { neon } from "@neondatabase/serverless";
import { drizzle } from "drizzle-orm/neon-http";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStore } from "@nicia-ai/typegraph";
const sql = neon(env.NEON_DATABASE_URL);
const db = drizzle({ client: sql });
const backend = createPostgresBackend(db);
const store = createStore(graph, backend);
// backend.capabilities.execution.interactiveTransactions === false (auto-detected)
```
Use `neon-http` for reads and for the fused schema-managed writes listed
above. Run schema migrations, and any write outside that envelope, through
`neon-serverless`, regular `pg`, or another transactional driver.
### PGlite (Postgres-in-WASM)
[PGlite](https://pglite.dev/) is a full Postgres compiled to WebAssembly that runs
in-process — in Node, Bun, Deno, or the browser — with no server and no native
addon. It's ideal for local development, embedded apps, and running the real
Postgres dialect (including pgvector) in tests without Docker.
`@electric-sql/pglite` is an optional peer dependency. Vector support additionally
needs `@electric-sql/pglite-pgvector` (PGlite ≥ 0.5 ships pgvector as a separate
package):
```bash
npm install @electric-sql/pglite @electric-sql/pglite-pgvector
```
The batteries-included helper constructs the engine, loads pgvector, runs the
schema DDL, and returns a ready backend — the Postgres analog of
`createLocalSqliteBackend`:
```typescript
import { createLocalPgliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres/pglite";
import { createStore } from "@nicia-ai/typegraph";
// In-memory by default, with pgvector enabled.
const { backend, db, client } = await createLocalPgliteBackend();
const store = createStore(graph, backend);
// backend.close() disposes the PGlite engine.
```
```typescript
// Persistent on disk:
const { backend } = await createLocalPgliteBackend({ dataDir: "./pgdata" });
// No embeddings? Skip the extension (no pgvector dependency needed):
const { backend } = await createLocalPgliteBackend({ vector: false });
// Pass an explicit pgvector extension object:
import { vector } from "@electric-sql/pglite-pgvector";
const { backend } = await createLocalPgliteBackend({ vector });
```
If you construct PGlite yourself, pass its Drizzle database straight to
`createPostgresBackend` — the execution fast path detects PGlite and routes it
correctly:
```typescript
import { PGlite } from "@electric-sql/pglite";
import { vector } from "@electric-sql/pglite-pgvector";
import { drizzle } from "drizzle-orm/pglite";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const client = await PGlite.create({ extensions: { vector } });
await client.exec(generatePostgresMigrationSQL());
const backend = createPostgresBackend(drizzle(client));
```
PGlite is single-connection and serial: there is no pooling, so concurrent
`store.transaction()` calls queue rather than run in parallel. It complements,
rather than replaces, a Docker-based Postgres for CI — PGlite exercises the SQL
dialect and pgvector, but not driver-specific behavior (node-postgres statement
naming, postgres-js, pgbouncer, real concurrency).
### PostgreSQL with Vector Search
For semantic search, enable pgvector. `createPostgresBackend` defaults to `pgvectorStrategy`, so no extra
wiring is required:
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
// Migration SQL enables the pgvector extension
await pool.query(generatePostgresMigrationSQL());
// Runs: CREATE EXTENSION IF NOT EXISTS vector;
const db = drizzle(pool);
const backend = createPostgresBackend(db);
```
pgvector stores embeddings in per-field typed `vector(N)` tables (provisioned by `createStoreWithSchema` at boot
— the generated migration SQL creates no embedding table) with HNSW or IVFFlat indexes, and supports the
`cosine`, `l2`, and `inner_product` metrics.
See [Semantic Search](/semantic-search) for query examples.
### Refreshing planner statistics after bulk loads
`importGraph()` refreshes planner statistics automatically after an import
that created or updated rows, and `store.materializeIndexes()` does the
same on SQLite after creating indexes (pass `refreshStatistics: false` to
opt out). On PostgreSQL, `materializeIndexes()` builds with
`CREATE INDEX CONCURRENTLY` and skips the automatic refresh — call
`store.refreshStatistics()` after materializing.
`bulkCreate` and `bulkInsert` on nodes and edges also refresh
automatically when a single autocommit call writes 1,000 rows or more. Tune or disable this
with the `autoRefreshStatistics` store option:
```typescript
// Refresh after any autocommit bulkCreate of 5,000+ rows
const store = createStore(graph, backend, { autoRefreshStatistics: 5000 });
// Never refresh automatically after bulkCreate
const store = createStore(graph, backend, { autoRefreshStatistics: false });
```
Bulk writes inside a `store.transaction(...)` block never auto-refresh —
statistics collected mid-transaction cannot see the uncommitted rows —
so refresh manually after the transaction commits. The same applies to
loops of small `bulkCreate` batches that never individually reach the
threshold, and to backend-level batch inserts — the loop example below
covers that pattern.
PostgreSQL's query planner relies on table statistics to choose
between multi-column indexes on `typegraph_edges` (forward vs reverse vs
cardinality), and when those statistics are stale the planner can pick a
reverse-index scan with a filter — turning a 0.5ms forward traversal into a
5ms one. SQLite's planner is similarly sensitive: without `sqlite_stat1`
data, some FTS5 fulltext queries fall back to a plan that's roughly 30×
slower. Autovacuum / background statistics collection will catch up
eventually, but refreshing explicitly gives correct latencies immediately.
```typescript
for (const batch of batches) {
await store.nodes.Document.bulkCreate(batch);
}
await store.refreshStatistics();
```
The implementation runs `ANALYZE` against the TypeGraph-managed tables in
the configured backend — the call is safe regardless of custom table names
or fulltext / embedding configuration. Cloudflare D1 and Durable Object SQLite
reject the performance-only `PRAGMA analysis_limit` tuning statement through
their authorizer. TypeGraph recognizes only that `SQLITE_AUTH` failure and
continues with scoped `ANALYZE`; workerd permits `ANALYZE`, so planner statistics
are still refreshed but without bounded sampling. Unexpected PRAGMA or ANALYZE
failures stay visible through the existing caller warning or rejection. If you
need to bypass the API for an unusual deployment (for example issuing `ANALYZE`
over a separate admin connection), call `backend.execute()` with raw SQL as the
escape hatch.
### pgbouncer / transaction-pool mode
By default, the node-postgres / neon-serverless fast path issues server-side
prepared statements (`client.query({name, text, values})`) so PostgreSQL
caches the parsed plan per session. This is incompatible with pgbouncer in
transaction-pool mode: pgbouncer routes successive statements over different
backend connections, so a `name` registered on one connection isn't visible
on the next. Pass `prepareStatements: false` to fall back to unnamed
positional queries:
```typescript
const backend = createPostgresBackend(db, {
prepareStatements: false, // pgbouncer transaction-pool compatibility
});
```
The in-process cache that maps SQL text → statement name is LRU-bounded
(default 256 entries, override via `preparedStatementCacheMax`). Eviction
never recycles a name, because a live connection may still retain that name for
its original SQL. Therefore this setting does not bound server-side prepared
statement memory. For a high-cardinality stream of SQL text, use
`prepareStatements: false` instead.
### Adopted schema transactions
`store.withEvolvedTransaction(nativeTx, plan, callback, { waitBudgetMs })`
requires an initialized adapter Store and a live caller-owned transaction.
Plan the extension outside that transaction with `store.planEvolution()`.
For change plans, the default exclusive schema-fence wait budget is 5,000 ms; a `SchemaFenceTimeoutError`
requires rollback and retry of the complete application transaction. Omit
`waitBudgetMs` for no-op plans, which use ordinary adoption without the exclusive
fence and refuse that option.
Interactive PostgreSQL adapters validate the active session, retain the existing
schema advisory lock → schema row → recorded-write lock order, and use
transaction-scoped advisory locks. This lock lifetime is suitable for transaction
poolers such as Hyperdrive. The adapter restores temporary timeout settings before
the callback. Noninteractive HTTP drivers cannot adopt schema transactions.
SQLite schema adoption requires an active transaction on the backend's exact
native connection with an observable `inTransaction` state, as provided by
better-sqlite3. Drivers without that evidence refuse schema adoption; ordinary
transaction support alone does not imply support for this operation. A deferred
SQLite transaction acquires the writer slot before validating the schema plan.
Adapters default to a DML-only schema provisioning policy. Plans requiring new
vector slots or identity work refuse before taking a mutating fence, running
DDL, or changing schema rows. A privileged adapter configured with
`schemaProvisioning: "transactional"` can apply those plans: it revalidates
storage on the pinned caller session and provisions identity relations, vector
tables, and durable contribution markers inside that same transaction. The
caller must roll back the entire native transaction if any step fails.
```typescript
const backend = createPostgresBackend(db, {
schemaProvisioning: "transactional",
});
```
Use a connection with permission to run the required DDL for this adapter;
keep the default policy for a runtime role limited to DML.
Bootstrap base storage before this request path; missing bootstrap tables
refuse rather than being created lazily. Database permissions still determine
whether transactional DDL succeeds. Generic eager index materialization,
including concurrent PostgreSQL indexes, remains an explicit post-commit
operation on the refreshed Store.
Custom adapters must implement `adoptSchemaWriteTransaction` with the same
session-bound fencing, finite-wait, and CAS guarantees to support change plans.
See [Graph Extensions](/graph-extensions) for callback and receipt usage.
### Authoritative command sessions
Store create paths use the backend's `commands` port for writes whose
decision and mutation must share one command boundary. First-party paths pass
an explicit command context: a root port owns any internal transaction it
needs and cannot inherit caller coordination, while a transaction-scoped
backend uses the active caller or Store transaction. A
transaction command may additionally carry a coordination token only after it
has acquired the graph's advisory lock; the token is bound to that graph and
transaction session and cannot authorize work on another connection.
On PostgreSQL, the lock statement also observes the effective transaction
isolation and binds it to the same token. Match-key convergence therefore
accepts only read committed or serializable based on database state, not the
caller-requested option or the server's assumed default.
`GraphBackend.commands` is a required member as of the authoritative command
port release. Custom backends must expose `{ session, execute }` and implement
the `node.create`, `edge.create`, and `edge.converge-create` commands, or return
a typed `unsupported` result for dimensions they do not provide. The former
optional managed-create and specialized edge-insert hooks are no longer a
complete backend implementation; migrate those branches into the command
port before upgrading.
For a custom backend, the migration shape is:
```typescript
const commands: GraphCommandPort = {
session: "transaction", // use "root" for a single-statement backend
execute(command, context) {
// Apply every requested dimension, or explicitly refuse the command.
switch (command.kind) {
case "node.create": {
return { outcome: "unsupported", entity: "node", dimensions: ["claims"] };
}
case "edge.create": {
return {
outcome: "unsupported",
entity: "edge",
dimensions: ["endpointPredicate"],
};
}
case "edge.converge-create": {
return { outcome: "unsupported", entity: "edge", dimensions: ["convergence"] };
}
}
},
};
const backend: GraphBackend = { ...members, commands };
```
Every command port caller must provide the explicit context. TypeGraph-owned
write paths use the command helper, which verifies that any coordination token
belongs to the active graph and transaction session and carries a supported
effective isolation before executing convergence. The portable PostgreSQL
graph-lock path records that isolation automatically. A custom implementation
of `lockSchemaVersionAndGraphWrite` must return the normalized
`GraphCommandIsolation` observed by its combined lock statement. When
decorating a first-party backend with `deriveBackend`, a same-session
`commands` override retains the session identity. A wrapper that changes
session or forwards to a different connection is a new command boundary and
cannot reuse a token from the original port.
These are four different execution guarantees; do not use “atomic” as a
catch-all:
- **Interactive transaction** (`store.transaction(...)`) pins one session and
can make several Store operations commit or roll back together. The
`runOptionallyInTransaction` callback receives
`{ mode: "interactive-transaction" }` when this boundary was opened, or
`{ mode: "sequential" }` on a backend without transaction support.
- **Static internal adapter batch** is an adapter implementation detail (for
example, a D1 batch or a bind-budgeted multi-row insert). It may make one
precompiled set of statements atomic, but it is not a public Store
transaction and does not make an arbitrary sequence of Store calls atomic.
- **Certified atomic SQL program** is the backend-authoring transport seam for
a closed, ordered sequence of statements. A backend earns this capability by
passing the framework-agnostic conformance runner: result slots and bound
parameters must be preserved, a failure in a later statement must leave no
primary or sidecar writes, and an empty program must be a no-op. Certification
is separate from semantic mutation eligibility; a transport alone does not
authorize a mutation family. Bundled recognized PostgreSQL drivers provide
this boundary either through Neon HTTP's transaction batch or a pinned
interactive transaction; an unrecognized driver leaves it unavailable.
- **Authoritative one-statement command** is the `commands.execute` port. A
command returns a created/found/rejected/unsupported result after the
database statement itself owns the decision and mutation. It is the
transactionless path for eligible durable edge `matchIdentity` convergence;
it is not a promise that every command or side effect can be fused.
Operational Identity, single-edge claim/cardinality checks, and undeclared
dynamic `matchOn` convergence remain interactive-transaction contracts. A
custom or non-transactional backend must refuse those dimensions rather than
silently falling through to a sequence of independent statements. Eligible
direct edge batches on bundled roots are a narrower static-program contract:
the insert and cardinality sidecars execute in one native atomic exchange. A
declared durable edge `matchIdentity` is different for endpoint convergence:
its canonical key has a database arbiter, so the eligible root create/found
command can be authoritative in one statement.
Backend implementations may also expose the optional
`findEdgesByMatchIdentity` read capability for bounded merge planning. It must
match the complete `(graphId, kind, name, key)` tuple and return tombstoned
owners as well as active rows; omitting it keeps the portable full-clone path.
Custom Drizzle operation strategies can opt in by supplying the corresponding
owner-query builder. A strategy without that builder does not expose the
capability, so callers can detect and retain the portable path.
### Connection Pooling
For production, always use connection pooling:
```typescript
import { Pool } from "pg";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Maximum pool size
idleTimeoutMillis: 30000, // Close idle connections after 30s
connectionTimeoutMillis: 2000, // Timeout for new connections
});
// Handle pool errors
pool.on("error", (err) => {
console.error("Unexpected pool error", err);
});
// Graceful shutdown
process.on("SIGTERM", async () => {
await pool.end();
process.exit(0);
});
```
### API Reference
#### `createPostgresBackend(db, options?)`
Creates a PostgreSQL backend adapter. Accepts any Drizzle PostgreSQL database
instance, regardless of the underlying driver. Tested with `drizzle-orm/node-postgres`,
`drizzle-orm/postgres-js`, `drizzle-orm/neon-serverless`,
`drizzle-orm/neon-http`, and `drizzle-orm/pglite`. The neon-http driver is auto-detected and
`capabilities.execution.interactiveTransactions` is set to `false` (HTTP can't hold a session); use
`drizzle-orm/neon-serverless` if you need transactional writes.
```typescript
function createPostgresBackend(
db: AnyPgDatabase,
options?: {
tables?: PostgresTables;
/**
* Override the fulltext strategy. Defaults to `tsvectorStrategy`.
* Pass a custom `FulltextStrategy` to swap the fulltext stack, or
* `false` to disable fulltext support entirely — the backend then
* advertises no `capabilities.fulltext` and omits the fulltext
* CRUD/search methods, mirroring `vector: false`.
*/
fulltext?: FulltextStrategy | false;
/**
* Override the vector search strategy. Defaults to
* `pgvectorStrategy`. Pass a custom `VectorStrategy` to change the
* storage / index engine, or `false` to disable vector support.
*/
vector?: VectorStrategy | false;
/**
* Override specific backend capabilities. Useful for HTTP-style
* drivers or test scenarios. neon-http already has
* `execution.interactiveTransactions: false` auto-applied — pass
* this to override that or to disable other capabilities for custom
* drivers.
*/
capabilities?: BundledBackendCapabilityOverrides;
/**
* Use server-side prepared statements on the node-postgres /
* neon-serverless fast path. Default `true`. Set to `false` when
* pooling through pgbouncer in transaction-pool mode (named
* statements are invisible across pooled connections).
*/
prepareStatements?: boolean;
/**
* LRU cap on the number of distinct SQL strings tracked for
* prepared-statement naming. Default 256. Worst-case server-side
* footprint is roughly `cap × pool size` prepared statements.
* Ignored when `prepareStatements` is `false`.
*/
preparedStatementCacheMax?: number;
},
): GraphBackend;
```
Pass `{ fulltext: false }` when the graph has no `searchable()` fields and
you would rather skip the fulltext table (`typegraph_node_fulltext`) and its
GIN index than carry them unused:
```typescript
const backend = createPostgresBackend(db, { fulltext: false });
```
#### `createPostgresTransactionBackend(tx, options?)`
Creates a full backend on a Drizzle PostgreSQL transaction opened by the
application. Use it when TypeGraph's tables share a transaction with other
application tables, especially when TypeGraph uses prefixed table names. Pass
the same `PostgresBackendOptions` as `createPostgresBackend`:
```typescript
import {
createPostgresTransactionBackend,
createPostgresTables,
} from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const graphTables = createPostgresTables({ nodes: "app_graph_nodes" });
await db.transaction(async (tx) => {
const backend = createPostgresTransactionBackend(tx, {
tables: graphTables,
});
// Use the backend or a Store built from it within this callback.
});
```
The factory requires a transaction handle and serializes TypeGraph statements
on its single pinned connection, including concurrent reads started by the
same Store operation. Backends created for the same transaction handle share
one queue. The application owns commit and rollback and must await all work
using these backends before its transaction callback returns.
`createPostgresBackend(tx)` also routes a PostgreSQL transaction handle to the
transaction-scoped backend automatically. Use `createPostgresTransactionBackend`
when you want the transaction-scoped intent to be explicit; a regular database
handle passed to `createPostgresBackend(db)` still creates the pooled backend.
#### `createLocalPgliteBackend(options?)`
Creates an in-process PGlite backend with automatic engine construction,
schema DDL, and optional pgvector loading. The returned backend owns the PGlite
engine; call `backend.close()` when the process or test is done.
```typescript
async function createLocalPgliteBackend(options?: {
/**
* PGlite data directory. Omit for an in-memory database, pass a filesystem
* path for persistence, or use a runtime-specific scheme such as `idb://`.
*/
dataDir?: string;
tables?: PostgresTables;
/**
* Omit to load @electric-sql/pglite-pgvector, pass `false` to disable vector
* support, or pass a PGlite Extension object to control the extension import.
*/
vector?: false | Extension;
/**
* Override the fulltext strategy. Defaults to `tsvectorStrategy`. Pass
* `false` to disable fulltext support entirely — the backend then
* advertises no `capabilities.fulltext` and omits the fulltext CRUD/search
* methods, and the installation DDL never creates the fulltext table.
*/
fulltext?: FulltextStrategy | false;
}): Promise<{
backend: GraphBackend;
db: PgliteDatabase;
client: PGlite;
}>;
```
#### `generatePostgresMigrationSQL()`
Returns complete fresh-installation SQL for creating TypeGraph tables and
stamping the current deployment-wide base-schema marker in PostgreSQL. It
includes the pgvector extension. The vector-disabled local PGlite factory uses
the same installation builder internally while omitting only that extension.
```typescript
function generatePostgresMigrationSQL(
tables?: PostgresTables,
fulltextStrategy?: FulltextStrategy | false,
): string;
```
#### `generatePostgresDDL(tables?)`
Returns individual DDL statements (CREATE TABLE, CREATE INDEX) as an array. Useful when you
need per-statement control, for example to execute them in separate transactions or log them
individually. This low-level array deliberately omits the deployment-wide
marker row, so joining it does not produce a complete installation. Use
`generatePostgresMigrationSQL()` for a database that will be opened through
`createVerifiedStore()` or the DML-only graph-template APIs.
```typescript
function generatePostgresDDL(
tables?: PostgresTables,
fulltextStrategy?: FulltextStrategy | false,
): string[];
```
#### `generatePostgresDropSQL(tables?, fulltextStrategy?)`
Returns one `DROP TABLE IF EXISTS` statement for the base and fulltext tables
that `generatePostgresDDL()` would create. Use it to clean up an isolated,
prefixed PostgreSQL table set after closing every backend connected to it.
Pass the same tables and fulltext strategy used at installation. The statement
does not use `CASCADE`: PostgreSQL refuses the drop if an application-owned
object depends on one of these tables. It does not drop graph-scoped vector
tables materialized later at runtime, so a working copy using those tables
needs additional graph-scoped cleanup.
```typescript
function generatePostgresDropSQL(
tables?: PostgresTables,
fulltextStrategy?: FulltextStrategy | false,
): string;
```
### Upgrading deployment-wide base storage
Skip this section when `createStoreWithSchema()` or
`createAdapterStoreWithSchema()` owns schema preparation: the bundled SQLite
and PostgreSQL adapters adopt each numbered base-schema release automatically
on the first privileged open. No separate bootstrap command is needed. The
deployment invariant is ordering: that privileged open must finish before any
DML-only runtime worker starts. Base-schema version 1 includes the durable graph
template relation and edge match-identity storage. It is required even for
graphs without a `matchIdentity` declaration because every edge write names the
two nullable columns.
When database DDL is managed externally, apply the matching migration before a
runtime worker opens the new graph schema. Apply the marker write last: it is
the durable proof that every preceding step succeeded. The examples use the
default TypeGraph table names. Replace every occurrence consistently when the
adapter uses custom table names.
For SQLite, run this migration exactly once. SQLite has no portable `ADD COLUMN
IF NOT EXISTS`, so a migration tool must record whether it has already applied
the two `ALTER TABLE` statements. Fresh and published schemas include the
nullable-pair `CHECK` below. Privileged adoption accepts an externally managed
table that already has both columns without that defensive constraint: SQLite
does not expose structural CHECK metadata or support adding one without a full
table rebuild, while TypeGraph writes always bind both values or neither.
```sql
CREATE TABLE IF NOT EXISTS "typegraph_graph_templates" (
"template_id" TEXT PRIMARY KEY NOT NULL,
"schema_hash" TEXT NOT NULL,
"schema_doc" TEXT NOT NULL,
"created_at" TEXT NOT NULL
);
ALTER TABLE "typegraph_edges"
ADD COLUMN "match_identity_name" TEXT;
ALTER TABLE "typegraph_edges"
ADD COLUMN "match_identity_key" TEXT
CHECK (("match_identity_name" IS NULL) = ("match_identity_key" IS NULL));
CREATE UNIQUE INDEX IF NOT EXISTS "typegraph_edges_match_identity_uq"
ON "typegraph_edges" (
"graph_id", "kind", "match_identity_name", "match_identity_key"
);
CREATE TABLE IF NOT EXISTS "typegraph_base_schema_versions" (
"installation" INTEGER PRIMARY KEY NOT NULL,
"version" INTEGER NOT NULL,
"updated_at" TEXT NOT NULL,
CONSTRAINT "typegraph_base_schema_versions_singleton_check"
CHECK ("installation" = 1)
);
INSERT INTO "typegraph_base_schema_versions"
("installation", "version", "updated_at")
VALUES (1, 1, CURRENT_TIMESTAMP)
ON CONFLICT ("installation") DO UPDATE SET
"version" = excluded."version",
"updated_at" = excluded."updated_at"
WHERE "typegraph_base_schema_versions"."version" <= excluded."version";
```
For PostgreSQL, the adoption statements are idempotent:
```sql
CREATE TABLE IF NOT EXISTS "typegraph_graph_templates" (
"template_id" TEXT PRIMARY KEY NOT NULL,
"schema_hash" TEXT NOT NULL,
"schema_doc" JSONB NOT NULL,
"created_at" TIMESTAMPTZ NOT NULL
);
ALTER TABLE "typegraph_edges"
ADD COLUMN IF NOT EXISTS "match_identity_name" TEXT;
ALTER TABLE "typegraph_edges"
ADD COLUMN IF NOT EXISTS "match_identity_key" TEXT;
DO $$
BEGIN
IF NOT EXISTS (
SELECT 1
FROM pg_constraint
WHERE conrelid = to_regclass('"typegraph_edges"')
AND conname = 'typegraph_edges_match_identity_pair_check'
) THEN
ALTER TABLE "typegraph_edges"
ADD CONSTRAINT "typegraph_edges_match_identity_pair_check"
CHECK (
("match_identity_name" IS NULL) = ("match_identity_key" IS NULL)
);
END IF;
END $$;
CREATE UNIQUE INDEX IF NOT EXISTS "typegraph_edges_match_identity_uq"
ON "typegraph_edges" (
"graph_id", "kind", "match_identity_name", "match_identity_key"
);
CREATE TABLE IF NOT EXISTS "typegraph_base_schema_versions" (
"installation" INTEGER PRIMARY KEY NOT NULL,
"version" INTEGER NOT NULL,
"updated_at" TIMESTAMPTZ NOT NULL,
CONSTRAINT "typegraph_base_schema_versions_singleton_check"
CHECK ("installation" = 1)
);
INSERT INTO "typegraph_base_schema_versions"
("installation", "version", "updated_at")
VALUES (1, 1, NOW())
ON CONFLICT ("installation") DO UPDATE SET
"version" = excluded."version",
"updated_at" = excluded."updated_at"
WHERE "typegraph_base_schema_versions"."version" <= excluded."version";
```
The conditional update makes marker publication monotonic: replaying an older
migration can never claim that storage prepared by a newer TypeGraph release is
older. The fresh-installation generators use `DO NOTHING` instead because they
are not upgrade planners; an existing stale marker remains stale until the
numbered privileged adoption lifecycle runs.
`createVerifiedStore`, `assertSchemaCurrent`, and the DML-only graph-template
APIs read this marker and throw `BaseSchemaMigrationError` when it is missing,
stale, or newer than the running library. They never attempt repair. A plain
`createStore` remains a synchronous zero-I/O attach; if it reaches an edge
write on legacy storage, the write fails with `ConfigurationError` and
`details.code === "EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE"` rather than a raw
missing-column error.
Provisioning the columns does not authorize re-keying existing data. Adding,
removing, renaming, or changing the fields of a declared `matchIdentity` remains
a breaking graph-schema change while that edge kind has any physical rows,
including tombstones. Export the affected edges, hard-delete them, publish the
new schema, and import them again so every row receives a key under the new
declaration.
### Base-schema version 5: byte-ordered `graph_id` indexes (PostgreSQL)
Version 5 adds one index to each relation `listGraphIds` seeks (`nodes`, `edges` and
`schema_versions`), ordering `graph_id` by bytes instead of by the database collation so that a page's
cursor, prefix and limit bound the walk. SQLite already keeps text indexes in byte order, so its step
only advances the marker. The index is a single-column `graph_id` index: PostgreSQL deduplicates the
repeated values, so it stays small (about 7 MB beside a 97 MB `nodes` heap of one million rows) and
adds 1 to 2% to writes on the relation it lands on (single creates and 1,000-row bulk writes alike).
The privileged open builds the three indexes with a plain `CREATE INDEX`, which blocks writes to the
table while it runs (about 0.1 second per million `nodes` rows on the measurement hardware). This
happens inline at boot even when `systemIndexes: "skip"` is set: that option only defers system index
materialization, not base-schema adoption. For a large deployment, build the indexes first with
`CONCURRENTLY`; the adoption step is `IF NOT EXISTS` and then finds them in place. Run each statement
outside a transaction, and never run the same concurrent build from two sessions at once. Use the
adapter's table names throughout:
```sql
CREATE INDEX CONCURRENTLY IF NOT EXISTS "typegraph_nodes_graph_id_bytes_idx"
ON "typegraph_nodes" ("graph_id" COLLATE "C");
CREATE INDEX CONCURRENTLY IF NOT EXISTS "typegraph_edges_graph_id_bytes_idx"
ON "typegraph_edges" ("graph_id" COLLATE "C");
CREATE INDEX CONCURRENTLY IF NOT EXISTS "typegraph_schema_versions_graph_id_bytes_idx"
ON "typegraph_schema_versions" ("graph_id" COLLATE "C");
```
`CREATE INDEX CONCURRENTLY` can leave an invalid index behind if it is interrupted; drop it and rerun.
`listGraphIds` checks only that each index exists and is valid, not its definition, so an index you
create by hand under one of these names with a different definition is trusted and makes the walk slow
rather than wrong. Create them exactly as shown.
Advancing the marker to 5 is a one-way step: a library release that predates version 5 refuses a
database stamped 5, so roll forward rather than back once any process has adopted it.
Externally managed DDL applies the same statements, then advances the marker to 5 with the
monotonic `INSERT ... ON CONFLICT` shown above. Until the indexes exist `listGraphIds` still returns
correct pages, by reading and de-duplicating the anchor relations instead of walking them.
## Drizzle-Free Entrypoints
TypeGraph keeps its public core and backend contracts independent of Drizzle:
- `@nicia-ai/typegraph/core` exports graph definition helpers and their
schema-derived types for packages that only define or share schemas.
- `@nicia-ai/typegraph/backend` exports the complete backend, dialect,
SQL-fragment, fulltext, and vector strategy contracts for adapter authors.
- `@nicia-ai/typegraph/sqlite/local` and
`@nicia-ai/typegraph/postgres/pglite` create managed Stores without exposing
adapter-native handles.
Application code can continue importing the complete portable Store API from
`@nicia-ai/typegraph`. Use the `/adapters/drizzle/...` entrypoints only when the
application deliberately owns a Drizzle connection or needs native transaction
interop.
Custom insert builders must apply the same born-ended validity rule as the
built-in adapters. Import its public owner instead of duplicating the bound
comparison:
```typescript
import { resolveStampedValidityLowerBound } from "@nicia-ai/typegraph/backend";
const validFrom = resolveStampedValidityLowerBound(
params.validFrom,
params.validTo,
writeInstant,
);
```
Use the same `writeInstant` for the decision and the row's creation/update
stamp. This keeps custom node and edge inserts, plus node resurrection paths
that reset the validity window, aligned with Store and interchange semantics at
the zero-width boundary. Edge resurrection retains its stored lower bound and
does not use this stamping helper.
## Managed Store Entrypoints
For local applications that do not need direct database access, TypeGraph can
own the connection, provision its schema, and return the complete typed Store:
- `@nicia-ai/typegraph/sqlite/local` — Node-only SQLite through the
native better-sqlite3 addon
- `@nicia-ai/typegraph/postgres/pglite` — in-process PostgreSQL through
PGlite's WebAssembly runtime
```typescript
import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local";
import { createLocalPgliteStore } from "@nicia-ai/typegraph/postgres/pglite";
const sqliteStore = await createLocalSqliteStore(graph, { path: "./graph.db" });
const postgresStore = await createLocalPgliteStore(graph, { vector: false });
```
These entrypoints expose no adapter-native database handle. The returned
`Store` keeps the complete graph API, including graph-owned
`store.transaction(...)`, but intentionally omits `withTransaction` and
`withRecordedTransaction`, which require a caller-owned adapter handle. The
Store owns its connection, so call `store.close()` during shutdown. Its
declaration surface is safe for strict TypeScript consumers that do not install
unused database drivers.
PGlite vector support is enabled by default and loads the optional
`@electric-sql/pglite-pgvector` package. Install that package when using vector
fields, or pass `{ vector: false }` as above for a smaller non-vector setup.
Both managed entrypoints also accept `fulltext: false`, which skips the
fulltext table at bootstrap and returns a backend with no
`capabilities.fulltext`.
Both factories accept `store` and `schemaManagement` groups, so the managed
path supports the same hooks, history/revision tracking, custom SQL schema,
query defaults, and migration policy as `createStoreWithSchema`:
```typescript
import { createSqlSchema } from "@nicia-ai/typegraph";
const store = await createLocalSqliteStore(graph, {
path: "./graph.db",
pragmas: { busyTimeoutMs: 10_000 },
store: {
history: true,
schema: createSqlSchema({
nodes: "app_nodes",
edges: "app_edges",
fulltext: "app_fulltext",
uniques: "app_uniques",
}),
},
schemaManagement: { systemIndexes: "skip" },
});
```
When a custom SQL schema is supplied, the managed factory provisions those
same physical table names; no separate Drizzle table configuration is needed.
`drizzle-orm` is an optional peer dependency for these two managed
entrypoints: they load it only when their factory is called and, when it is
absent, reject with a typed `ConfigurationError` (`MISSING_PEER_DEPENDENCY`)
naming the package and the install command (`npm install drizzle-orm`) rather
than a bare module-resolution stack. The explicit `/adapters/drizzle/...`
entrypoints below expose Drizzle-native backends, connections, or schema
builders — or, for `/adapters/drizzle/engine`, the factory that assembles a
backend from a caller-supplied engine profile — and load `drizzle-orm` when
the module is evaluated. Importing one without the peer installed therefore
surfaces the raw module-resolution error, which names the same package.
## Drizzle Adapter Entrypoints
TypeGraph exposes Drizzle adapters through public entrypoints:
- `@nicia-ai/typegraph/adapters/drizzle/indexes` — Drizzle schema-builder helpers for TypeGraph index declarations
- `@nicia-ai/typegraph/adapters/drizzle/sqlite` — Generic SQLite adapter (any Drizzle SQLite driver)
- `@nicia-ai/typegraph/adapters/drizzle/sqlite/local` — Batteries-included better-sqlite3 wrapper (Node.js only)
- `@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql` — Batteries-included libsql wrapper (Node.js, Workers, browser)
- `@nicia-ai/typegraph/adapters/drizzle/postgres` — PostgreSQL adapter (any Drizzle Postgres driver)
- `@nicia-ai/typegraph/adapters/drizzle/postgres/working-copy` — PostgreSQL table-backed working-copy manager
- `@nicia-ai/typegraph/adapters/drizzle/postgres/pglite` — Batteries-included PGlite (Postgres-in-WASM) wrapper
- `@nicia-ai/typegraph/adapters/drizzle/engine` — `createSqlBackend`, `deriveEngineProfile`, the bundled builders, `SqlEngineProfile`
Import from the entrypoint matching your database:
```typescript
import { createSqliteBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql";
import { createPostgresBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createLocalPgliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres/pglite";
```
### Engine profiles
`createPostgresBackend` and `createSqliteBackend` are each `createSqlBackend`
applied to a profile built by `buildPostgresEngineProfile` /
`buildSqliteEngineProfile`, both exported alongside `createSqlBackend` and
`deriveEngineProfile` from the engine entrypoint:
```typescript
import {
buildPostgresEngineProfile,
createSqlBackend,
} from "@nicia-ai/typegraph/adapters/drizzle/engine";
const backend = createSqlBackend(buildPostgresEngineProfile(db, options));
```
Most callers adapting a bundled backend want `deriveEngineProfile`, which
builds a variant of a bundled profile — a different lock spelling, a looser
declared capability, a replaced resource-audit verdict — without hand-copying
every other field. A profile written from scratch is not constructible today:
the assembly constructor is unexported, and `createSqlBackend` refuses a
hand-built assembly. See [Authoring an engine profile](/backend-authoring) for
the derivable-field table, the refusals a custom profile can hit, a worked
example, and [what is not derivable yet](/backend-authoring#what-is-not-derivable-yet).
A profile owns everything that genuinely differs between engines: dialect
tokens, the execution adapter, transaction framing, its `fenceSql` lock
spelling (see [Write fence declaration](#write-fence-declaration-writefence)),
provisioning DDL, strategies, and limits. `createSqlBackend` owns everything that is the
same for every SQL engine: deriving the final capabilities, resolving the
write-fence decision once, assembling the mirrored member groups, auditing
the backend's resource shape, and applying the trust marks. A backend minted
this way earns the marks its own declarations back: the schema-fenced-insert
mark only when the resolved fence plan actually fences writers, the
root-autocommit mark only when the profile declares single-statement
durability, and the atomic-program registrations only when its capabilities
support root atomic batching — whether the profile is for a PostgreSQL-wire
engine with a different locking story or an embedded engine with a different
transaction model.
`createSqlBackend` refuses a profile whose resolved capabilities omit
`writeFence` — every mark and registration it applies assumes a
resolvable write-fence decision, and a profile that does not declare one
cannot back that decision soundly. `buildPostgresEngineProfile` and
`buildSqliteEngineProfile` are the reference profiles to read when modeling a
new one.
A profile's `provisioning.catalog` supplies the backend's optional `catalog`
member: physical-schema introspection — table and index existence, each
column's normalized type family and raw declared type (a `CatalogColumn` is
`{ name, kind, declaredType }`; `declaredType` is required, and every custom
`columnTypes` implementation must populate it), and this engine's
index-build facts — for the handful of store paths that need to read the
engine catalog directly instead of compiling a portable query: index
materialization (`store.materializeIndexes()` refuses only once its
empty-candidate short circuit and the status-table ensure step have already
run; `store.materializeSystemIndexes()`, which has no candidate short
circuit, refuses only once that same status-table ensure step has run), the
recorded-time schema check, and the recorded-time migration's column read. A
profile that leaves `catalog` unset builds a backend with no `catalog`
member at all; those paths then refuse with a `ConfigurationError` naming
`catalog` rather than guessing at engine-specific SQL.
A dialect also declares `subgraphMembershipStrategy`, naming a decision the
dialect adapter makes, not one a profile supplies directly — the dialect
adapters are a fixed record keyed by `SqlDialect`, and each adapter's
capabilities (`DialectCapabilities`) declares `subgraphMembershipStrategy`, so a
profile inherits whichever of the two dialects its own `dialect` field names. It
is the plan-shape choice behind `store.subgraph()`'s reachable-node filter:
`"materialized-ids"` fetches the traversal closure once and filters both the
node and edge queries against that fixed id list (the shape PostgreSQL uses,
trading one extra round trip for a parameter-driven plan), while `"inline-cte"`
embeds the recursive closure in each fetch instead (the shape SQLite uses, where
an in-process traversal is cheap and a growing parameter list would pressure the
bind budget). This is a control-flow and prepared-plan decision, not SQL text a
shared token could express identically on both shapes, so it lives on
`DialectCapabilities` rather than in the query compiler.
`instantiateStatement` — a member of the profile's `graphTemplateRuntime` bag, and so one of the
fields `deriveEngineProfile` can override — is a required builder cloning a durable schema template
into a fresh graph. Given the template and target graph's ids and schema hashes
(`InstantiateGraphTemplateSqlParams`: `templateId`, `templateSchemaHash`, `graphId`, `schemaHash`,
and the three physical table names it reads), it must return the statement that inserts the target
graph's `schema_versions` row from the template's stored document and copies the template's
contribution-marker rows into the target graph — taking the target graph's write lock, the same key
the schema-commit fence takes, co-atomically with the insert on an engine that fences with locks. An
engine whose dialect can compose a data-modifying CTE beside the schema INSERT (PostgreSQL) folds
the marker copy and the lock into that one statement; an engine that cannot (SQLite) instead
supplies the optional `copyContributionMarkers` dep, which runs the marker copy as a second
statement once the schema row is confirmed. The bundled
`postgresInstantiateGraphTemplateStatement` and `sqliteInstantiateGraphTemplateStatement` builders
(`graph-template-sql.ts`) are what `createPostgresBackend` and `createSqliteBackend` supply to
their own profiles; neither is exported, so a from-scratch profile reaches the same shape only by
copying a bundled profile and adapting its statement, while a derived profile can replace the whole
`graphTemplateRuntime` bag through `deriveEngineProfile`.
`FenceSql` (see [Write fence declaration](#write-fence-declaration-writefence))
declares `advisoryLockExpression` and `isolationFactExpression` as the two
composable, no-`SELECT` forms an `advisory`-mechanism backend author
supplies; TypeGraph derives the standalone-statement counterparts
(`acquireKeyed`, `acquireKeyedWithIsolation`, `isolationFact`) from them. A
statement that
must compose a lock or an isolation read INSIDE a larger query it builds
itself — a CTE, a data-modifying statement — embeds the bare expression
directly, rather than running the derived standalone form as its own
preceding statement. The schema write fence's fused schema + graph-write
statement (`postgres-schema-write-fence.ts`) is the one site that needs
this: it puts the expression in its own CTE's `SELECT ... AS "lock_token"`,
resolving the fence target's OWN `FenceSql` — the bundled `postgresFenceSql`
for a bundled backend, a derived profile's own override otherwise — so a
custom spelling backs this fused statement exactly as it backs every
ordinary lock site.
The graph-template instantiation statement's `locked AS (SELECT ...)` CTE is
a DIFFERENT case, not a `FenceSql` consumer at all: it composes the baked
single-argument `advisoryLockSingleExpression` directly.
The ONE lock form with no override point is `advisoryLockSingleExpression`,
the ONE-argument form on a bare key: PostgreSQL stores it in a lock space
distinct from every namespaced two-argument lock, and the schema-commit fence
and graph-template instantiation both take it on the same key so the two
mutually exclude. It is not a `FenceSql` member — both bundled builders bake it
in directly, and a custom profile has no way to replace it.
## Cloudflare D1
TypeGraph supports Cloudflare D1 for edge deployments, with some limitations.
Cloudflare D1 has no interactive transaction primitive, so it cannot commit
TypeGraph schema versions: `commitSchemaVersion` / `setActiveVersion` need to
hold one transaction across their compare-and-swap read and activating
write, and D1 has no session to hold it on. Apply the base DDL with Wrangler
/ drizzle-kit. `capabilities.execution.unitOfWork` reports `"batch"`. A
singleton node create, update, `upsertById`, or delete fuses on a kind with
no declared unique constraint (a create takes a generated or a
caller-supplied id) — except a node delete, which fuses even when the kind
DOES carry a declared unique constraint, because the atomic delete program
releases that claim in the same statement. A singleton edge create fuses
when the kind's cardinality is `"many"`, and edge update and delete fuse the
same way (`EdgeCollection` has no `upsertById`). So do
`bulkInsert`/`bulkCreate`/`bulkDelete`/`bulkReplaceById`/`bulkUpsertById`,
and a constrained write inside an atomic program's claim envelope. Each of
these asserts the active schema version inside the statements
`D1Database.batch()` runs together, so these succeed on a schema-managed
Store. A write that cannot fuse either fails closed with
`BATCH_WRITE_UNSUPPORTED` naming a proven reason (a probe-then-write
constraint check, an interactive callback, Operational Identity, history, or
a schema commit), or — for a write that simply doesn't fit the fused shape,
such as a singleton create, update, or `upsertById` on a uniquely-constrained
kind, or a supplied-id tombstone resurrection — fails closed with the plain
`SCHEMA_WRITE_FENCE_UNSUPPORTED` limitation and no named reason; see
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares)
for the full reason table. Use a raw `createStore()` only when the
application accepts unfenced writes for the remaining paths:
```typescript
import { drizzle } from "drizzle-orm/d1";
import { createStore } from "@nicia-ai/typegraph";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
export default {
async fetch(request: Request, env: Env) {
const db = drizzle(env.DB);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
// Use store...
},
};
```
This raw Store does not validate or fence a committed TypeGraph schema version.
For schema commits and multi-statement schema-managed writes on Cloudflare, use
**Durable Objects** (below), whose SQLite storage exposes an interactive
transaction runner.
**Important:** D1 has no interactive transaction primitive
(`D1Database.batch(...)` is transactional, but batch-only — not an
interactive runner), so `store.transaction()` refuses on D1 before invoking
its callback. See
[Limitations](/limitations) for details. For a transactional Cloudflare
SQLite store, use **Durable Objects** (below) instead.
For the same reason, a write guarded by a **declared constraint** — edge
cardinality other than `many`, a `disjointWith` axiom, a shared-scope unique, or
dynamic `getOrCreateByEndpoints` convergence — is refused on D1 with
`CONSTRAINT_WRITE_FENCE_UNSUPPORTED` rather than committed unfenced. See
[Declared constraints require an interactive transaction](#declared-constraints-require-an-interactive-transaction).
## Cloudflare Durable Objects (SQLite)
A store backed by `drizzle(ctx.storage)` inside a Durable Object is
**auto-detected** as `transactionMode: "do-sqlite"` and reports
`capabilities.execution.interactiveTransactions: true` — no `executionProfile` hint needed.
Unlike D1, Durable Objects expose an interactive storage transaction runner,
so adapter stores can provide fully atomic `store.transaction()` and
`store.withTransaction()` operations.
The runtime authorizer forbids temporary tables, so the same profile reports
`capabilities.graphAnalytics.supported: false`. Traversal algorithms such as
`shortestPath`, `reachable`, and `weightedShortestPath` automatically use their
inline fallback; temporary-table-only analytics such as
`weaklyConnectedComponents` throw `UnsupportedBackendCapabilityError`.
The authorizer also rejects SQLite's `analysis_limit` tuning PRAGMA. Statistics
refresh catches that specific authorization error and still runs scoped
`ANALYZE`; this affects refresh cost only, not query results.
The same profile advertises Cloudflare's 100-bound-parameter query limit.
TypeGraph uses that hard ceiling for its managed write batches and
recorded-history flushes; capability overrides may lower it but cannot raise
it. Literal `.in()` and `.notIn()` query lists are packed into one JSON-bound
parameter, so the list itself does not exhaust the Durable Object budget.
```typescript
import { drizzle } from "drizzle-orm/durable-sqlite";
import { createAdapterStoreWithSchema } from "@nicia-ai/typegraph";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
export class MyObject {
constructor(private ctx: DurableObjectState) {}
async handle() {
const db = drizzle(this.ctx.storage);
const backend = createSqliteBackend(db);
// Boots schema/DDL outside any storage transaction (no DDL in the
// business transaction); the schema-version commit uses the
// do-sqlite runner.
const [store] = await createAdapterStoreWithSchema(graph, backend);
// Atomic across TypeGraph + the product's own relational tables:
await store.transaction(async (tx) => {
await tx.nodes.Document.update(documentId, props);
if (tx.sqlAvailability !== "available") {
throw new Error(`Native transaction unavailable: ${tx.sqlAvailability}`);
}
const sqlTx = tx.sql;
await sqlTx.insert(documentVersions).values(versionRow);
});
}
}
```
TypeGraph delegates to the async storage runner
`ctx.storage.transaction(async …)` (Drizzle's own `db.transaction()` on
Durable Objects is `ctx.storage.transactionSync` and cannot span an
`await`, so it is not used). See the
[Cross-Store Transactions recipe](/recipes#cross-store-transactions-drizzle--typegraph)
for the caller-owned (`withTransaction`) and graph-owned (`tx.sql`) shapes.
## Backend Capabilities
Check what features a backend supports:
```typescript
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
if (store.capabilities.execution.interactiveTransactions) {
await store.transaction(async (tx) => {
/* ... */
});
} else {
// Handle non-transactional execution
}
if (store.capabilities.vector?.supported) {
// Vector similarity queries available
}
```
`store.capabilities` is the portable runtime source of truth; adapter authors
can inspect the same object as `backend.capabilities`. The shape is:
| Field | Meaning |
| -------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| `execution` | Execution boundaries: `interactiveTransactions`, exact-resource `atomicBatch` support, and derived `unitOfWork` |
| `windowFunctions` | SQL window functions such as `ROW_NUMBER()` are available |
| `orderedAggregates?` | Ordered scalar and record collection aggregation; absent means unsupported |
| `constraintClaims?` | The backend carries the claim relations that fence declared constraints without a lock (see below) |
| `durableEdgeMatchIdentity?` | Edge writes persist and atomically arbitrate a schema-declared endpoint/property identity |
| `graphAnalytics?.{supported,mathFunctions}` | Static support for whole-graph temporary-table iteration, plus availability of deferred transcendental-math algorithms |
| `vector?.metrics` / `vector?.indexTypes` / `vector?.maxDimensions` | Vector strategy capabilities (present once a vector strategy is configured) |
| `fulltext?.{supported,languages,phraseQueries,prefixQueries,highlighting}` | Fulltext strategy capabilities |
| `recursiveTraversal?.{supported,reason}` | Whether the engine can compute a bounded transitive closure of a relation in one round trip — a recursive CTE, or a graph-native expansion operator. **Absent means supported** |
| `writeFence?.{mechanism,drain}` | How this engine excludes concurrent writers, and how far a caller can drain a table lock — see [Write fence declaration](#write-fence-declaration-writefence) |
`expr.collect()` requires `orderedAggregates: true`. Bundled PostgreSQL supports it. Supported preparable synchronous
SQLite clients and the dedicated async `createLibsqlBackend()` factory are probed when the backend
is created; the query itself adds no discovery statement. Other unprobed SQLite
connections default to unsupported. If you have
verified that your engine supports aggregate-local ordering, declare
`capabilities: { orderedAggregates: true }` in the bundled backend options. Older or unsupported
engines must retain `false`; collection queries are refused before execution.
SQLite introduced `NULLS FIRST` / `NULLS LAST` ordering in version 3.30 and aggregate-local ordering
in [version 3.44](https://www.sqlite.org/releaselog/3_44_0.html).
Collection compilation avoids depending on JSON object subtype preservation during sorting.
Existing reads continue to work when ordered aggregates are unavailable.
Custom dialect adapters implement `orderedScalarJsonArray` with one required object argument:
`{ value, valueType, orderBy, filter }`. Migrate positional implementations by destructuring that
object, applying `filter` as an aggregate `FILTER (WHERE ...)` before wrapping the aggregate in the
empty-input `COALESCE`, and leaving it off when `filter` is `undefined`. The aggregate must preserve
included NULL operands and return `[]` for empty input. The `filter` key itself is required in the
adapter contract, even though its value may be `undefined`, which requires old positional
implementations to migrate explicitly.
The dialect adapter contract also requires `orderedRecordJsonArray` for
`expr.collect({ field: scalarExpression }, options)`. Custom adapters must add this method when
upgrading. It builds an ordered JSON array of flat records with the named projected scalar fields,
honors the required ordering tuple and optional aggregate-local filter, preserves admitted SQL NULL
fields within each record, and returns `[]` for empty input. The same `orderedAggregates` capability
governs scalar and record collections; a custom adapter must supply both SQL emitters before
declaring that capability.
The former top-level `capabilities.transactions` override is not interpreted
as an alias. Bundled factories refuse it with `LEGACY_CAPABILITY_OVERRIDE`,
including for JavaScript and already-compiled callers, because transaction
availability and atomic batching are now independent facts. Move the value to
`capabilities.execution.interactiveTransactions`; root `atomicBatch` support
is discovered from the bundled transport and cannot be claimed through factory
overrides. Bundled PostgreSQL transaction factories may expose
`atomicBatch: "session"` on the exact already-open transaction object. That
declaration is paired with fresh transport and semantic registrations and is
never inherited by an ordinary derived backend.
`execution.unitOfWork` is derived, never declared by a factory or override:
`"optimistic-retry"` when `interactiveTransactions` is `true` AND the resolved
write fence is `{ mechanism: "row", conflict: "commit-time" }` (see
[Write fence declaration](#write-fence-declaration-writefence) below);
`"interactive"` when `interactiveTransactions` is `true` otherwise; else
`"batch"` when `atomicBatch` is not `"none"` (an HTTP-only driver such as
`drizzle-orm/neon-http`, which cannot hold an open session but does support a
native atomic program); else `"none"`. Only the root capability derivation
ever resolves the write-fence plan needed for the `"optimistic-retry"` arm —
a derived or session-scoped capabilities object (a `store.transaction`
session, a projected backend) has no way to re-resolve that plan for
itself, but it carries the root's answer forward instead of losing it: it
reads whether its own source object was already `"optimistic-retry"` and
keeps the tier for as long as `interactiveTransactions` stays `true`,
falling back to `"interactive"` only for a capabilities object whose source
never carried the tier to begin with.
Two further internal readers key off the `"batch"` value: the batch-tier
write verdict (`resolveBatchWriteVerdict`) that produces
`BATCH_WRITE_UNSUPPORTED` refusals, and the autocommit single-statement
eligibility gate that decides whether a supplied-id singleton create can
fuse its schema fence into one statement. Even absent `"optimistic-retry"`,
`unitOfWork` exists so any caller can tell the execution shapes apart without
re-deriving the same distinction from `interactiveTransactions` and
`atomicBatch` separately.
Under `"optimistic-retry"`, every TypeGraph-owned transaction that acquires a
fence row replays a real commit-time conflict as a whole unit, up to
`OPTIMISTIC_RETRY_ATTEMPTS` (3) attempts, and only exhausting that budget (or
a non-retryable failure) surfaces `TransactionConflictError` to the caller —
see [Retrying on conflict](/schemas-stores#retrying-on-conflict). That covers
every store-owned write (collection create/update/delete, bulk paths,
`importGraph`, identity maintenance, contribution rebuild, index
materialization) as well as the two backend-owned transactions that acquire
the schema-commit fence row directly, outside the store's own write path:
graph-template instantiation and a schema commit (`commitSchemaVersion` and
its three siblings, via `runSchemaWriteTransaction`). A nested write running
inside an existing transaction (`store.transaction`, an adopted transaction)
never retries on its own: it cannot restart a transaction it does not own, so
its conflict propagates unchanged to the outermost store-owned write, or to
`store.transaction` itself. This tier therefore changes behavior only for a
transaction that opens its own top-level connection.
An `"optimistic-retry"` backend requires `node:async_hooks`' `AsyncLocalStorage`
to detect a retried unit nested inside another one; on a runtime where it is
unavailable, the first retried unit `runRetriedUnit` opens is refused with
`OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` rather than degrading to independent,
unsafe per-unit retries, while interactive backends are unaffected and keep
retrying (`store.transaction`'s own `retry` option) with no async-context
support at all.
`graphAnalytics.supported` describes the backend shape, not mutable PostgreSQL
session state. A hot standby or a role without the database `TEMP` privilege can
still reject the working-table transaction that the iterative graph algorithms
open: a standby refuses the read-write transaction itself, and a role without
`TEMP` refuses the `CREATE TEMP TABLE` inside it. Both refusals reach the caller
as `UnsupportedBackendCapabilityError`, with the PostgreSQL error retained as
its `cause`.
### Durable edge match identity capability
`capabilities.durableEdgeMatchIdentity: true` is a correctness promise. A
custom backend making it must provide all of these guarantees:
- Every edge write carrying `InsertEdgeParams.matchIdentity` stores both the
name and key with the row. They are either both absent or both present.
- A database constraint atomically owns uniqueness over `(graph_id, kind,
match_identity_name, match_identity_key)`. Soft deletion keeps the key;
physical hard deletion releases it.
- `commands.execute()` handles a durable `edge.converge-create` as one database
decision and returns the authoritative `created` or `found` row. Returning
`unsupported` fails closed with
`DURABLE_EDGE_MATCH_IDENTITY_COMMAND_UNSUPPORTED`; TypeGraph does not fall
back to a read-then-write race.
- Storage exists before runtime writes. Implement
`ensureEdgeMatchIdentityStorage` for privileged schema adoption, or provision
the columns, pair constraint, and unique arbiter independently before setting
the capability.
`insertEdgesDurableBatchReturning` is an optional throughput member. When
implemented, every input must carry a durable identity, conflicts are omitted
from the returned rows, and returned rows identify exactly which inputs were
created. Omitting it preserves correctness through per-row authoritative
commands, but loses the set-oriented bulk/import fast path.
`findEdgesByHeterogeneousEndpointSet` is likewise an optional set-read
optimization. An input carrying `opposite` requests an exact directed endpoint
pair, not every edge incident to the first endpoint. One call may contain only
incident inputs or only exact-pair inputs; mixing the two modes is refused. A
backend that omits the member retains the exact per-pair fallback.
### Validity-end clearing capability
Custom backends must advertise `capabilities.clearValidTo: true` only when both
`updateNode` and `updateEdge` apply `clearValidTo: true` by storing SQL `NULL` in
`valid_to`. The built-in SQLite and PostgreSQL adapters do. An explicit clear on
a backend without that promise is refused with `ConfigurationError` code
`CLEAR_VALID_TO_UNSUPPORTED` before coalescing or writes, so the result
does not depend on whether the target row is already open. Omission still means
preserve; custom backends that do not support clearing remain compatible with
all writes that omit the option.
### Recorded-table migration DDL (`recordedTableDdl`)
`GraphBackend.recordedTableDdl` is an optional, synchronous callback used only by the offline
timestamp-only recorded-time preview migration. The migration calls it twice, once with temporary
table names and once with the final names, and expects DDL for `recordedClock`, `recordedNodes`, and
`recordedEdges`. The backend owns this callback because table creation, indexes, and named
constraints are dialect-specific and must not pull Drizzle into portable entrypoints.
A custom backend can omit the callback unless it created data in the old preview schema. If
`migrateLegacyRecordedTime` discovers that schema and the callback is absent, it throws
`UnsupportedBackendCapabilityError` with `details.capability: "recordedTableDdl"`. When the engine
names primary-key constraints, the temporary and final callback results must either both name the
constraint or both omit it; a one-sided result throws `ConfigurationError` code
`RECORDED_DDL_CONSTRAINT_NAME_MISMATCH`.
The callback only describes DDL. It must not execute statements or inspect the catalog, because the
migration invokes it inside its transaction. See
[Migrating Preview Recorded Time](/schema-management#migrating-preview-recorded-time) for the
operator workflow.
### Recursive traversal capability
Both bundled backends declare `capabilities.recursiveTraversal: { supported: true }`. **Absent
means supported** — mirroring `returning`, not `constraintClaims`: every existing custom backend
already runs the six recursive-CTE emission sites unconditionally, so absence meaning unsupported
would refuse traversals that work today.
```typescript
const capabilities: Partial = {
recursiveTraversal: { supported: false, reason: "engine has no WITH RECURSIVE / equivalent" },
};
```
A backend that genuinely lacks the primitive declares `{ supported: false, reason }`. A factory
refuses a contradictory declaration — `supported: false` with no `reason`, or `supported: true`
with a dangling `reason` — with `ConfigurationError` details code `CAPABILITY_DECLARATION_CONTRADICTION`.
Five operations refuse when unsupported: variable-length (`traverse`) queries, `store.subgraph()`,
historical identity class reads, identity-expanded historical queries, and the identity
window-ledger read — each throwing `ConfigurationError` code `RECURSIVE_TRAVERSAL_UNSUPPORTED`
with `details.operation` naming the site and `details.reason` echoing the declaration.
`weightedShortestPath` is the one exception: on a backend with temporary statements but no
recursion, it **falls back** to a per-hop predecessor walk instead of refusing, issuing
`pathLength + 1` extraction statements for the path a recursive CTE would have returned in one
round trip. The unweighted `shortestPath` (along with `reachable`, `canReach`, and `neighbors`)
emits no recursive CTE at all — it routes through the iterative working-table or inline path
instead — so it neither refuses nor falls back regardless of this declaration.
### Write fence declaration (writeFence)
TypeGraph serializes a family of writes — Operational Identity's mutations, and the
TypeGraph-owned recorded-clock allocation behind `history` / `revisionTracking` — behind a
per-graph fence rather than trusting the engine's default isolation. `capabilities.writeFence`
declares what this backend can provide, as two independent facts, and `resolveWriteFencePlan` is
the one place that declaration turns into a plan every lock site consumes instead of re-deriving:
```typescript
const capabilities: Partial = {
writeFence: { mechanism: "advisory", drain: "table-lock" },
};
```
`mechanism` is how the backend excludes concurrent writers. `writeFence` is a discriminated union on
`mechanism`, and `drain` is a field of the `"advisory"` and `"row"` shapes only — `"engine-serialized"`
and `"caller-serialized"` declarations carry no `drain` key at all:
| `mechanism` | Meaning |
| --- | --- |
| `"advisory"` | A keyed `pg_advisory_xact_lock`-style lock a caller takes explicitly. Needs `fenceSql` (below) and a `drain`. |
| `"row"` | A keyed exclusion spelled by TypeGraph itself against the never-dropped fences relation, for an engine with no advisory-lock primitive. Needs a `drain` and a `conflict` (below); a `fenceSql.isolationFactExpression` is optional (absent means recorded capture and match-key convergence fail closed on an unknown isolation fact, exactly as they do for a target that supplies neither). |
| `"engine-serialized"` | The engine serializes writers by construction — SQLite's single writer slot. No lock statement, no `fenceSql`, no `drain`. |
| `"caller-serialized"` | A deployment-level promise, not an engine fact — see below. No lock statement, no `drain`; a `fenceSql` the backend still carries is used only for its isolation-fact read (recorded capture's isolation guard). |
`drain` (on `mechanism: "advisory"` or `"row"` only) is a separate fact: whether a caller that
already excluded other writers can additionally take a relation-wide lock on the table a drain site
protects:
| `drain` | Meaning |
| --- | --- |
| `"table-lock"` | Yes — a `LOCK TABLE`-style statement is available and the drain site takes it. |
| `"quiescent"` | The resource is already exclusive for some other reason (an advisory or row lock layered under a deployment's own `caller-serialized` promise, for instance), so the drain site takes NO statement — one it does not need rather than one it cannot spell. |
| `"none"` | Neither — a drain site refuses, naming the drain. |
`conflict` (on `mechanism: "row"` only) is the engine fact for what happens when two writers
acquire the SAME fence row:
| `conflict` | Meaning |
| --- | --- |
| `"wait"` | A lock-based engine — the second acquirer's statement blocks until the first commits, exactly like an advisory lock. |
| `"commit-time"` | An optimistic-concurrency engine — both acquirers proceed and the loser's COMMIT fails. Correctness comes from the retry owner replaying the whole unit, never from waiting, so `conflict: "commit-time"` derives the `"optimistic-retry"` execution tier (see [Backend Capabilities](#backend-capabilities) above) and requires it: declaring it on a non-interactive backend is refused the same way an out-of-place `drain` is. |
`"engine-serialized"` and `"caller-serialized"` satisfy every drain site unconditionally — a writer
slot and an in-process serialization promise are each already a stronger exclusion than any `drain`
value could add, so attaching one to either mechanism is refused (see **Runtime validation** below)
rather than silently ignored; attaching `conflict` to anything but `"row"` is refused the same way.
`resolveWriteFencePlan` resolves one of five plans:
- `{ kind: "lock", drain, sql }` — take the declared advisory lock (`sql`, the target's own
spelling), and, when `drain === "table-lock"`, the table lock a drain site needs.
- `{ kind: "row", drain, conflict, sql }` — take the SAME `sql.acquireKeyed` /
`sql.acquireKeyedWithIsolation` a `"lock"` plan's site calls, spelled instead against the fences
relation; `conflict` is the one fact a `"row"` site (and the execution tier) reads that a
`"lock"` site never needs.
- `{ kind: "engine-serialized" }` — no lock needed; the engine serializes writers by construction.
- `{ kind: "caller-serialized" }` — no lock needed; the deployment's own promise excludes concurrent
writers (see below).
- `{ kind: "unfenced" }` — no declaration at all. Every fence that guards a read-then-write across
statements refuses rather than running unfenced. Only a predicate carried *inside* the statement
it guards degrades, since one statement cannot race itself.
Resolution order: (1) the declared `writeFence` value, if present; (2) absent, AND the backend was
built by `createSqliteBackend` / `createPostgresBackend` — derived from `dialect`, which is exactly
what every lock site used to compute inline (this derivation never resolves `"row"`: it is the two
bundled dialects' own `"advisory"`/`"engine-serialized"` split); (3) absent on anything else —
`unfenced`, because an undeclared custom backend is by definition uncertified and inferring lock
support from `dialect` alone is the unsound inference this capability replaces.
The two bundled backends resolve exactly these declarations — copy the one matching your engine:
- PostgreSQL: `writeFence: { mechanism: "advisory", drain: "table-lock" }`
- SQLite: `writeFence: { mechanism: "engine-serialized" }` (no `drain`: the writer slot already
excludes every drain site's writer, so a drain site under it always takes no statement — the same
behavior `drain: "quiescent"` describes for `"advisory"`/`"row"`, without a `drain` field to spell it)
A backend that declares `mechanism: "advisory"` also supplies `fenceSql`: `lockTables` (only needed
when `drain: "table-lock"`) plus the two composable, no-`SELECT` forms `advisoryLockExpression` /
`isolationFactExpression` a statement embeds inside a larger query it builds itself (see the
schema-write-fence discussion above) — the complete `FenceSql` bag. `resolveWriteFencePlan`'s `lock`
arm derives the standalone-statement forms every ordinary lock site actually calls — `acquireKeyed`;
`acquireKeyedWithIsolation` (the lock plus the session's isolation-level fact, read in the same
statement it locks in); and `isolationFact` — from those two expressions, so a backend author never
spells both forms separately. `mechanism: "row"` needs no `advisoryLockExpression` at all: TypeGraph
spells its own `acquireKeyed` / `acquireKeyedWithIsolation` against the fences relation (see below),
and a `fenceSql.isolationFactExpression` — when supplied — rides the SAME acquisition statement, so a
`"row"` target's isolation fact is read on the exact connection that took the row. The bundled
PostgreSQL spelling is exported as `postgresFenceSql` from `@nicia-ai/typegraph/adapters/drizzle/postgres`
— pass it straight through as `fenceSql` when wrapping that backend (under either `"advisory"` or
`"row"`), or supply a custom `FenceSql` matching a different engine's lock syntax. A backend that
declares `mechanism: "advisory"` with a `fenceSql` missing a member the resolved `mechanism`/`drain`
combination needs is refused at construction with details code `WRITE_FENCE_SQL_UNAVAILABLE`, naming
the missing member; `"row"` is refused the same way only for `drain: "table-lock"` without
`lockTables` — its acquisition statement needs no author-supplied spelling at all, so a missing
`tableNames.fences` instead refuses the first time a keyed site actually acquires the row, not at
construction; `"engine-serialized"` and `"caller-serialized"` need no `fenceSql` to take a lock at
all.
#### The fences relation
A `"row"`-mechanism backend needs one relation, `typegraph_fences(key TEXT PRIMARY KEY, generation
BIGINT NOT NULL)` (`INTEGER NOT NULL` on SQLite) — part of TypeGraph's base schema on both bundled
dialects, so a fresh install already has it and `generateSqliteMigrationSQL` /
`generatePostgresMigrationSQL` add it to an existing database. It is **never dropped, never cleared
by `clear()`, and never row-deleted** — the same durability contract `schema_versions` and
`recorded_clock` carry. Every acquisition is one portable statement TypeGraph spells itself, never
the profile: `INSERT INTO typegraph_fences (key, generation) VALUES (key, 1) ON CONFLICT (key) DO
UPDATE SET generation = generation + 1 RETURNING generation`, keyed on `${namespace}:${key}` —
the SAME advisory namespaces and per-position keys an `"advisory"`-mechanism backend locks on,
verbatim, so the lock-order contract carries over unchanged to an engine using the fences relation
instead of `pg_advisory_xact_lock`. A custom backend supplies the relation's physical name through
`tableNames.fences` (defaulted to `typegraph_fences` by both bundled factories) exactly as it names
every other TypeGraph-owned table.
#### Runtime validation
TypeScript's discriminated union only holds a caller who goes through the type checker — a plain
JavaScript backend author, or a value round-tripped through JSON or a config file, can still supply
an unrecognized `mechanism` string, an unrecognized `drain` or `conflict` string, a `drain` attached
to `"engine-serialized"` / `"caller-serialized"`, or a `conflict` attached to anything but `"row"`.
`resolveWriteFencePlan` validates every declaration — whether it came from `capabilities.writeFence`
directly or from the first-party dialect fallback — before shaping a plan from it, and refuses with
`ConfigurationError` details code `WRITE_FENCE_DECLARATION_INVALID`, naming the invalid `field`
(`"mechanism"`, `"drain"`, or `"conflict"`) and, for an unrecognized value, the `accepted` list. An
unrecognized `drain` never falls through to behaving like `"quiescent"` — it is refused outright,
the same as an unrecognized `mechanism`. The same validator refuses `conflict: "commit-time"`
outright when the target's own `capabilities.execution.interactiveTransactions` is `false`: that
value is honored only by the `"optimistic-retry"` execution tier, which never derives without an
interactive transaction to replay inside, so accepting the declaration there would silently drop it
rather than apply it.
#### `caller-serialized`: the promise split into two halves
`caller-serialized` is for a deployment that knows its database has no other concurrent writer, but
whose engine is neither an advisory-lock engine nor a single-writer one — a PostgreSQL-wire engine
with no working `pg_advisory_xact_lock` / `LOCK TABLE`, for example. The promise has two halves,
and TypeGraph only enforces the first:
- **In process**, TypeGraph enforces it: every root member the backend classifies as mutating —
collection writes, `store.transaction` / `transactionWithNative`, schema commits, identity and
contribution maintenance, index materialization, table/DDL provisioning, `clearGraph`, import,
and the raw-SQL members (`execute`, `executeRaw`, `executeStatement`,
`executeTemporaryStatement`) that can carry an arbitrary write — runs through one per-backend
serialized queue, so two concurrent calls through one pool cannot race each other. A root write
awaited from inside a `store.transaction` callback is refused rather than left to deadlock behind
the transaction's own queue slot.
- **Outside the process**, the deployment enforces it: no other client writes to this database
while this backend is open. TypeGraph cannot see or verify that half; declaring
`caller-serialized` is asserting it.
Adopting an externally owned transaction (`store.withTransaction(externalTx)`, backed by
`adoptTransaction`) is refused outright on a `caller-serialized` backend, with `ConfigurationError`
details code `CALLER_SERIALIZED_REFUSES_ADOPTION`: an adopted transaction's lifetime belongs to the
caller, not to this backend's write-unit queue, so there is no honest way to hold a queue slot open
for it — queuing it would block every other queued write until the caller's own transaction ends,
and leaving it unqueued would let its writes interleave with the queue's own, silently breaking the
promise `caller-serialized` makes. Open the transaction through this backend's own `transaction()` /
`transactionWithNative()` instead, or do not declare `caller-serialized` on a backend that needs
cross-store adoption.
`createPostgresBackend` accepts `writeFence: { mechanism: "caller-serialized" }` — a claim about the
deployment — while still refusing `mechanism: "engine-serialized"` outright, because that value is
a claim about the *engine*, which this factory's own engine does not back.
Constructing Operational Identity, or `history: true` / `revisionTracking: true`, against an
`unfenced` backend is refused immediately at `createStore` — never mid-flush — with
`ConfigurationError` details code `IDENTITY_REQUIRES_WRITE_FENCE` (identity) or
`RECORDED_CLOCK_REQUIRES_WRITE_FENCE` (recorded-clock allocation), and the refusal message names
the exact declaration line to add.
A `lock` plan whose `drain` cannot back a site's `requires: "drain"` — `drain: "none"` — is
refused with details code `WRITE_FENCE_UNAVAILABLE`, naming `details.operation` and the drain that
could not be satisfied. `"engine-serialized"` and `"caller-serialized"` satisfy either `requires`
value (`"keyed"` or `"drain"`) without consulting `drain`.
The PostgreSQL schema fence refuses too, and it is worth knowing why it is not on the
degradable side. The per-graph advisory lock plus `SELECT ... FOR UPDATE` a schema commit takes,
and the `FOR SHARE` a managed write takes on that same row, each fence a read-then-write sequence
that spans **statements**: `commitSchemaVersion` reads the active version and then writes the
flip, and a managed write holds its `FOR SHARE` for the remainder of the transaction so the
version it asserted stays true through the writes that follow. Skipping those locks would not
give a slower-but-correct path; it would assert a version and then let the very change the
assertion was checking for land before the write. So a PostgreSQL-dialect backend that resolves
`unfenced` is refused at the schema commit with `WRITE_FENCE_UNAVAILABLE`, naming the operation.
The one part that *does* degrade is the fence folded into a managed insert's own statement. That
predicate is evaluated inside the INSERT that depends on it, and one statement cannot race
itself: with no locking clause the fence subquery still yields no row when the expected version
is no longer active, so the INSERT still writes nothing. This is how SQLite has always run the
path, on the strength of its writer slot.
This matters for a PostgreSQL-wire engine that implements neither `pg_advisory_xact_lock` nor the
`FOR UPDATE` / `FOR SHARE` clauses, and whose engine merges concurrent transactions rather than
serializing them.
`writeFence` has no arm for "no exclusion mechanism at all" — every `mechanism` value claims
something real. An engine with neither locks nor a writer slot nor a caller promise to make has
nothing honest to declare, and `unfenced` is how that shows up downstream — but neither bundled
factory will hand you that backend. `createPostgresBackend` and `createSqliteBackend` both build on
`createSqlBackend`, which refuses at construction, with `ConfigurationError` details code
`ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION`, when the resolved capabilities carry no
`writeFence` at all — including `capabilities: { writeFence: undefined }` passed to either factory,
which no longer builds a backend the way it once did:
```typescript
createPostgresBackend(db, {
capabilities: { writeFence: undefined },
});
// throws ConfigurationError: ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION
```
`unfenced` is reachable only outside that gate: a hand-assembled `GraphBackend` that never goes
through `createSqlBackend`, or a custom `SqlEngineProfile` whose `declaredCapabilities.writeFence`
some other override clears before it reaches a call site — never through either bundled factory.
Getting there is a way of admitting the engine truly has nothing to declare; TypeGraph then refuses
a schema-managed store built on it at `createStore`, naming the missing capability, rather than
running a fence the engine cannot enforce.
If the deployment instead knows it is the only writer of this database — a pool clamped to one
connection, or a single-writer topology otherwise enforced outside TypeGraph — declare
`writeFence: { mechanism: "caller-serialized" }` instead (see above): that is
the honest way to spell a deployment convention. Do not reach for `mechanism: "engine-serialized"`
for the same purpose — that declaration means the *engine* serializes writers by construction, and
a deployment convention is not a construction. `createPostgresBackend` refuses that particular
claim outright for this reason.
### Capability bundles
A **capability bundle** groups a set of `GraphBackend` members that one operation family needs
together, with one verdict resolver and one member accessor, so a caller never re-derives "does
this backend support X" from a scattered `undefined` check. Seven pilot bundles ship in this
release:
| Bundle | Kind | Disposition |
| ------------------------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claims` | gated | Bidirectional cross-check between the `constraintClaims` declaration and the core members; disagreement in either direction refuses with `CONSTRAINT_CLAIM_SURFACE_MISMATCH` |
| `statementExecution` | gated | Core `executeStatement` absent refuses with `IDENTITY_REQUIRES_STATEMENT_EXECUTION` |
| `recordedRevisionOrigins` | gated | Core `ensureRevisionOriginsTable` absent refuses with the operation's own typed error |
| `batchPointRead` | graduated | `getNodes` absent falls back to per-id `getNode`; `getEdges` absent falls back to per-id `getEdge` |
| `uniqueSidecarBatch` | graduated | `insertUniqueBatch` absent falls back to `issueClaimsIndividually`; `checkUniqueBatch` absent falls back to a per-key loop; `hardDeleteUniquesByNodeIds` absent refuses with the operation's own typed error |
| `contributionHealth` | graduated | `verifyContributions` / `repairContributions` / `rebuildContribution` absent each refuse with the operation's own typed error; `probeContributions` absent falls back to `{ entries: [] }` |
| `endpointSetRead` | graduated | `findEdgesByEndpointSet` absent refuses set-oriented `bulkFindFrom` / `bulkFindTo` with `ENDPOINT_SET_READ_UNSUPPORTED`; singleton reads remain available |
The port-mismatch rule that governs every bundle's member accessor is keyed to the disposition,
not blanket: a `refuse`-disposition row whose backend object cannot actually reach the member
throws that bundle's own `portSurfaceCode` (`CONSTRAINT_CLAIM_SURFACE_MISMATCH` for `claims`,
`BUNDLE_PORT_SURFACE_MISMATCH` for the other six); a `fallback`-disposition row whose port cannot
reach the member takes its declared fallback instead of throwing — the verdict said the member
was there, the object it binds against says otherwise, and a fallback row is defined to degrade
rather than assert.
This bundle model ships for **seven of the twenty-one** member-bearing operation families measured
in this workstream; the remaining fourteen are a named follow-up workstream, not a silent gap —
their members keep working exactly as before, unbundled, with an access-count ceiling that
prevents new scattered checks from accumulating ahead of that follow-up.
A backend author does not need to do anything for these seven bundles today: both bundled backends
already carry every core member each bundle's `dialects` scope requires. The atomic transport
conformance runner is the foundation for certifying a **third-party** backend: the author supplies
engine-specific statements, state observers, and exact-root provenance checks, while the runner
asserts the shared transport contract. Bundle verdicts remain a separate check against the declared
capabilities and the object the calls actually execute on. A backend must therefore make every
declared capability (`constraintClaims`, `contributions`, and execution support) truthful about
what the active backend object implements, not just which fields it sets.
Run the conformance fixture in the custom backend's own test suite, then pair
the earned declaration with the exact root transport in its factory:
```typescript
import {
decorateBackend,
registerAtomicMutationPrograms,
registerAtomicSqlProgram,
runAtomicMutationProgramConformance,
runAtomicTransportConformance,
} from "@nicia-ai/typegraph/backend";
const backend = createCustomBackend({
execution: {
interactiveTransactions: false,
atomicBatch: "root",
},
});
registerAtomicSqlProgram(backend, { executeAtomicBatch });
const authorCreatedWrapper = decorateBackend(backend, {});
await runAtomicTransportConformance({
...transportCases,
backend,
derivedBackends: [authorCreatedWrapper],
executeAtomicBatch,
});
registerAtomicMutationPrograms(backend, mutationPrograms);
const semanticCases = buildSemanticCases({ backend });
await runAtomicMutationProgramConformance({
backend,
derivedBackends: [authorCreatedWrapper],
equal: deepEqual,
cases: semanticCases,
});
```
Transport registration is exact-resource evidence only: a derived backend does
not inherit it, and a second registration on the same object is refused rather
than replacing the function production uses. A bundled PostgreSQL transaction
session earns a separate registration bound to its pinned client; it does not
inherit the root's registration.
Create wrappers with the exported `decorateBackend()` seam so the runner can
verify their lineage back to the registered root instead of accepting an
unrelated object as derivation evidence. The conformance fixture's mandatory
provenance checks prove registration, lineage, derived isolation, and—when
applicable—transaction isolation against the real objects supplied by the
backend author. A non-interactive root reports the
transaction-isolation check as skipped rather than claiming evidence it could
not obtain. Transport registration certifies mechanics, not graph semantics, and
therefore does not by itself opt a custom backend into any Store mutation
program.
The separate `registerAtomicMutationPrograms()` call is the semantic boundary:
each member declares one complete TypeGraph mutation family implemented by that
exact backend resource. Omitted families retain the portable path, and an empty profile or
a profile registered before its atomic transport is refused with
`ConfigurationError`.
The semantic executors must preserve the same schema fence, validation,
side-effect, refusal classification, rollback, postimage correlation, result
ordering, and bind-ceiling contracts as the bundled implementation. Registering
one family is not evidence for another. Derived and projected backends inherit
neither registration. An exact transaction session must be registered
independently before Store code can dispatch through it.
`runAtomicMutationProgramConformance()` is the executable semantic boundary.
For every reachable positive-limit variant in `mutationPrograms`, the fixture supplies
three real Store-level cases:
1. an ordered success whose return value and independently read committed state
both match their oracles;
2. a stale-schema-fence refusal that leaves the database unchanged; and
3. a family-specific typed refusal that either rolls back every sibling write
after native dispatch or explicitly refuses before dispatch without writing.
The runner resolves the profile from `backend`; it does not accept a detached
profile description, caller-supplied provenance verdict, or fixture-owned
dispatch counter. Before any fixture preparation can write, it validates the
complete case inventory and probes the author's actual derived backend objects.
It observes dispatch inside the exact registered executors and therefore refuses
a success that silently used the portable fallback,
a case bound to a different family claim, a missing or duplicate family case,
and a case that claims an unregistered family. A zero entry limit is an honest
opt-out and does not require an unreachable case. `mutateEdges` has separate
`resolvedSet` and `durableConvergence` variants because proving one does not
prove the other.
Every semantic case identifies the exact `backend` its callbacks use. The
runner checks that binding and the registered profile identity before any
preparation, again between preparation and execution, and after execution, so
a pre-dispatch refusal or a mid-run registry replacement cannot borrow another
root's certificate. The fixture callbacks should invoke public Store methods and inspect committed
rows through an independent database read. Supply at least one real wrapper or
derived backend created with `decorateBackend()`; the runner does not manufacture
a projection and mistake that tautology for author evidence. Do not instrument
or replace the registered executors—the runner owns dispatch evidence. Run
conformance with exclusive use of that exact root: unrelated same-variant writes
during the observation window cannot be distinguished from fixture traffic.
Mark each semantic refusal's `dispatch` as `"required"` or `"pre-dispatch"`
according to the Store contract, and do not use executor return rows as the
state oracle. Stale-fence cases always require native dispatch regardless of a
fixture value supplied by untyped JavaScript. Match
Store-level typed errors rather than raw driver sentinels. The runner is
framework-agnostic, so the same fixture runs in the custom backend's own test
suite. Pair it with the shared cross-backend Store integration suite; transport
conformance alone cannot prove graph semantics.
The profile is family-scoped:
| Member | Store operations authorized |
| ----------------------------- | ----------------------------------------------------------------------------------------- |
| `createNodes` / `createEdges` | Eligible direct `bulkInsert()` and `bulkCreate()` programs |
| `replaceNodes` | Eligible complete-document `nodes.bulkReplaceById()` programs |
| `deleteNodes` / `deleteEdges` | Eligible direct `bulkDelete()` programs |
| `updateNodes` / `updateEdges` | Eligible resolved update-only sets |
| `mutateNodes` / `mutateEdges` | Eligible mixed create/update sets; the edge family also owns durable endpoint convergence |
Executor limits such as `maxEntries`, `replaceNodes.maxEntries.plain`,
`replaceNodes.maxEntries.claimed`,
`createNodes.claimSupport.maxInputCostPerEntry`, and the two edge mutation
ceilings are part of the registration contract and must be nonnegative
integers; zero honestly declares that the backend's bind budget cannot admit
one member of that shape. TypeGraph validates those declarations before
publishing the exact-root profile. `claimSupport.families` explicitly
advertises `uniqueness` and/or `disjointness`; an empty list with a zero bound
honestly opts out of all claim work. The Store calls the exported
`atomicNodeClaimInputCost()` owner for each member and refuses the native path
when its complete normalized claim set exceeds the executor's declared bound.
Custom executors must use that same helper instead of reproducing its
dialect-reviewed bind formula. `deleteNodes.releasedClaimFamilies` similarly
declares which owner-side claim cleanup the delete program proves.
Bundled replacement executors also expose an `accepts(entries)` pre-dispatch
proof. It packs prepared members with the same bind-weighted planner used by
execution, so claimed batches are admitted by their actual work instead of an
unrelated fixed 32-entry ceiling; `false` is an explicit no-SQL fallback
verdict. Custom executors may provide the same exact admission seam when one
claimed-member ceiling would be needlessly pessimistic.
`replaceNodes.releasedClaimFamilies` declares which previous owner claims the
replacement releases before acquiring its complete postimage claims; the Store
does not infer that proof from `claimSupport`. Node
create/update/mutation executors advertise derived-storage support separately
through `projectionSupport.families`; omission or an empty list honestly opts
out, and the Store never infers projection safety from transport registration
alone. The supported families are `fulltext` and `embedding`.
On a transactionless root, dedicated
update-only and mixed mutation executors are independently reachable Store
families even when their entry ceilings are equal, so each requires its own
conformance evidence. On an interactive root, the collection-level
read/partition/write unit moves into a transaction and exact-root registration
does not follow; the root conformance inventory therefore excludes the mixed
variants while continuing to require direct create, delete, update, and durable
convergence evidence. Bundled PostgreSQL binds the same reviewed lowering to the
exact transaction session and exercises the mixed node and edge variants against
a real engine, including typed refusal rollback. The same session profile
registers `replaceNodes`, so a caller-owned PostgreSQL transaction keeps blind
replacement inside its savepoint-backed atomic program.
An exact `atomicBatch: "session"` conformance fixture includes those mixed
variants even though `interactiveTransactions` is true; nested-transaction
isolation is reported as inapplicable because the fixture resource is already
the open transaction. The transport runner accepts the same exact-session
resource and certifies its ordered slots, parameter preservation, rollback,
and empty-program behavior. Generic derived session objects still lose both
the declaration and the identity-bound registrations.
Backend authors implementing edge
refusal paths use the exported
`AtomicEdgeBatchEndpointRefusalError`,
`AtomicEdgeBatchCardinalityRefusalError`,
`AtomicEdgeConvergenceTombstoneRefusalError`, and
`AtomicEdgeDeleteIdentityRefusalError` signals; restricted node deletion uses
`AtomicNodeDeleteRestrictedRefusalError`. This preserves the Store's existing
typed diagnostic classification rather than exposing driver-specific sentinel
errors.
Execution support is intentionally not collapsed into one ordered “tier.” An
interactive transaction and an exact-root atomic batch are independent facts:
a backend may provide either, both, or neither. Each Store operation selects
the boundary its own semantics require instead of treating one mechanism as a
universal substitute for the other.
Endpoint-set reads have a small, independent conformance fixture for custom
backends. Import `runEndpointSetReadConformance` from the `backend` entrypoint
and provide the exact backend, one or more successful `FindEdgesByEndpointSetParams`
cases, and at least one refusal case. The runner checks that the backend exposes
`findEdgesByEndpointSet`, preserves the expected edge rows, and refuses invalid
requests using the adapter's typed error. It does not create a schema or assume
a driver, so the same fixture can run against any engine. A backend that omits
the member remains valid for singleton reads; Store bulk endpoint reads refuse
with `ENDPOINT_SET_READ_UNSUPPORTED`.
### Declared constraints require an interactive transaction
A **constrained write** — one whose correctness rests on a check-then-write that
no database key repeats at write time — runs its probe and its write under one
per-graph mutual exclusion. That fence is a transaction-scoped construct on both
dialects: SQLite's `BEGIN IMMEDIATE` writer slot, PostgreSQL's
`pg_advisory_xact_lock` (which outside a transaction is taken and dropped inside
its own implicit single-statement one, excluding nothing). A backend reporting
`capabilities.execution.interactiveTransactions: false` can supply neither, so such a write is
**refused** rather than run unfenced — a constraint enforced only when nothing
races is the defect the fence exists to close.
The refusal is a `ConfigurationError` with `details.code`
`CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, and `details.constraint` naming which
class needed the fence, because the way forward differs per class:
| `details.constraint` | The write that needs the fence | Way forward without a transactional backend |
| --- | --- | --- |
| `edgeCardinality` | Creating or resurrecting an edge whose `cardinality` is `one`, `unique`, or `oneActive` | Declare the edge `cardinality: "many"` and enforce the limit in application code |
| `edgeMatchKeyConvergence` | `getOrCreateByEndpoints` using an undeclared dynamic `matchOn` key | Declare the edge registration's durable `matchIdentity`, or use `create` with a caller-chosen id |
| `nodeDisjointness` | Creating a node under a kind that participates in a `disjointWith` axiom | Drop the axiom and keep ids distinct across those kinds yourself |
| `nodeUniquenessClaim` | **Updating or resurrecting** a node whose kind declares any unique constraint, of any scope — a transition reserves the new key *before* the row write it gates, and only a transaction can undo the pair together | Drop the constraint, or run updates on a transactional backend. Plain **creates** under a `scope: "kind"` unique are unaffected: their claim follows the row |
| `nodeUniquenessScope` | Creating **or updating** a node under a `scope: "kindWithSubClasses"` unique that actually expands past the node's own kind | Scope the constraint to `"kind"`, which the uniques primary key enforces on its own |
`importGraph` / `importGraphStream` is refused on the same backends whenever any
node kind of the graph owes a claim ahead of its row — that is, declares **any**
unique constraint or has a disjoint partner — or any edge kind is non-`many`.
The import writes both creates and updates, so the widest of those placements is
what decides it.
This affects **Cloudflare D1**, **`drizzle-orm/neon-http`**, and any SQLite
backend built with `transactionMode: "none"`. Durable Objects are unaffected —
`do-sqlite` reports `capabilities.execution.interactiveTransactions: true` and fences normally.
Unconstrained writes on those backends are untouched and keep working exactly as
before: a `cardinality: "many"` edge created, updated and deleted; any node
delete, including one whose kind participates in a disjointness axiom (a delete
re-derives no cross-kind verdict); a node whose uniques are all `scope: "kind"`;
and an undeclared `getOrCreateByEndpoints` that *finds* an existing edge in the default
`ifExists: "return"` mode, or resurrects a `many` one — that resurrection is an
id-keyed `UPDATE` that re-derives nothing. With `coalesceUnchangedUpserts`
enabled, confirming that a single `ifExists: "update"` endpoint replay is
unchanged requires the endpoint match-key convergence fence and therefore
refuses on these backends. Outside the native durable-convergence envelope,
the bulk `getOrCreateByEndpoints` form returns an all-live default-`"return"`
batch from one set-oriented root read because that outcome writes nothing.
Inside the native envelope, the authoritative upsert runs first; it preserves
the logical `"found"` outcome in one exchange but may take incumbent-row locks
and produce write amplification. If any member may write, the whole batch
retains that refusal on transactionless roots unless it matches the narrow native
durable-convergence envelope: schema-declared
`matchIdentity`, `cardinality: "many"`, declared match fields, default
`ifExists: "return"`, and no temporal mutation. That eligible form is one
closed atomic exchange; dynamic match fields, update mode, constrained
cardinality, temporal options, and all transaction-scoped or derived roots
retain the refusal or fallback path required by their contracts.
An otherwise eligible tombstoned winner cannot use the native path: the native
attempt rolls back and transactionless convergence refuses with the typed
`CONSTRAINT_WRITE_FENCE_UNSUPPORTED` (`edgeMatchKeyConvergence`) error. Use a
transaction-capable backend when schema-aware resurrection is required.
### Claim relations, and what they do not promise
Underneath the lock, a declared constraint is also reserved in a **claim
relation** whose primary key admits one live claimant per axis: `uniques` (for
uniqueness scopes and `disjointWith` pairs) and `typegraph_edge_claims` (for
`cardinality: "one" | "unique" | "oneActive"`). Both bundled backends carry them
and report `capabilities.constraintClaims: true`. The claim is what makes those
constraints hold for TypeGraph writers that hold no per-graph lock at all —
`importGraph` is the one in the box. The protocol is application-maintained:
raw SQL that writes only `nodes` or `edges` bypasses the corresponding claim
write and can violate the declaration. An out-of-band writer is fenced only if
it participates in the same claim protocol in the same transaction.
Three properties of that mechanism are worth knowing before you rely on it:
- **A claim row's lock is held to the end of the transaction, including on
refusal.** A caller that catches a typed constraint error and keeps going —
import's per-row recovery, or your own `try`/`catch` inside
`store.transaction` — still holds the lock on the row it was refused at, and
any other writer of that axis waits until the transaction ends. This is
inherent to every row-lock fence, not specific to this one.
- **Above READ COMMITTED, PostgreSQL reports a serialization failure instead of
the typed error.** At `REPEATABLE READ` or `SERIALIZABLE`, `INSERT … ON
CONFLICT DO UPDATE` raises `40001` rather than resolving the conflict, so the
losing writer sees a serialization failure to retry rather than
`UniquenessError`. SQLite has no such mode. This is unchanged from earlier
versions, which already reserved single-kind uniqueness through the same
statement.
- **Pre-existing violations are neither repaired nor refused at boot.** A
database that already held two live claimants of one axis before the claim
relations existed keeps holding them; the next write that touches that axis is
refused with the ordinary typed error naming the incumbent.
`store.verifyConstraintFences()` is the read-only diagnostic that makes that
state legible ahead of time:
```typescript
for (const violation of await store.verifyConstraintFences()) {
// violation.target names the claim row two claimants contend for
console.warn(violation.family, violation.target.axis, violation.target.key);
}
```
It reports one entry per contended axis — `nodeUniqueness` and
`nodeDisjointness` carry the conflicting `owners` (each a `concrete_kind` /
`node_id` pair, because ids are unique only per kind), `edgeCardinality` carries
the conflicting `edgeIds`. It reads the nodes, edges and `uniques` relations, so
it finds violations that predate the claim tables; it writes nothing, and it
repairs nothing — choosing which claimant keeps the axis is a data-loss decision
that stays with you.
### SQLite ↔ PostgreSQL parity
The **query language is fully portable** between SQLite and PostgreSQL. Predicates (comparison, string/`ILIKE`,
null, `between`, array, object, JSON-path), fixed and variable-length (recursive) traversals, bounded
neighbor reads, per-edge-kind subgraph windows, one-statement query batches, aggregates
(`count`/`sum`/`avg`/`min`/`max` with `groupBy`/`having`), set operations (`UNION`/`UNION ALL`/`INTERSECT`/`EXCEPT`,
including traversal, subquery, `GROUP BY`/`HAVING`, and per-leaf `ORDER BY`/`LIMIT`/`OFFSET` leaves), ordering with
`NULLS FIRST`/`LAST`, cursor pagination, temporal queries, and the fulltext query modes (`websearch`, `phrase`,
`plain`, `raw`) all behave identically. A query you write against one backend compiles and runs the same way on the
other.
The remaining differences are **engine and runtime capability gaps** — they
stem from what each database or hosted authorizer implements, not from
TypeGraph choosing separate query semantics per backend:
| Capability | SQLite | PostgreSQL | Behavior on the unsupported side |
| ------------------------------------------------------ | ------------------------------------------------- | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------- |
| Whole-graph temporary-table analytics | ✓ standard connections / ✗ D1 and Durable Objects | ✓ connection-based drivers / ✗ `neon-http` | Throws `UnsupportedBackendCapabilityError`; traversal algorithms with an inline engine fall back automatically |
| Vector metric `inner_product` | ✗ | ✓ | Rejected at compile time on SQLite (`sqlite-vec`/`libsql-native` expose `cosine` + `l2`; `pgvector` adds `inner_product`) |
| Vector index type `ivfflat` | ✗ | ✓ | Index declaration is **skipped** on SQLite (`indexTypes`: `hnsw`/`none` vs `hnsw`/`ivfflat`/`none`) |
| Filtered approximate search **guarantees** a full page | ✓ `sqlite-vec` / ✗ `libsql-native` | ✗ (`pgvector` recovers, but is bounded) | Only `sqlite-vec` guarantees it; the others can return **fewer than `limit`** rows under heavy filtering — see below |
| Per-query fulltext `language` override | ✗ | ✓ | Throws on SQLite — FTS5's tokenizer is fixed at table-create time; `tsvector` accepts a regconfig per query |
| HNSW `efSearch` query tuning | ✗ | ✓ transactional HNSW drivers | Refused, never ignored: `UnsupportedBackendCapabilityError` with `details.capability` `vector.searchFrontierTuning` on **any** SQLite backend (vector and hybrid alike — neither `sqlite-vec`'s `vec0` KNN nor `libsql-native`'s DiskANN has a per-search frontier), and on transaction-less Postgres or a non-HNSW slot |
| Bounded planner-statistics sampling | ✓ standard connections / ✗ D1 and Durable Objects | Native `ANALYZE` sampling | Restricted SQLite skips `analysis_limit` but still attempts scoped `ANALYZE`. Performance only — same results |
| TypeGraph Identity Profile | ✓ transactional drivers | ✓ transactional drivers | Enabled graphs fail fast on non-atomic drivers; identity-disabled graphs retain their ordinary path |
| Constraint claim relations (`capabilities.constraintClaims`) | ✓ | ✓ | Identical relations and identical statements on both dialects. A third-party backend that omits them declares `constraintClaims` absent and keeps the per-graph lock as its only fence |
| Durable edge match identity (`capabilities.durableEdgeMatchIdentity`) | ✓ bundled adapters | ✓ bundled adapters | Both dialects persist the same canonical key and use a unique database arbiter. A custom backend must satisfy the full capability contract above or leave the capability absent |
| Managed node projection fusion | ✓ registered atomic bulk programs; singleton fallback | ✓ registered atomic bulk programs; singleton create fusion | Eligible node bulk creates and resolved updates group fulltext/vector transitions into the same atomic program as their row mutations on both dialects. PostgreSQL additionally fuses an eligible singleton generated-ID create into one SQL statement when every active strategy supplies an inserted-node builder |
| Managed node claim fusion (`capabilities.atomicNodeInsertClaims`) | ✗ portable transactional fallback | ✓ PostgreSQL/PGlite | SQLite keeps claim acquisition and insertion in the portable transaction. PostgreSQL transaction receivers fuse supported claim plans; a root non-transactional receiver is limited to exactly one generated-id, same-kind uniqueness claim with no other side effects |
| Managed edge cardinality fusion | ✗ portable transactional fallback | ✓ PostgreSQL/PGlite transaction receivers | SQLite keeps its guarded claim and edge insert in the portable transaction. PostgreSQL can combine endpoint liveness, one cardinality claim, and the insert in one statement after any required graph lock |
| Atomic SQL transport (`capabilities.execution.atomicBatch`) | ✓ on certified D1/libSQL roots; otherwise `none` | ✓ on bundled recognized PostgreSQL drivers, including neon-http | `root` means the exact backend owns the atomic boundary; `session` means the exact object is already bound to an open transaction and the outer transaction owns commit/rollback. Both require identity-keyed executor registration. Neon HTTP uses its native transaction batch; session-capable `pg`, postgres-js, neon-serverless, and PGlite drivers can execute programs on one pinned Drizzle transaction. Unrecognized drivers remain `none`. A custom backend must pass the framework-agnostic conformance runner before opting in; omitted support keeps the portable path |
| Eligible registered managed writes | ✓ bundled SQLite roots, including D1 and libSQL | ✓ bundled PostgreSQL roots, including neon-http | Eligible singleton generated-ID nodes and `cardinality: "many"` edges use one authoritative create statement. Eligible node updates may carry fulltext/vector replacements; unconstrained non-durable-identity edge updates, direct edge deletes, and plain restricted node deletes use one authoritative read/gate plus one registered atomic mutation. Generated-, caller-, or mixed-ID node `bulkInsert`/`bulkCreate` batches compose supported multi-claim/cross-scope claim sets with projections in one schema-fenced native program; direct edge programs also maintain durable match identity and cardinality claims. Direct edge `bulkDelete` and plain restricted node `bulkDelete` use the same mutation profile. Eligible mixed `bulkUpsertById` sets, including node projections, use the profile on serverless roots and on exact bundled PostgreSQL transaction sessions; a generic derived backend still loses the evidence. A custom backend may opt in per family only after registering its exact transport and semantic executor. Unregistered or otherwise ineligible families, projected/identity-enabled node deletes, over-budget claimed members, cascade/disconnect deletes, and other managed writes retain the existing path |
| Typed constraint error above READ COMMITTED | n/a (no such isolation mode) | ✗ at `REPEATABLE READ` / `SERIALIZABLE` | PostgreSQL raises `40001` from the claim's upsert instead of resolving the conflict, so the loser retries a serialization failure rather than reading `UniquenessError` |
| Claim row lock released before end of transaction | ✗ | ✗ | Held to commit/rollback on both dialects, refusal included — a caller that catches a constraint error blocks other writers of that axis for the rest of its transaction |
| Recursive traversal (`capabilities.recursiveTraversal`) | ✓ | ✓ | Identical on both bundled backends. A third-party backend declaring `{ supported: false, reason }` refuses the five recursion-dependent operations with `ConfigurationError` code `RECURSIVE_TRAVERSAL_UNSUPPORTED`; `weightedShortestPath` degrades to a predecessor walk instead — see above. Unweighted `shortestPath` is unaffected — it never emits a recursive CTE |
| Write fence (`capabilities.writeFence`) | ✓ `engine-serialized` (single writer slot) | ✓ `lock` (advisory + table locks) | Identical guarantee, different mechanism. A custom backend that declares no `writeFence` resolves `unfenced` and is refused at construction for Operational Identity or TypeGraph-owned recorded-clock allocation |
| Capability bundles (`CAPABILITY_BUNDLES`) | Identical | Identical | Both bundled backends implement every pilot bundle's core/extra members on both dialects it scopes to. A third-party backend with a port gap refuses (gated core, or a `refuse`-disposition extra) or degrades (a `fallback`-disposition extra) per that bundle's own registry row |
| Engine-native lineage (`backend.lineage`) | ✗ (recorded-relations lineage via `history`) | ✗ (recorded-relations lineage via `history`) | Neither bundled profile declares its own `lineage`. `resolveLineage` derives it from the store's recorded relations whenever `history: true` is on, identically on both dialects, so a graph-merge diff against such a store is pruned the same way regardless of backend. Without `history`, no lineage source resolves and the diff is full; the anchor is the durable revision anchor when `revisionTracking: true`, otherwise the compatibility content fingerprint |
| Engine-native recorded time (`backend.recordedTime`) | ✗ (TypeGraph-owned recorded relations via `history`) | ✗ (TypeGraph-owned recorded relations via `history`) | Neither bundled profile declares `recordedTime`, so `resolveRecordedTimeOwnership` derives `"typegraph-relations"` for both — `history: true` captures into TypeGraph's own recorded relations and clock, identically on both dialects, and every recorded-time integration suite and the parity snapshot run unchanged. A backend that supplies `recordedTime` (and the co-required `lineage`) reads and writes recorded time through its own engine instead: no TypeGraph capture, clock, or recorded relations, `RecordedInstant` anchors in the `e1:` form, and several TypeGraph-relation-specific surfaces refused — see [Engine-native recorded time](/queries/temporal#engine-native-recorded-time) and [Supplying `recordedTime`](/backend-authoring#supplying-recordedtime). No bundled backend implements this today; the capability is proven by a PostgreSQL-family simulation (shared by the always-running `tests/backends/postgres/pglite-engine-native-recorded-time.test.ts` and the `POSTGRES_URL`-gated `tests/backends/postgres/engine-native-recorded-time.test.ts`, both built on `engine-native-recorded-time-simulation.ts`) that dresses TypeGraph's own recorded relations as a temporal-table expression, labeled as a simulation rather than a real third engine |
Identity support also has a **driver** dimension inside each dialect:
| Driver | Atomic identity support | Behavior |
| ------------------------------------------------------------ | ----------------------- | ------------------------------------------------------------------------------------------------------------------- |
| Managed SQLite, libSQL, Durable Objects | ✓ | Full profile |
| PostgreSQL `node-postgres`, `postgres-js`, neon-serverless, PGlite | ✓ | Full profile; identity-affecting writes serialize per graph, limiting each graph to one identity writer at a time |
| Cloudflare D1 | ✗ | Enabled graphs fail at store construction with `ConfigurationError` details code `IDENTITY_REQUIRES_ATOMIC_BACKEND` |
| `drizzle-orm/neon-http` | ✗ | Same fail-fast error; identity-disabled graphs retain the ordinary single-statement path |
### Filtered approximate search
Every approximate (ANN) vector search carries at least one row filter: the liveness predicate that hides
soft-deleted and out-of-validity rows. A `.where(...)` predicate narrows it further. Engines differ in where they
apply that filter relative to the index traversal, which decides whether a page can come back short. Read it from
`backend.capabilities.vector.filteredApproximateSearch`:
```typescript
const filtered = backend.capabilities.vector?.filteredApproximateSearch;
if (filtered?.guaranteesFullPage !== true) {
// An approximate search here may return fewer than `limit` rows.
}
```
**Check `guaranteesFullPage`, not `mode`.** `mode` names the mechanism the strategy asks the engine for; only
`guaranteesFullPage` tells you whether a short page is possible.
| `mode` | Strategy | `guaranteesFullPage` | Meaning |
| ------------------- | --------------- | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `"filter-pushdown"` | `sqlite-vec` | `true` | The filter constrains the `vec0` KNN candidate set itself. `limit` matching rows come back whenever `limit` exist. |
| `"iterative-scan"` | `pgvector` | `false` | The index is re-entered for more candidates (`hnsw.iterative_scan` / `ivfflat.iterative_scan`, applied automatically on pgvector ≥ 0.8). Much better recall than a post-filter, but **not a guarantee**: the scan stops at `hnsw.max_scan_tuples` / `ivfflat.max_probes`. And on **pgvector < 0.8** there is no iterative scan at all — the backend detects that, warns once, and the search stays `ef_search`-bounded. |
| `"post-filter"` | `libsql-native` | `false` | DiskANN's `vector_top_k` is a table function with no filter pushdown and no way to re-enter the index. TypeGraph over-fetches `4 × (limit + offset)` neighbors and filters afterwards, so once more than that headroom is filtered out **the search silently returns fewer than `limit` rows while more matches exist**. |
Heavy tombstone drift — routine in a temporal store — is what turns a bounded search from a theoretical caveat into
a short page. When a full page matters, use an exact search (`approximate: false`), which scans and so applies the
filter to every row; or declare the field's index as `"none"` so it is always brute-forced.
Vector and fulltext capabilities are populated from the configured strategy, so the matrix above reflects the
bundled strategies (`sqlite-vec`/`libsql-native`/`pgvector`, `fts5`/`tsvector`). A custom strategy advertising
different `metrics`/`indexTypes`/`filteredApproximateSearch`/`searchFrontierTuning` shifts these rows accordingly —
always check `backend.capabilities` at runtime rather than hard-coding the dialect.
`searchFrontierTuning` is **required** on a vector strategy's capabilities, so a strategy must state whether it has a
per-search ANN frontier knob rather than inheriting silence. It is a discriminated union: `{ tunable: true, parameter,
indexType, requiresTransactionScope }` names the engine parameter `efSearch` maps to (`pgvector`: `hnsw.ef_search`, on
an `hnsw` slot, needing a transaction to scope it), while `{ tunable: false, reason }` names why the engine has no such
knob and is what makes `efSearch` a typed refusal there. A hand-written strategy that omits the field no longer
compiles.
Both bundled backends advertise `windowFunctions: true`. Relation `topPerPartition()` refuses execution
with `UnsupportedBackendCapabilityError` when a custom backend sets `windowFunctions: false`.
Vector, fulltext, and hybrid relevance-ranking
queries use `ROW_NUMBER()` internally and throw `ConfigurationError` before SQL generation if a custom backend profile
sets `windowFunctions: false` — there the window output *is* the result (the relevance k-cutoff / rank ordinal), so
there is no correct fallback.
`bulkFindByIndex({ limitPerInput })` also uses `ROW_NUMBER()` when available, but it does **not** throw on a
windowless profile: the per-input cap is a transfer optimization with identical row semantics either way, so it
degrades gracefully — fetching all matching ids and capping per group in application code. The unbounded
`bulkFindByIndex` path needs no window and is always available.
:::note[JSON is native on both backends]
SQLite stores JSON as text and queries it with the built-in JSON functions (`json_extract`, `json_each`, …);
PostgreSQL uses native `JSONB`. The dialect layer hides this difference, so JSON-path predicates and **B-tree
expression indexes on scalar JSON properties** (`defineNodeIndex` / `defineEdgeIndex`) are at full parity. The one
JSON-related difference is performance, not capability: PostgreSQL can use a single GIN index to accelerate
array/object **containment** predicates (`contains()` / `containsAll()` / `hasKey()` / `pathEquals()`), whereas on
SQLite those run as `json_each()` scans — correct results, just not index-accelerated. See
[Indexes](/performance/indexes) for the full breakdown.
:::
:::note[Transactions are driver-dependent, not backend-dependent]
Both backends report `execution.interactiveTransactions: true` by default. The exception is symmetric and lives in
specific drivers:
Cloudflare D1 (SQLite) and `drizzle-orm/neon-http` (Postgres) are non-transactional, so they downgrade to
`execution.interactiveTransactions: false`. Operations that require atomicity (`commitSchemaVersion`,
`setActiveVersion`, Operational Identity) throw on those drivers regardless of backend. A
schema-managed Store's write that cannot fuse its schema fence into its own statement fails closed
the same way, because it has no other way to hold the transaction-scoped fence; `store.transaction()`
refuses on those roots regardless. See
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for which
writes fuse. Eligible operations
with a certified atomic SQL program remain available independently of this interactive transaction capability.
:::
:::note[Aggregate set operations are a builder limitation, not a parity gap]
`GROUP BY`/`HAVING` leaves are supported by the set-operation compiler on **both** backends, but the query builder
does not expose `.union()`/`.intersect()`/`.except()` on `.aggregate()` queries. That limit applies equally to SQLite
and PostgreSQL, so it is not a portability difference.
:::
## Connection Management
Connection ownership follows the entrypoint:
- **Managed Store factories** (`/sqlite/local` and `/postgres/pglite`) own the
connection and provisioned resources. `await store.close()` releases them.
- **Owned local backend factories** (`createLocalSqliteBackend` and
`createLocalPgliteBackend`) also own their resources. A Store delegates
`close()` to its backend, so `await store.close()` releases them.
- **Bring-your-own adapter factories** (`createSqliteBackend`,
`createPostgresBackend`, and `createLibsqlBackend`) leave connection ownership
with the caller. Their Store's `close()` does not close the supplied client or
pool.
When you bring your own connection, you are responsible for:
1. **Creating connections** with appropriate configuration
2. **Connection pooling** for production use
3. **Closing connections** on shutdown
```typescript
// You create the connection
const sqlite = new Database("app.db");
const db = drizzle(sqlite);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
// You close the connection
process.on("exit", () => {
sqlite.close();
});
```
Here `store.close()` leaves `sqlite` open because the application supplied the
connection. Close the driver or pool through its own API.
### Serialized connections
Some drivers run every statement through **one** connection. Two long-lived
interchange streams cannot share such a connection — an export snapshot holds a
read transaction for the whole stream while an import writes one per chunk — so
TypeGraph refuses the second one with a typed error instead of letting it hang
(see
[Interchange serialized-connection guard codes](/errors#interchange-serialized-connection-guard-codes)).
Recognizing a serialized connection means recognizing the *driver*, from the
shape of the client object. That is deliberately conservative: a driver
TypeGraph cannot positively identify is left unmarked, because refusing a pooled
connection would refuse work that succeeds.
| Driver / configuration | Detected | Notes |
| --------------------------------------------------------------------------------------------- | ------------------ | ---------------------------------------------------------------------------------------------------------------- |
| better-sqlite3, bun:sqlite, sql.js, local libSQL (`file:` / `:memory:`), Durable Object storage | ✓ automatic | One handle, one connection |
| PGlite | ✓ automatic | One in-process WASM connection |
| Bare `pg` / neon-serverless `Client`, a checked-out `PoolClient` | ✓ automatic | One owned socket |
| `pg` `Pool` capped at one (`{ max: 1 }`, `{ max: "1" }`, `{ poolSize: "1" }`) | ✓ automatic | pg-pool does not coerce the cap, so the string forms are the same one-connection pool |
| postgres-js capped at one (`{ max: 1 }`, `?max=1`, `PGMAX=1`) | ✓ automatic | Same reasoning on the postgres-js side |
| Default-size pools, `neon-http`, D1, RDS Data API, remote libSQL (`http` / `ws`) | — deliberately not | Each statement gets an independent connection; refusing would refuse work that succeeds |
| `expo-sqlite`, `op-sqlite`, `sqlite-proxy`, `pg-proxy`, a bespoke adapter | ✗ **declare it** | Serialized in fact, but the client exposes no shape TypeGraph can attribute to a known driver |
| Bun `SQL` (Postgres) at `{ max: 1 }` | ✗ **declare it** | The cap is readable, but nothing identifies the driver, and a cap on an unknown client is not evidence |
| postgres-js with a non-numeric string cap other than one, e.g. `?max=5` | ✗ **declare it** | Opens exactly one connection today only because postgres-js does not coerce the value — marking it would encode an upstream bug that will one day be fixed |
For the rows marked **declare it**, tell TypeGraph what it cannot see. The
option is on `createSqliteBackend` and `createPostgresBackend` — the two
factories that resolve it. The batteries-included wrappers
(`createLibsqlBackend`, `createLocalSqliteBackend`, `createLocalPgliteBackend`)
do not take it, because each already detects its own connection.
`{ mode: "shared", resource: pool }` is incorrect for a `pg.Pool` that can open
multiple connections, even if several backends use that pool. Each transaction
checks out its own connection; marking the pool as one resource makes independent
snapshot exports and imports contend for a single lease and refuses concurrent
operations that the pool can run. Leave the declaration absent for such a pool.
```typescript
const sql = postgres(process.env.DATABASE_URL + "?max=5");
const backend = createPostgresBackend(drizzle(sql), {
// This client really does run every statement on one connection.
serializedResource: { mode: "shared", resource: sql },
});
```
Two backends that name the **same** object are one serialized resource, exactly
as two wrappers over a detected client are. Naming a *different* object than the
one TypeGraph detected is refused with a `ConfigurationError`
(`details.reason: "serialized-resource-conflict"`) rather than silently
preferred: two wrappers over one connection given two different sentinels would
stop being seen as a pair, which is the failure the guard exists to prevent.
The refusal names each side by constructor (`details.declaredKind` /
`details.detectedKind`) instead of carrying the two handles, because `details`
is what `toLogString()` serializes and a driver handle there would log whatever
that driver stores — a `pg.Pool` keeps its `connectionString`.
The reverse declaration escapes a detection that is wrong for your topology:
```typescript
const backend = createSqliteBackend(db, {
serializedResource: { mode: "independent" },
});
```
**Scope.** `{ mode: "independent" }` lifts the *shared-resource* refusal between
two distinct backend objects. It does not lift the object-identity refusal, under
which one SQLite backend exporting into **itself** is refused with
`INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`. That one is a fact about a single
handle holding a single open snapshot transaction, not a claim about connection
topology, so no declaration can make it false — pass a second backend instead.
That surviving refusal is SQLite-only, so on PostgreSQL the declaration lifts
the refusal for one backend exporting into itself as well: a client that hands
out independent connections — which is exactly what the declaration claims —
runs the snapshot and the writes it contends with on different ones.
## Database roles & least privilege
`createStoreWithSchema()` and `createStore()` divide cleanly along DDL
privilege, so a production deployment can run its application under a
least-privilege, DML-only database role.
- **`createStoreWithSchema(graph, backend)` runs DDL.** It bootstraps the
base tables on a fresh database, applies safe auto-migrations, and
adopts release-added deployment-wide base storage on pre-provisioned
databases, even when the persisted graph schema is unchanged. The first
adoption creates or repairs the graph-template and edge match-identity
storage, then stamps a version marker; a warm base-schema check is one
`SELECT` with no base-adoption DDL. It also
durably materializes strategy-owned runtime storage — both fulltext and
each `embedding()` field's per-`(kind, field)` vector table, plus a
durable marker for each. It also brings TypeGraph's own base-relation
**system indexes** up to the running library version: bootstrap DDL only
runs on the very first boot, so an index shipped in a newer version
reaches an already-initialized database through this step (built with
`CREATE INDEX CONCURRENTLY` on PostgreSQL; a database whose indexes all
exist settles from the catalog with no DDL). Contribution and system-index
preparation have their own catalog checks and may still issue DDL, so the
role it runs under **must hold `CREATE` / DDL privileges**. Run it once at startup, outside request
handlers and transactions. (`store.evolve()` likewise provisions any
embedding field it introduces, so it too needs DDL privileges.)
Deployments that never run `createStoreWithSchema` (manual-DDL boot with
a plain `createStore` attach) adopt new system indexes by calling
`store.materializeSystemIndexes()` once under a DDL-capable role after
upgrading; deployments that must not run index builds inline at boot
(large tables behind a readiness probe) pass `systemIndexes: "skip"` to
`createStoreWithSchema` and materialize out-of-band the same way.
- **`createStore(graph, backend)` is a synchronous, zero-I/O attach.**
It does not create tables, repair DDL, or record that runtime storage
is materialized — it issues **no DDL ever**. Use it only to attach to a
database a prior `createStoreWithSchema` boot already initialized. A
fulltext operation or an **embedding write** against a database that was
never initialized — a `create({ embedding })` or embedding update/delete
— throws `StoreNotInitializedError` rather than silently emitting
`CREATE TABLE` on the hot path. (Vector *reads* are not marker-gated:
`store.search.vector`, `store.search.hybrid`, and a query-builder
`.similarTo()` predicate compile to SQL against the per-field table
directly, so on an un-provisioned database they surface the engine's own
missing-relation error instead — `no such table: tg_vec_…` on SQLite,
`relation … does not exist` on Postgres. Same cause, same fix; use
`createVerifiedStore` to catch it at attach rather than at first query.)
This is what lets a least-privilege role run vector ops: the table
already exists. Graphs with no `searchable()` or `embedding()` fields
are unaffected.
The Store is also raw and unversioned: its writes do not participate in
the schema-version fence. Direct backend writes have the same semantics.
Quiesce those writers yourself before changing schemas.
- **`createVerifiedStore(graph, backend)` is the same zero-DDL attach
with a verification gate.** It reads the active schema row, folds the
persisted graph extension, and refuses to construct the Store unless
the database is at the same schema version as the code graph. Throws
`BaseSchemaMigrationError` when deployment-wide base storage is missing,
stale, or newer than the library, `MigrationError` on graph-schema drift
(safe or breaking), `ConfigurationError` when no graph schema has been
initialized, and `StoreNotInitializedError`
when the schema is current but runtime-contribution markers are
missing. The runtime-side counterpart of `createStoreWithSchema` for
least-privilege deployments. If you only need the gate without
building a Store (e.g. a readiness probe), call `assertSchemaCurrent`.
Its managed writes require a transactional backend with the schema-write
fence; non-transactional and unsupported custom backends can attach for
reads but fail closed on the first write.
The adapter equivalents (`createAdapterStoreWithSchema` and
`createVerifiedAdapterStore`) carry the same managed metadata. So does
`createAdapterStore(..., { reconciled })` with a cached reconciliation snapshot,
and Stores returned by `evolve()` or rebound from an already-managed Store.
Check `store.introspect().schemaVersion !== undefined` at runtime. Calling
`store.clear()` deletes the schema rows and resets that Store to raw semantics;
reopen it through a managed factory before resuming version-fenced writes.
- **`store.verifyContributions()` diagnoses contribution storage;
`store.repairContributions()` repairs safe findings under a privileged
role.**
Every gate above trusts the marker row without probing the catalog, so
a database whose strategy-owned tables were dropped out of band opens
clean and fails at the first dependent read or write. This method compares each contribution
currently expected by the active graph and backend strategies with its
marker and the catalog. It does not audit retired marker rows, and a
never-attempted contribution with neither marker nor table is omitted, so
an empty result is not initialization proof. It is read-only (`SELECT`
only, no DDL) so the least-privilege role can run it, and it is deliberately
not part of any open path. For a readiness check, construct the Store with
`createVerifiedStore()` first and then run this diagnostic; otherwise use
it as an operator check. The repair method re-audits current declarations,
preserves data while repairing `missing-marker` and
`failed-materialization`, and reports `stale` or `orphaned-marker` as
`requires-rebuild`. Run repair through the DDL-capable migration role, not
the least-privilege runtime role. Follow the per-state table in
[The store opens clean but a fulltext or vector read fails](/troubleshooting#the-store-opens-clean-but-a-fulltext-or-vector-read-fails)
rather than applying one repair to every entry.
- **`store.probeContributions()` is the read-only readiness check;
`store.rebuildContribution()` is the destructive last resort.**
The two bracket `repairContributions()` into one escalation ladder:
probe (writes nothing) → repair (non-destructive) → rebuild
(destructive, but scoped to the calling graph). The probe reports one
`ready` / `degraded` entry per
search projection and is safe on a read path, on a replica, and under
the least-privilege role — it shares the detection logic of the other
two rather than reimplementing it, so it cannot disagree with the gate
the hot path actually consults. The rebuild is the only repair for a
`stale` contribution, whose table exists at a shape the current
`createDdl` no longer produces; it deletes and refills only the calling
graph's rows in the shared fulltext table, escalating to drop → recreate
when that table holds no other graph's rows (under a database-scoped DDL
advisory lock, since that DDL is database-global), and runs the whole
sequence inside one transaction under the schema-write fence. It refuses
with `ContributionRebuildUnsupportedError` for vector storage, whose
embeddings exist only in the table it would drop
(`reason: "vector-source-unavailable"`), and for a `stale` shape whose
storage still holds other graphs' rows
(`reason: "shared-storage-in-use"`, naming them in
`details.otherGraphIds`). Run rebuilds through
the DDL-capable migration role, in a maintenance window: the
transaction is held for the whole refill, and on PostgreSQL a drop's
`ACCESS EXCLUSIVE` lock blocks both searches and writes to any kind with
`searchable()` fields until it commits. Reach it from a `createStore()`
Store — the managed factory's boot step refuses to open while a
contribution is `stale`. See
[Contribution health: probe, repair, rebuild](/troubleshooting#contribution-health-probe-repair-rebuild).
Strategy contributions declare an ownership `scope`: `"graph"` (the
default for older custom strategies) provisions one physical contribution
per graph, while `"deployment"` provisions shared physical storage once
under TypeGraph's reserved deployment marker and records a separate
graph-local activation marker. The built-in full-text strategies use
deployment scope; vector slots remain graph-scoped. A subsequent graph open
reads both attestations and performs no DDL, so it can run under a
DML-only role without disabling full-text search. The reserved marker key is
exported as `DEPLOYMENT_CONTRIBUTION_GRAPH_ID`; graph definitions must not
use that id.
### Contribution capability parity
`backend.capabilities.contributions` declares how far up the ladder a
backend goes. Each rung is separate because a backend can genuinely stop
at any of them, and a rung a backend cannot serve refuses with a typed
error rather than returning something that looks like success.
| Backend | `supported` | `probe` | `rebuild` |
| --- | --- | --- | --- |
| SQLite (better-sqlite3, bun:sqlite, libSQL, Durable Objects) | ✅ | ✅ | ✅ |
| SQLite with `transactionMode: "none"` | ✅ | ✅ | ❌ no schema fence |
| PostgreSQL (`pg`, `postgres-js`, PGlite, `neon-serverless`) | ✅ | ✅ | ✅ |
| PostgreSQL over `neon-http` | ✅ | ✅ | ❌ no schema fence |
| Custom fulltext strategy without `dropDdl` | ✅ | ✅ | ❌ no teardown DDL |
| Fulltext disabled (`fulltext: false`) | ✅ | ✅ | ✅ with a schema fence |
`rebuild` requires two things at once: a fulltext strategy that declares
`dropDdl` on its contribution, and a transactional schema fence
(`schemaWriteTransaction`) to run the sequence under. The HTTP-only
PostgreSQL drivers cannot hold a session across statements, so they have
no fence — the same reason they already report
`capabilities.execution.interactiveTransactions === false`. A third-party strategy predating
`dropDdl` keeps working for every other operation and is reported as not
rebuildable rather than being dropped through a synthesized statement
TypeGraph guessed at. Vector contributions are never rebuildable on any
backend; that is a property of what TypeGraph stores, not of the engine. A
backend built with `fulltext: false` has no fulltext contribution at all, so
the first condition is vacuously satisfied and `rebuild` reduces to whether
the backend has the transactional schema fence — the same value it would
report if fulltext were still active on a driver with that fence.
**`fulltext: false` stops creating and maintaining the fulltext table; it
never drops one.** On a database that already carries fulltext rows,
disabling fulltext leaves them in place and unmaintained: a hard delete
performed while fulltext is off leaves an orphaned row behind in the
fulltext table, because `hardDeleteNode`'s cascade has no active strategy to
build a delete statement from. Re-enabling fulltext later therefore requires
the destructive contribution rebuild — `store.rebuildContribution("fulltext")`,
which drops and recreates the fulltext table — **not**
`store.search.rebuildFulltext()`: that method pages live nodes to recompute
their content, and a hard-deleted node has no row left in the node table for
it to page, so it never revisits, and therefore never clears, the orphan.
### Recommended deployment shape
Run schema/DDL changes as a **privileged, one-time migration step**, then
run the application under a **least-privilege runtime role** that holds
only `SELECT` / `INSERT` / `UPDATE` / `DELETE`:
```typescript
// 1. Migration step — privileged role with DDL/CREATE.
//
// createStoreWithSchema is mandatory here: it bootstraps tables,
// applies safe auto-migrations, commits the schema_versions row,
// and writes the durable contribution markers. The runtime gate
// checks all of those.
const [/* store */] = await createStoreWithSchema(graph, adminBackend);
// Optional prerequisite if you manage DDL externally with
// drizzle-kit. Generated SQL creates the tables but does NOT
// initialize the schema row or contribution markers — still run
// createStoreWithSchema afterwards (it skips bootstrap when tables
// already exist and commits the row + markers):
//
// import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// await adminPool.query(generatePostgresMigrationSQL());
// await createStoreWithSchema(graph, adminBackend);
```
```typescript
// 2. Runtime — least-privilege, DML-only role. Zero DDL.
// createVerifiedStore fails fast if the privileged migrator is behind.
const runtimePool = new Pool({ connectionString: process.env.APP_DATABASE_URL });
const backend = createPostgresBackend(drizzle(runtimePool));
const [store] = await createVerifiedStore(graph, backend);
```
If the runtime role has no DDL privileges and you boot it with
`createStoreWithSchema()` anyway, the first cold boot fails with a
permission error on the bootstrap or contribution-marker DDL — see
[Troubleshooting](/troubleshooting).
## Environment-Specific Setup
### Development
```typescript
// In-memory for fast tests
const { backend } = createLocalSqliteBackend();
// Or file-based for persistence during development
const { backend } = createLocalSqliteBackend({ path: "./dev.db" });
```
### Testing
```typescript
// Fresh in-memory database per test
beforeEach(() => {
const { backend } = createLocalSqliteBackend();
store = createStore(graph, backend);
});
```
### Production
Single-role setup — `createStoreWithSchema` bootstraps and migrates on
boot, so the role needs DDL privileges. To run the application under a
least-privilege, DML-only role instead, split the migration step out as
described in [Database roles & least privilege](#database-roles--least-privilege).
```typescript
// PostgreSQL with pooling
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20,
ssl: { rejectUnauthorized: false }, // For managed databases
});
const db = drizzle(pool);
const backend = createPostgresBackend(db);
const [store] = await createStoreWithSchema(graph, backend);
```
## Next Steps
- [Schemas & Types](/core-concepts) - Define your graph schema
- [Semantic Search](/semantic-search) - Vector embeddings and similarity search
- [Limitations](/limitations) - Backend-specific constraints
# Query Builder Overview
> A fluent, type-safe API for querying your graph
TypeGraph provides a fluent, type-safe query builder for traversing and filtering your graph. This
page introduces the query categories and how they compose together.
## Query Categories
Every query builder method falls into one of these categories:
| Category | Purpose | Key Methods |
|----------|---------|-------------|
| [Source](/queries/source) | Entry point - where to start | `from()` |
| [Filter](/queries/filter) | Reduce the result set | `whereNode()`, `whereEdge()` |
| [Traverse](/queries/traverse) | Navigate relationships | `traverse()`, `optionalTraverse()`, `to()` |
| [Recursive](/queries/recursive) | Variable-length paths | `recursive()` |
| [Shape](/queries/shape) | Transform output structure | `select()`, `project()`, `map()`, `aggregate()` |
| [Expressions](/queries/expressions) | Typed database calculations | `expr`, `project()`, expression callbacks |
| [Aggregate](/queries/aggregate) | Summarize data | `groupBy()`, `count()`, `sum()`, `avg()` |
| [Order](/queries/order) | Control result ordering/size | `orderBy()`, `limit()`, `offset()` |
| [Temporal](/queries/temporal) | Time-based queries | `temporal()` |
| [Compose](/queries/compose) | Reusable query parts | `pipe()`, `createFragment()` |
| [Combine](/queries/combine) | Set operations | `union()`, `intersect()`, `except()` |
| [Execute](/queries/execute) | Run and retrieve | `execute()`, `first()`, `count()`, `exists()`, `paginate()`, `stream()`, `batch()` |
## Query Flow
A typical query follows this flow:
```text
Source → Filter → Traverse → Filter → Shape → Order → Execute
↑__________________|
(repeat as needed)
```
Each step is optional except Source and Execute. You can filter, traverse, and filter again as many
times as needed before shaping and executing.
## Basic Example
```typescript
const results = await store
.query()
.from("Person", "p") // Source
.whereNode("p", (p) => p.status.eq("active")) // Filter
.traverse("worksAt", "e") // Traverse
.to("Company", "c") // Traverse (target)
.whereNode("c", (c) => c.industry.eq("Tech")) // Filter
.select((ctx) => ({ // Shape
person: ctx.p.name,
company: ctx.c.name,
role: ctx.e.role,
}))
.orderBy("p", "name", "asc") // Order
.limit(50) // Order
.execute(); // Execute
```
## Type Safety
The query builder is fully typed. TypeScript infers result types based on your schema and selection:
```typescript
// TypeScript infers: Array<{ name: string; email: string | undefined }>
const results = await store
.query()
.from("Person", "p")
.select((ctx) => ({
name: ctx.p.name, // string (required in schema)
email: ctx.p.email, // string | undefined (optional in schema)
}))
.execute();
// Invalid property access is caught at compile time:
.select((ctx) => ({
invalid: ctx.p.nonexistent, // TypeScript error!
}))
```
For new database-side projections and calculations, use typed
[database expressions](/queries/expressions). `project()` compiles its callback to SQL, while
`map()` transforms decoded rows in JavaScript. Existing `select()` callbacks retain their
compatibility behavior.
## When to Use Queries vs Store API
**Use the query builder** when you need:
- Filtering based on node properties
- Traversing relationships between nodes
- Aggregating data across multiple nodes
- Complex predicates with AND/OR logic
**Use the [Store API](/schemas-stores#store-api)** for simple operations:
- Get a node by ID
- Create a new node
- Update a node's properties
- Delete a node
## Predicates Reference
Predicates are the building blocks for filtering. Each data type has its own set of predicates:
| Type | Documentation |
|------|--------------|
| String | [String Predicates](/queries/predicates/#string) |
| Number | [Number Predicates](/queries/predicates/#number) |
| Date | [Date Predicates](/queries/predicates/#date) |
| Array | [Array Predicates](/queries/predicates/#array) |
| Object | [Object Predicates](/queries/predicates/#object) |
| Embedding | [Embedding Predicates](/queries/predicates/#embedding) |
## Performance Tips
### Filter Early
Apply predicates as early as possible to reduce the working set:
```typescript
// Good: Filter at source
store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.active.eq(true))
.traverse("worksAt", "e")
.to("Company", "c");
// Less efficient: Filter after traversal
store
.query()
.from("Person", "p")
.traverse("worksAt", "e")
.to("Company", "c")
.whereNode("p", (p) => p.active.eq(true));
```
### Be Specific with Kinds
Unless you need subclass expansion, use exact kinds:
```typescript
// More efficient: Exact kind
.from("Podcast", "p")
// Less efficient: Includes all subclasses
.from("Media", "m", { includeSubClasses: true })
```
### Always Paginate Large Results
```typescript
const page = await store
.query()
.from("Event", "e")
.orderBy("e", "date", "desc")
.limit(100)
.select((ctx) => ctx.e)
.execute();
```
## Next Steps
Start with the fundamentals:
1. [Source](/queries/source) - Starting queries with `from()`
2. [Filter](/queries/filter) - Reducing results with predicates
3. [Traverse](/queries/traverse) - Navigating relationships
4. [Shape](/queries/shape) - Transforming output with `select()`
# Temporal
> Time-based queries with temporal()
TypeGraph tracks temporal validity for all nodes and edges. Use temporal queries to view the graph
at a point in time, audit changes, or access historical data.
## Temporal Modes
The `temporal()` method controls which versions of data are returned:
| Mode | Description |
|------|-------------|
| `"current"` | Only currently valid data (default behavior) |
| `"asOf"` | Data as it existed at a specific timestamp |
| `"includeEnded"` | All versions, including historical |
| `"includeTombstones"` | All versions, including soft-deleted |
## Current State (Default)
By default, queries return only currently valid, non-deleted data:
```typescript
// Returns only current, non-deleted nodes
const currentPeople = await store
.query()
.from("Person", "p")
.select((ctx) => ctx.p)
.execute();
```
This is equivalent to:
```typescript
.temporal("current")
```
## Point-in-Time Queries (asOf)
Query the graph as it existed at a specific moment:
```typescript
const yesterday = new Date(Date.now() - 24 * 60 * 60 * 1000).toISOString();
const pastState = await store
.query()
.from("Article", "a")
.temporal("asOf", yesterday)
.whereNode("a", (a) => a.id.eq(articleId))
.select((ctx) => ctx.a)
.execute();
```
This returns nodes and edges that were valid at the specified timestamp, even if they've since been updated or deleted.
### Use Cases for asOf
- **Auditing**: See what data looked like at a specific time
- **Debugging**: Reproduce issues by querying historical state
- **Compliance**: Generate point-in-time reports
- **Recovery**: Find old values before an erroneous update
```typescript
// What did the user's profile look like last week?
const lastWeek = new Date(Date.now() - 7 * 24 * 60 * 60 * 1000).toISOString();
const historicalProfile = await store
.query()
.from("User", "u")
.temporal("asOf", lastWeek)
.whereNode("u", (u) => u.id.eq(userId))
.select((ctx) => ctx.u)
.first();
```
## Shared-Coordinate Views (store.asOf)
`.temporal("asOf", T)` pins a single query. When several reads should share one
temporal coordinate, pin it once with `store.asOf(T)` and reuse the returned
**read-only view** — TypeGraph's as-of database value, in the style of Datomic
`(d/as-of db t)` and SQL:2011 `FOR SYSTEM_TIME AS OF`.
```typescript
const past = store.asOf("2024-01-01T00:00:00.000Z");
// Every read on `past` observes the graph as it was valid at that instant.
const alice = await past.nodes.Person.getById(aliceId);
const jobs = await past.edges.worksAt.findFrom(alice);
const peers = await past.reachable(aliceId, { edges: ["knows"] });
const team = await past.subgraph(aliceId, { edges: ["reportsTo"] });
const names = await past
.query()
.from("Person", "p")
.whereNode("p", (p) => p.department.eq("Engineering"))
.select((ctx) => ctx.p.name)
.execute();
```
The view pins the `nodes` / `edges` collections (`getById`, `getByIds`, `find`,
`count`, `findFrom`, `findTo`), `query()`, `subgraph()`, and the graph
algorithms (`reachable`, `canReach`, `shortestPath`, `neighbors`, `degree`).
For the other modes, use `store.view({ mode, asOf })`:
```typescript
// A view over every version, including soft-deleted ones.
const audit = store.view({ mode: "includeTombstones" });
const everyVersion = await audit.nodes.Document.find();
```
A view is **read-only**: writes stay on the live `store`, and a view collection
rejects `create` / `update` / `delete` with a `ConfigurationError`. `search` is
refused on a non-`"current"` view (the fulltext / vector index reflects current
state only). `asOf` must be a canonical UTC ISO-8601 timestamp
(`YYYY-MM-DDTHH:mm:ss.sssZ`).
See the [`store.asOf` / `store.view`
reference](/schemas-stores#temporal-views-storeasof-and-storeview) for the full
surface.
## Recorded Time (Bitemporal)
The modes above query **valid time** — *when a fact was true in the world*
(`validFrom` / `validTo`). Recorded time (also called **system time**) is the
second axis — *when a fact was recorded by TypeGraph*. With the built-in
captured relation, TypeGraph can run **bitemporal graph reads** for
TypeGraph-managed writes: you can ask "what did TypeGraph reconstruct as true,
as of a captured commit instant?" — including seeing values that were later
corrected.
Recorded-time capture is **opt-in** per store, because it writes a history row
for every committed TypeGraph collection change:
```typescript
const store = createStore(graph, backend, { history: true });
```
With `history: true`, every committed TypeGraph node/edge write is captured into
recorded-time relations (`typegraph_recorded_nodes` /
`typegraph_recorded_edges`) stamped with a per-graph monotonic commit instant.
Enable it on a **fresh graph**: there is no backfill, so an entity that already
exists is first recorded the next time it is written through TypeGraph. Capture
requires a transactional backend with statement execution (the built-in SQLite /
PostgreSQL backends).
Advanced hosts can bind an already-populated recorded relation for reads without
using TypeGraph's writer wrapper:
```typescript
import { createSqlSchema, recordedRelation } from "@nicia-ai/typegraph";
const recordedRead = recordedRelation({
schema: createSqlSchema({
recordedNodes: "audit_nodes",
recordedEdges: "audit_edges",
}),
});
const store = createStore(graph, backend, { recordedRead });
```
That option only supplies the read source for `asOfRecorded(T)` reconstruction.
It does not capture writes, advance TypeGraph's recorded clock, or make
`store.recordedNow()` available. If TypeGraph should own capture, use
`history: true`. `recordedRead` must be created by `recordedRelation({ schema })`
with a `createSqlSchema(...)` schema; the store validates those factory
descriptors at runtime and rejects combining them with `history: true`.
### Reading at a recorded instant
`store.asOfRecorded(T)` reconstructs the graph as TypeGraph recorded it at
instant `T`. `T` is a `RecordedInstant`: a branded, versioned string containing
both a per-graph logical revision and a physical wall-time high-water mark. It
originates from `store.recordedNow()` (below), or from
`asRecordedInstant(...)` when an anchor previously returned by TypeGraph has
round-tripped through untyped storage:
```typescript
import {
asRecordedInstant,
recordedInstantWallTime,
} from "@nicia-ai/typegraph";
const recorded = store.asOfRecorded(
asRecordedInstant("r1:0000000000000042:2024-06-01T12:00:00.000Z"),
);
const doc = await recorded.nodes.Document.getById(docId);
const cited = await recorded.edges.cites.getByIds(citationIds);
const reachable = await recorded.reachable(docId, { edges: ["cites"] });
```
Recorded collections also expose `scan()` for complete snapshot reconstruction.
Each call returns at most 1,000 entities in canonical `id` order; use the opaque
`nextCursor` to continue without retaining a separate identity inventory:
```typescript
const first = await recorded.nodes.Document.scan({ limit: 500 });
const second =
first.nextCursor === undefined ?
undefined
: await recorded.nodes.Document.scan({
limit: 500,
after: first.nextCursor,
});
const citations = await recorded.edges.cites.scan({ limit: 500 });
```
Scan cursors are forward-only and bound to the graph, entity kind, and both
temporal coordinates. Passing a cursor to another collection or recorded-time
view throws a `ValidationError` instead of silently skipping data. Iterate each
declared node and edge kind to reconstruct a complete historical graph snapshot.
A raw wall-clock string — `store.asOfRecorded(new Date().toISOString())` — does
**not** type-check, by design. Wall time does not identify which commit to read
when several commits share a millisecond. The anchor's logical revision provides
that order; its ISO component records a non-decreasing physical wall-time
high-water mark. To pin "as things stand right now" deterministically, use
`store.recordedNow()` (the recorded high-water mark), then guard the `undefined`
case before passing it to `store.asOfRecorded()`.
```typescript
await store.nodes.Document.update(docId, { title: "Revised" });
const checkpoint = await store.recordedNow(); // a stable anchor for this state
if (checkpoint === undefined) throw new Error("expected a recorded checkpoint");
console.log(recordedInstantWallTime(checkpoint)); // canonical UTC wall time
// ...later, however much the graph has changed:
const asOfCheckpoint = store.asOfRecorded(checkpoint);
```
`recordedNow()` is **graph-global**, not scoped to any one caller or write. It is
the single high-water mark for the whole graph, advanced by *every* committed
capture from *any* writer. So a change in `recordedNow()` across two reads means
"something committed to this graph in between" — **not** "the write I just made
landed." Do not use a `recordedNow()` advance as a per-writer "did my write
succeed?" signal: under any concurrent writer to the same graph it both misses
dropped writes (another writer moved the clock) and misfires on no-op writes. To
confirm a specific write committed, observe the write itself (e.g. its return
value, or run it inside `store.transaction(...)` and act on success), not the
global clock.
#### Logical revision and physical time
The canonical encoding is
`r1:<16-digit revision>:`. Revisions are strict and
monotonic within one graph. The physical component is sampled from the
application clock and clamped to the previous anchor only when that clock moves
backward. It may repeat, but never decreases. TypeGraph does not add one
millisecond per commit, so throughput cannot push recorded wall time beyond the
greatest wall time the graph has actually observed. After a backward clock
correction, the component remains at its prior high-water mark until wall time
catches up.
This non-decreasing physical component preserves cumulative diagonal replay for
default validity timestamps: a later recorded anchor cannot pin valid time
before an earlier commit's default `valid_from`.
The fixed-width revision prefix makes anchors lexicographically sortable within
a graph and gives each captured transaction a distinct addressable state. Use
`compareRecordedInstants(a, b)` rather than manually comparing strings, and
only compare anchors from the same graph. Recorded relations store the revision
as an integer, so their open interval ceiling is independent of the `r1` API
encoding and PostgreSQL range scans do not depend on text collation. Recorded
clocks remain per graph, and TypeGraph does not provide one cross-graph recorded
anchor.
Batch related writes in `store.transaction(...)`: one transaction allocates one
recorded instant. For event logs, align transactions with durable replay or
checkpoint boundaries, and cap transaction size separately so an initial sync
does not hold a write lock or capture buffer without bound.
Direct `store.asOfRecorded(T)` is **diagonal** bitemporal sugar: it uses the
anchor's logical revision for the recorded-time axis and its physical wall-time
component for the valid-time axis. To pin the two axes independently — *what was
valid at one instant, as TypeGraph captured it at another* — chain from a
valid-time view:
```typescript
// The state valid on Jan 1, as TypeGraph recorded it on Jun 1
// (e.g. after a correction was entered later).
const corrected = store
.asOf("2024-01-01T00:00:00.000Z")
.asOfRecorded(
asRecordedInstant("r1:0000000000000042:2024-06-01T12:00:00.000Z"),
);
const asKnownThen = await corrected.nodes.Invoice.getById(invoiceId);
```
Use `recordedInstantRevision(T)` for diagnostics and
`recordedInstantWallTime(T)` for display or logging. Do not split the versioned
anchor string manually.
`store.view({ mode }).asOfRecorded(T)` composes recorded time with any
valid-time mode — e.g. `includeTombstones` to reconstruct soft-deleted rows at a
recorded instant.
### The recorded view surface
A `RecordedStoreView` is a **narrow, reconstructing** read lens. It exposes only
reads that can be faithfully rebuilt from the recorded relations:
- **Point reads** — `nodes..getById` / `getByIds`, and the edge equivalents
- **`query()`** — a sealed query builder over the recorded relations
- **`subgraph()`** and the graph algorithms — `reachable`, `canReach`,
`shortestPath`, `degree`
Broad collection reads (`find` / `count` / `findFrom` / …), `search`, and
fulltext / vector predicates are **refused** with a `ConfigurationError` /
`UnsupportedPredicateError`: the fulltext and vector indexes reflect *current*
state only, so they cannot answer a recorded-time question. `T` must use the
canonical versioned RecordedInstant encoding; a plain ISO timestamp is rejected.
:::caution[Preview-schema migration]
Timestamp-only anchors and recorded tables created by the initial preview need
an explicit offline migration. Run `migrateLegacyRecordedTime({ backend })`
before opening the upgraded store, then translate externally persisted
checkpoints with `migrateRecordedAnchor({ backend, graphId, anchor })`. See
[Migrating preview recorded time](/schema-management#migrating-preview-recorded-time).
:::
### Engine-native recorded time
Everything above describes **TypeGraph-owned** recorded time: `history: true`
captures into TypeGraph's own recorded relations and clock. A backend can
instead track recorded time itself — declare `EngineProvisioning.recordedTime`
on it — and `history: true` then reads and writes through the engine's own
temporal storage; TypeGraph's capture relations, clock, and write-fence-gated
clock allocation are never engaged.
Which ownership a store reads under is **derived**, never an option you set:
it is `"engine-native"` exactly when the backend declares `recordedTime`,
`"typegraph-relations"` otherwise. Neither bundled SQLite nor PostgreSQL
profile declares it, so every example on this page runs under
`"typegraph-relations"` as shown; see [Supplying
`recordedTime`](/backend-authoring#supplying-recordedtime) for what a
third-party engine implements to opt in.
Under engine-native ownership:
- `store.recordedNow()`, `store.revisionNow()`, and `TransactionReceipt.recorded`
all come from the engine's own revision instead of TypeGraph's clock — one
call per transaction, not per graph. `TransactionReceipt.recorded` is
stamped only when a graph node/edge/identity write inside the transaction
actually changed a row — a delete of a missing id, an
`insertNodeIfAbsent` that found the row, and a coalesced no-op upsert all
leave it `undefined`, matching a read-only transaction. A transaction whose
only effect is a raw `tx.sql` statement also leaves it `undefined` even
though the engine's revision advances underneath it; use a graph collection
write when you need `receipt.recorded` to reflect the change. To observe
those writes, a receipted engine-native transaction routes every write
through an observing wrapper, so `transactionWithReceipt` does not use
session-scoped atomic batching where a plain `transaction` on the same
store would.
- `RecordedInstant` anchors use the engine form
`e1::` rather than
`r1:<16-digit revision>:`. The revision is an opaque,
engine-assigned token, never parsed as a number, so ordering two `e1:`
anchors (`compareRecordedInstants`) falls back to the timestamp component
only — two engine revisions minted within the same millisecond compare
equal even though they are distinct commits, unlike a `r1:` anchor's strict
per-commit counter. `recordedInstantWallTime(instant)` works for either
form; `recordedInstantRevision(instant)` throws for an `e1:` anchor, since
there is no TypeGraph numeric revision to return.
- `store.asOfRecorded(instant)` requires an instant minted under the SAME
store's own ownership form. An engine-native store refuses an `r1:`
instant, and a TypeGraph-owned store refuses an `e1:` instant, both with a
`ConfigurationError` (`RECORDED_INSTANT_OWNERSHIP_MISMATCH`) — an anchor
from one ownership form is never valid against the other, even against a
different store over the same data.
- No recorded relation is read or written. (A profile built on the bundled
schema factories still creates the recorded tables as part of its base DDL
— they just stay empty.) `recordedRead: recordedRelation({ schema })`
(above) and `migrateLegacyRecordedTime` are both refused: neither has a
TypeGraph-owned recorded relation to bind or migrate.
- `revisionTracking: true` is refused whether or not `history: true` is also
requested — there is no TypeGraph clock for it to advance; the engine's
own revision is the only tracking engine-native has, and it is available
only under `history: true`.
- Reconstructing identity at a recorded coordinate — `store.identityAtCoordinate`
at a past instant, and any query that reaches the historical identity
traversal — is refused: identity history reads TypeGraph's own recorded
relations directly, which an engine-native backend does not populate. Read
identity at the current coordinate instead, or use a TypeGraph-owned store
for historical identity reconstruction.
Everything else on this page — `asOfRecorded`'s diagonal composition with
`asOf`, the recorded view surface's read shape, `includeTombstones`
composition — behaves the same under either ownership form; only the anchor
grammar, the write mechanics, and the refusals above differ. See [Lineage and
pruned diffs](/graph-merge#lineage-and-pruned-diffs) for how graph-merge
derives a change delta under engine-native ownership — from the engine's own
`lineage`, never from recorded relations, since none exist to derive one
from.
### Writing with history enabled
Capture flushes at transaction commit, so writes must go through the store's
typed collections — use `store.transaction(...)` as usual:
```typescript
await store.transaction(async (tx) => {
await tx.nodes.Document.create({ title: "Draft" });
});
```
#### Raw SQL under history capture
The portable `HistoryStore` exposes neither raw SQL nor caller-owned
transaction adoption. If the store was deliberately created through
`createAdapterStore(..., { history: true })`, raw `tx.sql` is still disabled
(it would bypass capture), and `store.withTransaction(externalTx)` is replaced
by the callback form
`store.withRecordedTransaction(externalTx, async (tx) => { ... })`, which gives
capture a flush point before your transaction commits. Out-of-band database
writes and row-returning raw SQL paths are not audited by the built-in capture
wrapper; use TypeGraph collection writes when the recorded relation is the
source of truth.
The adapter history store's `.backend` is a runtime and type-level
`HistoryStoreBackend` projection. Capture-wrapped graph reads and writes remain
available. `executeRaw`, `executeStatement`, `executeDdl`, `trustedImport`,
`clearGraph`, and nested `transaction` are absent because each can mutate live
rows without a corresponding capture flush. The full guarded backend remains
internal to TypeGraph's query and transaction implementation.
`store.withTransaction` on a history-enabled store is a **compile error** (the
`externalTx` argument is rejected with a message naming
`withRecordedTransaction`); the runtime guard still throws `ConfigurationError`
if suppressed. Inside an `AdapterHistoryStore.transaction(...)`, the typed
context omits `tx.sql`, and `tx.sqlAvailability` reports `"history"` (or
`"revisionTracking"`) so portable code can branch without touching the runtime
guard. Suppressed JavaScript or TypeScript access still throws — see the
`tx.sqlAvailability` guidance in
[Cross-Store Transactions](/recipes/).
Both guards carry a branchable `details.code`; see
[Recorded-capture guard codes](/errors/#recorded-capture-guard-codes).
To write your own relational tables atomically with graph writes on a history
store, pass your transaction handle to `withRecordedTransaction` and write your
tables through **that** handle (not `tx.sql`):
```typescript
await db.transaction(async (pgTx) => {
const { receipt } = await store.withRecordedTransaction(pgTx, async (tx) => {
await tx.nodes.Document.update(documentId, props); // graph write
});
await pgTx.insert(streamCursors).values(cursorRow); // your own table
}); // one COMMIT / ROLLBACK across both layers
```
`withRecordedTransaction` returns a
[`TransactionOutcome`](/schemas-stores/#transaction-receipts): destructure
`{ result, receipt }`. `receipt.writes` counts the graph writes (drop
detection) and `receipt.recorded` is this transaction's recorded commit instant
— the per-transaction replay anchor. When the callback runs user code that also
bookkeeps, scope a sub-receipt with `tx.measure((scoped) => ...)`: writes through
the `scoped` context are attributed to the sub-receipt, while the surrounding
bookkeeping written through `tx` is not.
This is separate from `recordedRead`: a store created with a `recordedRead`
binding can reconstruct from a relation populated by another system, but
TypeGraph is not responsible for making that relation complete or atomic with
live writes.
#### Write cost: batch under `history: true`
Each **un-batched** write under `history: true` becomes its own transaction —
it allocates a recorded commit instant under a per-graph clock lock and flushes
one history row at commit. So a tight loop of single `create`/`update`/`delete`
calls pays that fixed cost once per call. Wrapping the same writes in one
`store.transaction(...)` allocates **one** recorded instant for the whole batch
and amortizes the overhead to roughly nothing.
Measured per-op latency, identical workload with capture off vs on (history
off → on; N = 400; reproduce with
`pnpm --filter @nicia-ai/typegraph-benchmarks bench:recorded-write`):
| Workload | SQLite | PostgreSQL |
| ------------------------------ | -----: | ---------: |
| create — un-batched (per op) | ~2.5× | ~5.5× |
| create — **batched in one txn** | ~1.5× | ~1.0× |
| update — un-batched (per op) | ~2.8× | ~6× |
| soft delete — un-batched | ~1.7× | ~1.9× |
The takeaway: capture is opt-in and cheap when you batch. Under `history: true`,
prefer `store.transaction(...)` for bulk writes; a loop of individual
collection writes is the one pattern that pays the per-write multiple. (Stores
created without `history: true` are unaffected — graph writes never touch the
capture path.) Batching also reduces recorded-clock consumption: one captured
transaction advances the per-graph clock once, even when it contains many
writes. See [Logical revision and physical time](#logical-revision-and-physical-time)
for the anchor format.
> **Performance.** Recorded reads reconstruct from the history relations rather
> than the live tables, so they are slower than current-state reads — most
> noticeably for full-graph `subgraph` / algorithm reconstructions on
> PostgreSQL. Reach for `asOfRecorded` for audit and point-in-time
> reconstruction, not hot-path reads.
## Including Historical Data (includeEnded)
View all versions, including superseded records:
```typescript
const history = await store
.query()
.from("Article", "a")
.temporal("includeEnded")
.whereNode("a", (a) => a.id.eq(articleId))
.orderBy((ctx) => ctx.a.validFrom, "desc")
.select((ctx) => ({
title: ctx.a.title,
validFrom: ctx.a.validFrom,
validTo: ctx.a.validTo,
version: ctx.a.version,
}))
.execute();
// Result shows all versions:
// [
// { title: "Final Title", validFrom: "2024-03-01", validTo: undefined, version: 3 },
// { title: "Draft v2", validFrom: "2024-02-15", validTo: "2024-03-01", version: 2 },
// { title: "Initial Draft", validFrom: "2024-02-01", validTo: "2024-02-15", version: 1 },
// ]
```
### Audit Trail
Build a complete change history:
```typescript
async function getAuditTrail(nodeId: string) {
return store
.query()
.from("Document", "d")
.temporal("includeEnded")
.whereNode("d", (d) => d.id.eq(nodeId))
.select((ctx) => ({
version: ctx.d.version,
title: ctx.d.title,
status: ctx.d.status,
validFrom: ctx.d.validFrom,
validTo: ctx.d.validTo,
updatedAt: ctx.d.updatedAt,
}))
.orderBy("d", "version", "asc")
.execute();
}
```
## Including Soft-Deleted Data (includeTombstones)
Include records that have been soft-deleted:
```typescript
const allIncludingDeleted = await store
.query()
.from("User", "u")
.temporal("includeTombstones")
.select((ctx) => ({
id: ctx.u.id,
name: ctx.u.name,
deletedAt: ctx.u.deletedAt, // Will have a value for deleted records
}))
.execute();
```
### Filtering Deleted Records
```typescript
// Find only deleted records
const deletedUsers = await store
.query()
.from("User", "u")
.temporal("includeTombstones")
.whereNode("u", (u) => u.deletedAt.isNotNull())
.select((ctx) => ({
id: ctx.u.id,
name: ctx.u.name,
deletedAt: ctx.u.deletedAt,
}))
.execute();
```
## Temporal Metadata Fields
When querying with temporal context, these fields are available:
| Field | Type | Description |
|-------|------|-------------|
| `validFrom` | `string \| undefined` | When this version became valid (`undefined` on an **open-left** row — see below) |
| `validTo` | `string \| undefined` | When this version was superseded (undefined if current) |
| `createdAt` | `string` | When the node was first created |
| `updatedAt` | `string` | When this version was written |
| `deletedAt` | `string \| undefined` | Soft-delete timestamp (undefined if not deleted) |
| `version` | `number` | Optimistic concurrency version number |
### Open-left rows (`validFrom` is `undefined`)
A row may have **no lower bound at all**, which means "valid since forever, as
far as this store knows". `asOf` and `current` treat such a row as valid at
every instant strictly before its `validTo`, or every instant if it has no end.
These writes produce one:
- a Store create or resurrecting upsert stating `validFrom: null`;
- an interchange record stating `validFrom: null` — a source row confirmed to
have no lower bound, round-tripped rather than re-stamped;
- a **born-already-ended** write: one that CREATES a row, or RESETS its window,
while stating a `validTo` at or before its own instant and no `validFrom`. The
row's start is unknown rather than after its end, so no bound is stored and the
row reads back at every `asOf` before that end. A `validTo` in the *future* is
unaffected — it still stamps the write instant, so the row stays invisible at
instants before it existed. Every **node** path that resets the window
qualifies, and reaches the same stored shape: a create on a fresh id, a create
on a tombstoned one, and a resurrecting `upsertById` / `bulkUpsertById`. An
**edge** never does: an edge create cannot land on a tombstone (a taken id
raises `Edge already exists`), and the two paths that resurrect one —
`bulkUpsertById` and `getOrCreateByEndpoints` — RETAIN the bound the row
carries and judge the stated `validTo` against it.
#### Rows written by older versions
Before that rule existed, a born-already-ended write stored the write instant as
`valid_from`, leaving a window that runs backwards — a row readable at **no**
coordinate at all. Upgrading does not rewrite such rows; they keep their window
and stay invisible until an operator repairs them explicitly with
`repairInvertedValidityWindows`, which normalizes them to the open-left shape
above. Prefer `relations: "live-and-recorded"`: repairing only the live axis
leaves the recorded twin inverted, so `asOfRecorded` reads keep returning the
invisible shape. See
[Repairing inverted validity windows](/schema-management#repairing-inverted-validity-windows)
for the operator checklist — run it with writers stopped, and re-baseline merge
branches afterwards.
```typescript
.select((ctx) => ({
...ctx.a, // All node properties
validFrom: ctx.a.validFrom,
validTo: ctx.a.validTo,
createdAt: ctx.a.createdAt,
updatedAt: ctx.a.updatedAt,
deletedAt: ctx.a.deletedAt,
version: ctx.a.version,
}))
```
## Temporal Traversals
Temporal modes apply to traversals as well:
```typescript
// See who worked at a company last year
const lastYear = new Date("2023-01-01").toISOString();
const pastEmployees = await store
.query()
.from("Company", "c")
.temporal("asOf", lastYear)
.whereNode("c", (c) => c.name.eq("Acme Corp"))
.traverse("worksAt", "e", { direction: "in" })
.to("Person", "p")
.select((ctx) => ({
name: ctx.p.name,
role: ctx.e.role,
}))
.execute();
```
`store.subgraph()` and `store.algorithms.*` accept the same `temporalMode`
and `asOf` options, defaulting to `graph.defaults.temporalMode`. See
[Temporal Behavior](/graph-algorithms#temporal-behavior) for the algorithm
surface and [`store.subgraph()` options](/schemas-stores#storesubgraphrootid-options)
for subgraph.
## Real-World Examples
### Version Comparison
Compare two versions of a document:
```typescript
async function compareVersions(docId: string, v1: number, v2: number) {
const versions = await store
.query()
.from("Document", "d")
.temporal("includeEnded")
.whereNode("d", (d) => d.id.eq(docId))
.select((ctx) => ctx.d)
.execute();
const version1 = versions.find((v) => v.version === v1);
const version2 = versions.find((v) => v.version === v2);
return { version1, version2 };
}
```
### Compliance Reporting
Generate a report as of a specific date:
```typescript
async function generateQuarterlyReport(quarterEnd: string) {
const activeContracts = await store
.query()
.from("Contract", "c")
.temporal("asOf", quarterEnd)
.whereNode("c", (c) => c.status.eq("active"))
.traverse("belongsTo", "e")
.to("Customer", "cust")
.select((ctx) => ({
contractId: ctx.c.id,
value: ctx.c.value,
customer: ctx.cust.name,
}))
.execute();
return {
asOf: quarterEnd,
totalContracts: activeContracts.length,
totalValue: activeContracts.reduce((sum, c) => sum + c.value, 0),
contracts: activeContracts,
};
}
```
### Undo/Recovery
Find the previous value before an update:
```typescript
async function getPreviousVersion(nodeId: string) {
const versions = await store
.query()
.from("Document", "d")
.temporal("includeEnded")
.whereNode("d", (d) => d.id.eq(nodeId))
.select((ctx) => ctx.d)
.orderBy("d", "version", "desc")
.limit(2)
.execute();
return {
current: versions[0],
previous: versions[1],
};
}
```
## Next Steps
- [Filter](/queries/filter) - Filtering with predicates
- [Traverse](/queries/traverse) - Graph traversals
- [Execute](/queries/execute) - Running queries
- [Bitemporal Time Travel](/examples/bitemporal-time-travel) - Valid time plus
recorded time in one runnable example
- [Agent Decision Replay](/examples/agent-decision-replay) - Reconstruct the
exact graph an agent saw
- [Breach Forensics](/examples/breach-forensics) - Traverse a reconstructed
access graph at the breach instant
# Troubleshooting
> Solutions to common issues and frequently asked questions
This guide covers common issues and their solutions when working with TypeGraph.
## Installation Issues
### "Cannot find module '@nicia-ai/typegraph'"
**Cause:** Package not installed or using wrong package name.
**Solution:**
```bash
npm install @nicia-ai/typegraph zod drizzle-orm
```
### "better-sqlite3 compilation failed"
**Cause:** Native module compilation requires build tools.
**Solutions:**
**macOS:**
```bash
xcode-select --install
```
**Ubuntu/Debian:**
```bash
sudo apt-get install build-essential python3
```
**Windows:**
```bash
npm install --global windows-build-tools
```
**Alternative:** Use `sql.js` for pure JavaScript SQLite (no compilation needed).
### Missing optional `drizzle-orm` peer
**Cause:** The managed SQLite or PGlite Store entrypoint was called without the optional
`drizzle-orm` peer installed.
**Solution:** Install the peer in the application that uses the managed entrypoint:
```bash
npm install drizzle-orm
```
The root package and other portable entrypoints do not require Drizzle. Explicit
`@nicia-ai/typegraph/adapters/drizzle/...` entrypoints load Drizzle when the module is evaluated,
so a missing peer there appears as the runtime's raw module-resolution error instead of
`MISSING_PEER_DEPENDENCY`. See [Managed Store Entrypoints](/backend-setup#managed-store-entrypoints).
### "Module not found: drizzle-orm/better-sqlite3"
**Cause:** An explicit Drizzle adapter import is missing `drizzle-orm`, or the application imported
the wrong Drizzle subpath.
**Solution:** First install `drizzle-orm`, then ensure the import matches the driver:
```bash
npm install drizzle-orm
```
```typescript
// Correct
import { drizzle } from "drizzle-orm/better-sqlite3";
// Incorrect
import { drizzle } from "drizzle-orm";
```
## Schema Definition Errors
### "Node schema contains reserved property names"
**Cause:** Using reserved keys (`id`, `kind`, `meta`) in your Zod schema.
**Solution:** Rename your properties:
```typescript
// Bad - 'id' is reserved
const User = defineNode("User", {
schema: z.object({
id: z.string(), // Error!
name: z.string(),
}),
});
// Good - use a different name
const User = defineNode("User", {
schema: z.object({
externalId: z.string(),
name: z.string(),
}),
});
```
TypeGraph automatically provides `id`, `kind`, and `meta` on all nodes.
### "Edge type already has constraints defined"
**Cause:** Defining `from`/`to` constraints on both the edge type and graph registration.
**Solution:** Define constraints in one place only:
```typescript
// Option 1: On the edge type (reusable across graphs)
const worksAt = defineEdge("worksAt", {
from: [Person],
to: [Company],
});
const graph = defineGraph({
edges: {
worksAt: { type: worksAt }, // No from/to here
},
});
// Option 2: On the graph (flexible per-graph)
const worksAt = defineEdge("worksAt");
const graph = defineGraph({
edges: {
worksAt: { type: worksAt, from: [Person], to: [Company] },
},
});
```
## Runtime Errors
### ValidationError: "Invalid input"
**Cause:** Data doesn't match the Zod schema.
**Solution:** Check the error details for specific issues:
```typescript
try {
await store.nodes.Person.create({ name: "" });
} catch (error) {
if (error instanceof ValidationError) {
console.log(error.details.issues); // Zod issues array
}
}
```
### NodeNotFoundError
**Cause:** Attempting to read/update/delete a non-existent node.
**Solution:** Check if the node exists first or handle the error:
```typescript
const node = await store.nodes.Person.getById(someId);
if (!node) {
// Handle missing node
}
// Or use error handling
try {
await store.nodes.Person.update(someId, { name: "New" });
} catch (error) {
if (error instanceof NodeNotFoundError) {
console.log(`Node ${error.details.id} not found`);
}
}
```
### RestrictedDeleteError
**Cause:** Attempting to delete a node that has edges, with `onDelete: "restrict"` (the default).
**Solution:** Either delete the edges first or use a different delete behavior:
```typescript
// Option 1: Delete edges first. Include ended-but-not-deleted edges if you
// are cleaning up historical validity windows too.
const edges = await store.edges.worksAt.findFrom(person, {
temporalMode: "includeEnded",
});
for (const edge of edges) {
await store.edges.worksAt.delete(edge.id);
}
await store.nodes.Person.delete(person.id);
// Option 2: Use cascade delete in schema
const graph = defineGraph({
nodes: {
Person: { type: Person, onDelete: "cascade" },
},
});
```
### DisjointError
**Cause:** Creating a node with an ID that's already used by a disjoint type.
**Solution:** Ensure IDs are unique across disjoint types or don't use explicit IDs:
```typescript
// If Person and Organization are disjoint:
// Bad - same ID for different types
await store.nodes.Person.create({ name: "Alice" }, { id: "entity-1" });
await store.nodes.Organization.create({ name: "Acme" }, { id: "entity-1" }); // Error!
// Good - let TypeGraph generate unique IDs
await store.nodes.Person.create({ name: "Alice" });
await store.nodes.Organization.create({ name: "Acme" });
```
## Query Issues
### "Alias 'x' is already in use"
**Cause:** Using the same alias twice in a query.
**Solution:** Use unique aliases:
```typescript
// Bad
store.query().from("Person", "p").traverse("knows", "e").to("Person", "p"); // Error! 'p' already used
// Good
store.query().from("Person", "p1").traverse("knows", "e").to("Person", "p2");
```
### Empty results when expecting data
**Causes and solutions:**
1. **Type mismatch:** Ensure you're querying the correct node type
```typescript
// Check the node type name matches exactly
.from("Person", "p") // Must match defineNode("Person", ...)
```
2. **Missing includeSubClasses:** When querying a superclass
```typescript
.from("Content", "c", { includeSubClasses: true })
```
3. **Strict predicate:** Check your filters aren't too restrictive
```typescript
// Debug by removing filters temporarily
const all = await store
.query()
.from("Person", "p")
.select((c) => c.p)
.execute();
console.log(all.length); // How many total?
```
### Slow queries
**Solutions:**
1. **Use the query profiler:**
```typescript
import { QueryProfiler } from "@nicia-ai/typegraph/profiler";
const profiler = new QueryProfiler();
profiler.attachToStore(store);
// Run your queries...
const report = profiler.getReport();
console.log(report.recommendations);
```
2. **Add indexes** based on profiler recommendations:
```typescript
import { defineNodeIndex } from "@nicia-ai/typegraph/indexes";
const nameIndex = defineNodeIndex(Person, { fields: ["name"] });
```
3. **Limit results:**
```typescript
.limit(100)
// Or use pagination
.paginate({ first: 20 })
```
## Database Connection Issues
### "Database is locked" (SQLite)
**Cause:** Multiple processes accessing the same SQLite file without WAL mode.
**Solution:** Enable WAL mode:
```typescript
const sqlite = new Database("myapp.db");
sqlite.pragma("journal_mode = WAL");
```
### Connection pool exhausted (PostgreSQL)
**Cause:** Too many concurrent connections.
**Solution:** Configure pool limits:
```typescript
import { Pool } from "pg";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Adjust based on your needs
idleTimeoutMillis: 30000,
});
```
### "relation 'typegraph_nodes' does not exist"
**Cause:** Migration not run.
**Solution:** Run the migration SQL:
```typescript
// PostgreSQL
import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
await pool.query(generatePostgresMigrationSQL());
// SQLite
import { generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
sqlite.exec(generateSqliteMigrationSQL());
```
### "permission denied" / cannot create relation on boot
**Cause:** `createStoreWithSchema()` is a privileged entry point. Its warm
base-schema check is one marker read with no base-adoption DDL, but bootstrap,
pending base adoption, graph migrations, contribution preparation, or system
index materialization can issue DDL. A DML-only role cannot safely own it.
**Solution:** Run schema/DDL changes as a privileged one-time migration
step, then attach at runtime with the zero-DDL
`createVerifiedStore()` (or `createStore()`) under the least-privilege
role. See
[Database roles & least privilege](/backend-setup#database-roles--least-privilege).
### `BaseSchemaMigrationError` from a zero-DDL runtime path
**Cause:** `createVerifiedStore`, `assertSchemaCurrent`, or graph-template
registration/instantiation found a missing, stale, or newer deployment-wide
base-schema marker. These paths deliberately do not repair physical storage.
**Solution:** For a missing or stale marker, run
`createStoreWithSchema(graph, adminBackend)` once under a DDL-capable role, or
apply the published external base-schema migration and stamp its marker last.
For a newer marker, deploy a TypeGraph release that supports that version.
The error details include `installedVersion`, `requiredVersion`, and `reason`.
### `MigrationError` from `createVerifiedStore` / `assertSchemaCurrent`
**Cause:** The runtime is using a code graph whose schema is ahead of
the database. The least-privilege runtime cannot migrate — by design,
it fails fast so requests don't run against a stale schema.
**Solution:** Run `createStoreWithSchema(graph, adminBackend)` under
the privileged role before promoting the new runtime build (apply any
generated migration SQL first if you manage DDL externally), then
restart the runtime. The thrown `MigrationError.message` includes the
diff summary and migration actions to apply.
### `ConfigurationError`: "no schema has been initialized"
**Cause:** A verifying attach (`createVerifiedStore` /
`assertSchemaCurrent`) ran before any privileged
`createStoreWithSchema()` boot — the database has no `schema_versions`
row (or no typegraph tables at all). The runtime deliberately refuses
to bootstrap under a least-privilege role. **Note:** running only the
generated migration SQL is not sufficient — it creates the tables but
does not write the schema row or contribution markers.
**Solution:** Run `createStoreWithSchema(graph, adminBackend)` once
under the privileged role. If you manage DDL externally with
drizzle-kit / `generatePostgresMigrationSQL()` /
`generateSqliteMigrationSQL()`, apply that first, then still run
`createStoreWithSchema()` to commit the schema row and contribution
markers. See
[Database roles & least privilege](/backend-setup#database-roles--least-privilege).
### `StoreNotInitializedError` on the first operation
**Cause:** The store was created with `createStore()` (a zero-I/O attach
that never materializes runtime storage) against a database that no
`createStoreWithSchema()` boot has initialized — commonly the runtime
started before the privileged migration step ran, or the wrong role/
database is configured. This covers fulltext operations and **embedding
writes**: a `store.nodes.*.create({ embedding })` (or embedding
update/delete) against an un-provisioned per-`(kind, field)` table throws
here rather than lazily issuing `CREATE TABLE` on the hot path. Vector
*reads* — `store.search.vector`, `store.search.hybrid`, and a
query-builder `.similarTo()` predicate — compile straight to SQL against
the per-field table, so they surface the engine's own missing-relation
error instead (`no such table: tg_vec_…` on SQLite, `relation … does not
exist` on Postgres) — same cause, same solution.
`createVerifiedStore()` catches every one of these cases at boot rather
than at the first hot-path operation.
A **`stale`** variant of this error on a vector field means something
different: the storage exists but was provisioned at a different shape —
typically the field's declared dimension changed after the table was
created. Boot deliberately leaves such a slot untouched (with a console
warning); run `store.reembedVectorField(kind, fieldPath)` to recreate the
storage at the new shape and re-embed.
**`ContributionUnavailableError`** with `state: "physical-storage-missing"`
means the physical fulltext table disappeared after initialization. Gated
fulltext operations preserve the driver error as `cause`, and transactional
backends roll back failed searchable writes. Query-builder fulltext predicates
compile directly to SQL and can still surface the engine's missing-relation
error. Run `store.rebuildContribution("fulltext")` to recreate the table and
repopulate it from the graph's nodes. A verified attach checks markers rather
than the physical catalog; use `probeContributions()` when startup must detect
out-of-band table loss.
**Solution:** Run `createStoreWithSchema(graph, adminBackend)` once
under the privileged role before the runtime attaches (it writes the
contribution markers that `createStore` / `createVerifiedStore` only
check), and prefer `createVerifiedStore()` over bare `createStore()` so
drift fails fast. See
[Database roles & least privilege](/backend-setup#database-roles--least-privilege).
A plain `createStore()` performs no reads and therefore cannot check the
base-schema marker at attach. If an edge write reaches legacy storage without
the match-identity columns, it throws `ConfigurationError` with
`details.code === "EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE"` instead of leaking
the database driver's missing-column error. The remedy is the same privileged
base-schema adoption.
### `SCHEMA_WRITE_FENCE_UNSUPPORTED` on the first managed write
**Cause:** The Store carries committed schema metadata (for example, it came
from `createStoreWithSchema`, `createVerifiedStore`, an adapter equivalent, or a
cached `{ reconciled }` snapshot), but its backend cannot run transactions or
does not implement the schema-write fence. Common examples are Cloudflare D1,
`drizzle-orm/neon-http`, and incomplete custom backends. The attach can still
succeed for reads; writes fail closed rather than racing a schema change.
**Solution:** Use a transactional backend (`neon-serverless`, regular
PostgreSQL, SQLite with transactions, or Durable Objects SQLite). If raw,
unfenced writes are an explicit application decision, construct the Store with
`createStore()` / `createAdapterStore()` without `{ reconciled }` and quiesce
writers yourself around schema changes.
### A convergence or claim write refuses on a non-transactional backend
This is expected when the operation needs an interactive transaction. Static
adapter batches are not a public transaction, and a sequence of independent
requests cannot safely implement Operational Identity, claim/cardinality
checks, or undeclared dynamic `matchOn` convergence. Use a backend with
`capabilities.execution.interactiveTransactions === true` and call `store.transaction(...)` for
those operations.
A declared edge `matchIdentity` is the exception for the eligible root
`getOrCreateByEndpoints` path: its persisted canonical key is backed by a
database unique arbiter, so the authoritative one-statement command can
return `created` or `found` without an interactive transaction. This exception
does not cover claims, history/revision sidecars, or other writes in the same
application workflow.
### The store opens clean but a fulltext or vector read fails
**Cause:** The durable physical contribution marker still says `initialized`
while the table it names is gone — a partial restore, a
hand-run `DROP`, or a schema-scoped restore that missed the
strategy-owned tables. Nothing on the open path probes the catalog:
boot and the runtime asserts short-circuit on a per-instance signature
cache and then on durable marker rows alone, which keeps the hot path free
of catalog round trips. The cost is that this database opens
completely clean and fails at the first read or write that depends on the affected slot.
For deployment-scoped storage such as the shared fulltext table, readiness is
the conjunction of that physical marker and a graph-local activation marker.
Vector tables remain graph-scoped and use one graph-local physical marker.
**Diagnosis:** `store.verifyContributions()` reports detected drift or a
recorded failed attempt among contributions currently expected by the active
graph and backend strategies. Each entry carries `owner`, `logicalName`,
`physicalName`, a `state`, and — for vector slots — `kind` and
`fieldPath`. When the marker recorded an error against its last
attempt, `lastError` carries it: `state` tells you which repair to run,
`lastError` tells you why it broke, which is often a different
question. The call is read-only — one existence query per contribution
table, no DDL and no writes — so it is safe on a live store under a
least-privilege role. It is deliberately **not** a boot step; call it
from a health check or an operator script.
**Solution:** Run `repairContributions()` on a Store backed by the privileged
DDL-capable connection:
```typescript
const result = await adminStore.repairContributions();
for (const entry of result.results) {
if (entry.status === "failed") {
console.error(entry.diagnostic, entry.error);
}
if (entry.status === "requires-rebuild") {
console.warn("manual rebuild required", entry.diagnostic);
}
}
```
The method performs its own fresh audit and resolves contribution declarations
from the committed graph and the active backend strategies. It never accepts a
diagnostic, physical table name, or DDL from the caller. A Store opened before
another writer evolved the graph catches up before enumerating vector slots
instead of repairing from its stale in-memory graph snapshot.
| `state` | Result | Data behavior |
| --- | --- | --- |
| `missing-marker` | `repaired` or `failed` | Runs idempotent DDL and re-stamps the marker; existing rows are preserved |
| `failed-materialization` | `repaired` or `failed` | Retries the current idempotent contribution DDL |
| `orphaned-marker` | `requires-rebuild` | The table and its data are already gone |
| `stale` | `requires-rebuild` | The stored physical shape does not match the current declaration |
`remaining` is a fresh post-repair diagnostic pass. An empty `remaining` array
means no current declaration remains unhealthy after the pass. Once it is
empty, a second call is idempotent and returns no results.
For a vector `requires-rebuild` entry, use
`reembedVectorField(kind, fieldPath, { embed })`. It drops and recreates the
slot, so pass an `embed` callback or the field comes back with zero embeddings.
For a fulltext `requires-rebuild` entry, use
`rebuildContribution("fulltext")` — the third rung of the ladder, described
below. Do not hand-edit the marker or run backend-owned DDL directly.
`repairContributions()` intentionally does not use the public diagnostic as an
instruction list and does not force marker writes. A warm backend re-reads the
marker, and the normal signature guard still refuses to bless stale storage.
**An empty result does not mean everything was checked.** The diagnostic
enumerates only current declarations. It ignores retired marker rows and treats
an expected contribution with neither marker nor table as never attempted, so
`[]` is not proof of initialization. A backend that cannot probe its own catalog
throws `ConfigurationError` rather than reporting a clean bill of health, but
vector slots on a backend without vector support are skipped silently and
correctly — that backend never materialized them, so reporting them would be a
false positive on every store it opens. For a readiness check, first attach with
`createVerifiedStore()` to establish schema and marker initialization, then run
this diagnostic. Also assert `backend.capabilities.vector?.supported` when
embedding storage is required rather than treating an empty array as proof that
it is intact.
### Contribution health: probe, repair, rebuild
The three contribution maintenance operations form one escalation ladder.
Each rung does strictly more, and costs strictly more, than the one below
it. Start at the top of this table and stop as soon as the projection is
`ready`.
| Rung | Call | Writes | Use when |
| --- | --- | --- | --- |
| 1. Probe | `store.probeContributions()` | Nothing | You want to know whether search is coherent right now. Safe on a read path, on a replica, and under a least-privilege role |
| 2. Repair | `store.repairContributions()` | Marker rows and idempotent `CREATE ... IF NOT EXISTS` | The probe reports `degraded` and `verifyContributions()` says `missing-marker` or `failed-materialization` — storage is intact and only the bookkeeping is wrong |
| 3. Rebuild | `store.rebuildContribution("fulltext")` | **Deletes and refills this graph's rows**; drops and recreates the shared storage only when no other graph has rows in it | `verifyContributions()` says `stale` or `orphaned-marker`, which repair reports as `requires-rebuild` |
**Rung 1 — the read-only probe.** One entry per search projection the
graph declares, so a caller can decide whether to issue a query without
running a write operation first:
```typescript
const health = await store.probeContributions();
for (const entry of health.entries) {
if (entry.state !== "ready") {
console.warn(`${entry.contribution} search is ${entry.state}`, entry.detail);
}
}
```
`entries` is empty when there is nothing to assess — a graph with no
`searchable()` or `embedding()` fields, or a backend with no contribution
machinery. It is never empty because a check was skipped: a backend that
provisions contributions but cannot probe its catalog throws
`ConfigurationError`, and declares the gap as
`capabilities.contributions.probe === false`. Route on `state`; `detail`
is a human-readable summary and not a stable format, so call
`verifyContributions()` for the structured per-table findings behind it.
`graphRevision` stamps the durable revision the assessment was taken at,
placing the probe in the graph's committed history. It is graph-global
like the clock it reads: an advance between two probes means something
committed in between, not that a particular caller's write landed. It is
absent unless the Store is revision-tracked (`revisionTracking: true` or
`history: true`) and a tracked write has already anchored the clock. A
store with no revision clock has no revision to stamp, and substituting a
wall-clock timestamp or the schema version would be a weaker guarantee
wearing the name of a stronger one — the schema version in particular does
not advance on data writes, so it could not order anything.
`state: "building"` is reserved. No shipped path publishes it; the
destructive rebuild is atomic, so a concurrent probe observes the state
before or after it and never a partial one. Treat it as "not `ready`".
**Rung 3 — the destructive rebuild.** A `stale` fulltext contribution
means the table exists at the shape a *previous* `createDdl` produced. The
ordinary ensure path cannot fix it: its `CREATE ... IF NOT EXISTS` no-ops
against the existing table, and re-stamping the marker there would leave
it blessing storage whose shape is wrong — precisely what the drift guard
exists to prevent. Only a drop makes the recreate meaningful, so the drop
is its own named operation rather than a flag on the ensure path:
```typescript
const result = await adminStore.rebuildContribution("fulltext");
// { rebuilt: ["typegraph_node_fulltext"], processed, repopulated, skipped }
```
**Opening a Store to run it.** A `stale` contribution makes
`createStoreWithSchema()` refuse: its boot step materializes runtime
contributions, and the drift guard will not run the current DDL against a
table provisioned at another shape. That refusal is deliberate and
persistent — it repeats on every restart until the shape is fixed, and it
leaves the `stale` verdict intact rather than downgrading it to a state
whose repair would bless the wrong shape. Reach the rebuild from a Store
opened without that boot step, which `createStore()` and
`createVerifiedStore()` are (they run no DDL by contract):
```typescript
const adminStore = createStore(graph, backend);
await adminStore.probeContributions(); // degraded, detail names `stale`
await adminStore.rebuildContribution("fulltext");
// createStoreWithSchema() now opens normally again.
```
**The rebuild is scoped to the graph you call it on.** The fulltext
projection is one physical table holding every graph's rows keyed by
`graph_id`, so the default teardown is the same
`DELETE ... WHERE graph_id` that `clear()` issues — this graph's index
content and nothing else — followed by the current `createDdl`, a refill
from this graph's node rows, and the marker stamp, all in one transaction
under the same per-graph fence as a schema commit. That path takes **no
table lock at all**: the delete is transactional and touches only rows this
graph owns, so it never makes another graph's writers wait. An interrupted
rebuild rolls back to the state it started from rather than leaving storage
attested but empty.
It escalates to dropping and recreating that shared table only when the
table holds no other graph's rows — the case where the drop takes nothing
with it. That escalation is the one repair for storage provisioned at a
shape the current DDL no longer produces, and because the DDL it issues is
database-global it runs under a database-scoped advisory lock
(`typegraph:contribution-ddl`; a no-op on SQLite, whose fence already holds
the single writer slot) rather than only the per-graph fence.
**When the shared table is in use by another graph, a `stale` rebuild
refuses.** Only recreating the storage repairs a `stale` shape, so if that
storage still holds rows belonging to other graphs the call throws
`ContributionRebuildUnsupportedError` with
`reason: "shared-storage-in-use"` rather than destroying content it cannot
reconstruct — those rows are derived from other graphs' nodes through their
own schemas — or re-stamping this graph's marker over a physical shape
nothing verified. `details.otherGraphIds` names the graphs that are in the
way. The sanctioned repair is a maintenance window with every graph on that
database offline: drop the table out of band, then run
`store.rebuildContribution("fulltext")` once per graph, each run recreating
the table from the current DDL and refilling that graph's own rows.
What the *recreate* path costs, and why it is still the right trade: the
transaction is held for the whole refill, and on PostgreSQL the rebuild
takes `LOCK TABLE ... IN ACCESS EXCLUSIVE MODE` on the shared table and
keeps it until commit. It takes that lock **before** deciding to drop, not
merely as a side effect of the `DROP TABLE` — the verdict "no other graph
has rows here, so dropping this destroys nothing" is only as good as the
exclusion it was computed under. Ordinary fulltext writes take no advisory
lock, so the contribution DDL lock excludes other *rebuilds* and nothing
else: a neighbouring graph's `INSERT` could commit between an unlocked
probe and the drop, and be destroyed by a rebuild that had already decided
it was alone. The sequence is therefore probe → `ACCESS EXCLUSIVE` →
re-probe, and only the re-probe's verdict authorizes a drop. A verdict that
flips under the lock loses the drop and keeps the lock (PostgreSQL holds
locks until commit), which is the rare and safe direction to be wrong in.
The cheap unlocked probe ahead of it exists only to keep the graph-scoped
path off the relation lock, and can only err toward keeping the table.
The window blocks more than searches — every write to a kind with
`searchable()` fields maintains the same table, so those block too, for
every graph on the database. On SQLite the rebuild holds the write lock for
the same span, so concurrent writers wait out their busy timeout and then
fail. Run it in a maintenance window on a large graph.
When the storage *shape* is fine and only the content is stale — a field
gained `searchable()` after data was written, or a `language` changed —
`store.search.rebuildFulltext()` is the incremental, resumable pass that
transacts per page instead.
Nothing is permanently lost for the graph you rebuild: its searchable text
is derived from node properties TypeGraph already stores. Other graphs on
the same database are not in reach either — their rows are kept by the
graph-scoped delete, and the drop that would take them never runs (a
`stale` shape that could only be repaired by that drop refuses instead).
Nodes whose stored `props` cannot be read as an object are counted in
`skipped` and are absent from the rebuilt index;
`store.search.rebuildFulltext()` reports their ids individually.
**Vector contributions cannot be rebuilt, and the call refuses rather
than trying.** `rebuildContribution("vector")` always throws
`ContributionRebuildUnsupportedError` with
`reason: "vector-source-unavailable"`. TypeGraph stores the vectors
callers supply and never the inputs that produced them, so the embeddings
exist only in the storage a rebuild would drop — dropping anyway would
destroy them and hand back storage that looks healthy and returns
nothing. `reembedVectorField(kind, fieldPath, { embed })` is the
sanctioned destructive path for vector storage precisely because it takes
the callback that can regenerate what the drop discards.
The same typed error covers two wiring gaps, and both refuse before
anything is dropped: `reason: "no-drop-ddl"` when the active fulltext
strategy declares no `dropDdl` on its contribution, and
`reason: "no-schema-fence"` when the backend exposes no
`schemaWriteTransaction` to make the sequence atomic (the HTTP-only
PostgreSQL drivers, and SQLite with transactions disabled). Both are
declared ahead of time as `capabilities.contributions.rebuild === false`.
## Semantic Search Issues
### "Extension not found" / "vector type not available"
**Cause:** Vector extension not installed. Only applies to PostgreSQL
(pgvector) and SQLite (sqlite-vec). libSQL / Turso has a built-in
native vector engine — there is nothing to load and it is wired
automatically by `createLibsqlBackend`.
**PostgreSQL:**
```sql
CREATE EXTENSION IF NOT EXISTS vector;
```
**SQLite:**
```typescript
import * as sqliteVec from "sqlite-vec";
sqliteVec.load(sqlite); // Must be called before creating backend
```
### "Dimension mismatch"
**Cause:** Query embedding has different dimension than stored embeddings.
**Solution:** Use consistent embedding dimensions:
```typescript
// Schema defines 1536 dimensions
const Document = defineNode("Document", {
schema: z.object({
embedding: embedding(1536),
}),
});
// Query embedding must also be 1536
const queryEmbedding = await generateEmbedding(text);
console.log(queryEmbedding.length); // Should be 1536
```
### "Inner product not supported" (SQLite / libSQL)
**Cause:** `inner_product` is PostgreSQL-only. Neither sqlite-vec nor
libSQL support the inner product metric (cosine and l2 only). Check
`backend.capabilities.vector.metrics` for the active backend.
**Solution:** Use cosine or L2:
```typescript
// Instead of:
d.embedding.similarTo(query, 10, { metric: "inner_product" });
// Use:
d.embedding.similarTo(query, 10, { metric: "cosine" });
```
## TypeScript Issues
### "Property 'x' does not exist on type"
**Cause:** Accessing a property not defined in your schema.
**Solution:** Ensure the property is in your Zod schema:
```typescript
const Person = defineNode("Person", {
schema: z.object({
name: z.string(),
email: z.string().optional(),
}),
});
// Now both properties are available with correct types
const person = await store.nodes.Person.getById(id);
person?.name; // string
person?.email; // string | undefined
```
### Type inference not working in select
**Cause:** Complex generic inference limitations.
**Solution:** Use explicit typing or simplify:
```typescript
// If inference fails, be explicit
.select((ctx) => ({
name: ctx.p.name as string,
company: ctx.c.name as string,
}))
```
## Still Having Issues?
1. **Check the [Limitations](/limitations)** page for known constraints
2. **Review [Architecture](/architecture)** to understand how TypeGraph works
3. **Search [GitHub Issues](https://github.com/nicia-ai/typegraph/issues)** for similar problems
4. **Open a new issue** with a minimal reproduction case
# Errors
> Error types and handling in TypeGraph
TypeGraph uses typed errors to communicate specific failure conditions. All errors extend the base
`TypeGraphError` class and include categorization, contextual details, and actionable suggestions.
## Error Categories
Every error is categorized to help determine the appropriate response:
| Category | Description | Typical Response |
|----------|-------------|------------------|
| `user` | Invalid input or misuse of API | Fix the input and retry |
| `constraint` | Graph constraint violated | Handle as business logic violation |
| `system` | Internal or infrastructure error | Log, alert, potentially retry |
```typescript
import { isUserRecoverable, isConstraintError, isSystemError } from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create(data);
} catch (error) {
if (isUserRecoverable(error)) {
// Show validation errors to user
return { error: error.toUserMessage() };
}
if (isConstraintError(error)) {
// Handle business rule violation
return { error: "This operation violates a constraint" };
}
if (isSystemError(error)) {
// Log and alert
console.error(error.toLogString());
throw error;
}
}
```
## Base Error
### `TypeGraphError`
Base error class for all TypeGraph errors.
```typescript
class TypeGraphError extends Error {
readonly code: string;
readonly category: ErrorCategory;
readonly details: Readonly>;
readonly suggestion?: string;
// Format error for end users (includes suggestion if available)
toUserMessage(): string;
// Format error for logging (includes code, category, and details)
toLogString(): string;
}
type ErrorCategory = "user" | "constraint" | "system";
```
**Properties:**
| Property | Type | Description |
|----------|------|-------------|
| `code` | `string` | Machine-readable error code |
| `category` | `ErrorCategory` | Error classification for handling |
| `details` | `Record` | Additional context about the error |
| `suggestion` | `string \| undefined` | Actionable guidance for resolution |
**Methods:**
| Method | Returns | Description |
|--------|---------|-------------|
| `toUserMessage()` | `string` | Human-readable message with suggestion |
| `toLogString()` | `string` | Detailed string for logging/debugging |
## Validation Errors
### `ValidationError`
Thrown when schema validation fails during node or edge creation/update. Includes structured issue
details with context about which entity failed.
```typescript
interface ValidationErrorDetails {
readonly issues: readonly ValidationIssue[];
readonly entityType?: "node" | "edge";
readonly kind?: string;
readonly operation?: "create" | "update";
readonly id?: string;
}
interface ValidationIssue {
readonly path: string;
readonly message: string;
readonly code?: string;
}
```
**Example:**
```typescript
try {
await store.nodes.Person.create({ name: "" }); // Empty name fails min(1)
} catch (error) {
if (error instanceof ValidationError) {
console.log(error.category); // "user"
console.log(error.details.kind); // "Person"
console.log(error.details.operation); // "create"
console.log(error.details.issues);
// [{ path: "name", message: "String must contain at least 1 character(s)" }]
console.log(error.toUserMessage());
// "Validation failed for Person create: name - String must contain at least 1 character(s)
//
// Suggestion: Check the data you're providing matches the schema..."
}
}
```
#### `INVERTED_VALIDITY_WINDOW`
A `ValidationError` whose issue carries the exported code
`INVERTED_VALIDITY_WINDOW` refused a valid-time window of negative width: the
write's `validTo` precedes the row's effective `validFrom`, so the row would have
stopped being true before it started and no `asOf` coordinate could observe it.
Branch on the code rather than on the message.
```typescript
import { INVERTED_VALIDITY_WINDOW_CODE, ValidationError } from "@nicia-ai/typegraph";
try {
// The stored validFrom is later than this end.
await store.edges.worksAt.update(edgeId, {}, { validTo: "2020-01-01T00:00:00.000Z" });
} catch (error) {
if (
error instanceof ValidationError &&
error.details.issues.some((issue) => issue.code === INVERTED_VALIDITY_WINDOW_CODE)
) {
// Supply an explicit validFrom for a historical window, or drop validTo.
}
}
```
Interchange import records the same refusal as a per-row error prefixed with the
code, so one bad row does not abort the import; trusted import refuses the whole
stream with `TrustedImportError` reason `invalid_stream`. A zero-width window
(`validTo === validFrom`) is legal and never raises this, and neither is a write
that STAMPS its own start while carrying only a historical `validTo`: any create,
and a node resurrection through `upsertById` / `bulkUpsertById`. Both store no
lower bound instead. An edge resurrection RETAINS the bound the row already
holds, so a `validTo` before that bound still raises this.
#### `IMMUTABLE_VALIDITY_LOWER_BOUND`
A `ValidationError` whose issue carries the exported code
`IMMUTABLE_VALIDITY_LOWER_BOUND` refused a `validFrom` the write could not apply.
A live row's lower bound is history: an in-place update never rewrites
`valid_from`, so a bound naming a different instant is refused rather than
accepted and silently dropped. The message names both instants — the one stated
and the one the row stores — so you can restate the stored bound without a
second read.
```typescript
import { IMMUTABLE_VALIDITY_LOWER_BOUND_CODE, ValidationError } from "@nicia-ai/typegraph";
try {
// The row is live and started at some other instant.
await store.nodes.Person.upsertById(id, props, { validFrom: "2020-01-01T00:00:00.000Z" });
} catch (error) {
if (
error instanceof ValidationError &&
error.details.issues.some(
(issue) => issue.code === IMMUTABLE_VALIDITY_LOWER_BOUND_CODE,
)
) {
// Omit validFrom, or restate the bound the row already holds.
}
}
```
What deliberately does not raise it:
- **Restating the stored bound.** Naming the instant the row already holds is
accepted; there is nothing to apply and nothing being ignored.
- **A create, or a resurrection.** Both write a fresh window, so a stated
`validFrom` is stored — that is the way to give a row a different lower bound.
- **`getOrCreateByEndpoints` returning an existing edge.** That branch performs
no write, so `validFrom` / `validTo` describe the row to create if none is
found. `clearValidTo` is refused on a live return-mode match because it names
a mutation; use `ifExists: "update"`.
- **A node upsert or endpoint-matched edge update with
`onImmutableLowerBound: "preserve"`.** This explicitly treats `validFrom` as
create/resurrection-only input. A live-row update keeps its stored lower
bound while still applying props and `validTo`; the default remains
`"refuse"` so an unqualified bound is never silently dropped. Edge updates
use the policy with `ifExists: "update"`; the bulk edge form sets it per item.
Under the default `"refuse"` policy, it reaches every path that accepts
`validFrom` against a live row: `upsertById`, `bulkUpsertById` (including a
repeated id in one batch, judged against the row the batch just queued),
`getOrCreateByEndpoints` / `bulkGetOrCreateByEndpoints` with
`ifExists: "update"`, and interchange import's `onConflict: "update"` legs —
where, as with the inverted-window refusal, it is recorded as a per-row error
prefixed with the code rather than aborting the import.
#### `ENTITY_ALREADY_EXISTS`
A `ValidationError` whose issue carries the exported code
`ENTITY_ALREADY_EXISTS` refused a create because the id is already taken.
`details.entityType` says whether a node or an edge was refused and
`details.kind` names its kind.
```typescript
import { ENTITY_ALREADY_EXISTS_CODE, ValidationError } from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create({ name: "Alice" }, { id: takenId });
} catch (error) {
if (
error instanceof ValidationError &&
error.details.issues.some((issue) => issue.code === ENTITY_ALREADY_EXISTS_CODE)
) {
// Use a different id, or update the existing entity.
}
}
```
The code is the same whichever layer noticed, on either backend. A node create
finds out from its own existence probe — but the probe and the INSERT are two
statements, and PostgreSQL does not serialize two write transactions under its
default READ COMMITTED isolation, so a concurrent create of the same NEW id can
commit in between and the engine refuses the INSERT instead. (SQLite's
`BEGIN IMMEDIATE` gives the writer slot to one transaction at a time, so its probe
always sees the winner's row.) An edge create has no existence probe at all, so
the engine's refusal is always what reports a taken edge id. All of these raise the
same error, so a caller retrying a generated id needs one branch, not several.
`details.id` names the taken id, and is present for every single-entity create.
It is absent only when the refused statement inserted more than one row: the
engine reports that the statement collided without saying which row did, and its
transaction is already aborted, so there is nothing left to probe. No race is
needed to reach that — a bulk create of edges, whose ids you supplied and which
nothing probes, is refused this way on every backend. Treat `details.id` as
optional if you create in bulk.
This is about identity, not values. A conflict on a declared `unique` constraint
raises `UniquenessError` instead, and a violated `unique: true` index declaration
surfaces as the engine's own failure — neither is reshaped into this error.
### `DisjointError`
Thrown when attempting to create a node that violates a disjointness constraint.
```typescript
// If Person and Organization are disjoint:
await store.nodes.Person.create({ name: "Alice" }, { id: "entity-1" });
try {
// Same ID, different disjoint type
await store.nodes.Organization.create({ name: "Acme" }, { id: "entity-1" });
} catch (error) {
if (error instanceof DisjointError) {
console.log(error.category); // "constraint"
console.log(error.details);
// { nodeId: "entity-1", attemptedKind: "Organization", conflictingKind: "Person" }
console.log(error.suggestion);
// "Use a different ID for the new node, or delete the existing node first..."
}
}
```
### `IdentityContradictionError`
Thrown when an identity mutation would make the assertion ledger contradictory —
for example asserting two nodes are the same after they were asserted different,
folding a same-class pair the ontology forbids, or importing an archive whose
assertions conflict with the target graph. Only raised on identity-enabled
graphs.
```typescript
try {
await tx.identity.assertSame(alice, aliceCopy);
} catch (error) {
if (error instanceof IdentityContradictionError) {
console.log(error.code); // "IDENTITY_CONTRADICTION"
console.log(error.category); // "constraint"
console.log(error.details);
// {
// operation: "assertSame", // "assertSame" | "assertDifferent" | "fold" | "import"
// a: { kind: "Person", id: "..." },
// b: { kind: "Person", id: "..." },
// reason: "different-assertion", // "different-assertion" | "same-class" | "disjoint-kinds"
// conflictingAssertionId: "...", // present when an existing assertion conflicts
// conflictingKinds: ["Person", "Organization"], // present when reason is "disjoint-kinds"
// }
console.log(error.suggestion);
// "Retract the conflicting identity assertion or correct the graph ontology before retrying."
}
}
```
### Identity validity errors
`IdentityValidityWindowError` refuses a future start, future end, inverted
window, or a second non-identical open window for one current semantic pair.
Its code identifies the reason: `IDENTITY_VALIDITY_FUTURE_START`,
`IDENTITY_VALIDITY_FUTURE_END`, `IDENTITY_VALIDITY_INVERTED`, or
`IDENTITY_VALIDITY_OPEN_WINDOW_CONFLICT`.
`IdentityEndpointValidityError` (`IDENTITY_ENDPOINT_VALIDITY`) means an
explicit assertion window extends outside an endpoint node's own validity or
deletion bounds. Future or inverted identity windows are user-category input
errors. A second non-identical open window and an endpoint-window conflict are
constraint-category errors. Both classes are package-root exports.
### `IdentityMergeConflictError`
Detected at merge **plan time** when the branches being merged carry opposing
or otherwise contradictory identity truth: one branch asserts a pair `same`
while another asserts it `different` (directly, or transitively through a
chain of `same` assertions no single branch ever wrote), a branch retracts an
assertion that a different branch reasserts under a new id (a retract/reassert
race — a branch that reasserts a pair it *also* retracted itself is
convergent, not a conflict, and merges cleanly), or a branch asserts an
identity relation over a node another branch deleted. Extends `MergeError`, so
an `instanceof MergeError` catch covers it alongside the other merge failures.
`merge()` and `IdentityMergeConflictError` are both exported from
`@nicia-ai/typegraph/graph-merge`, not the package root. `merge()` takes an
array of branches and never throws a `MergeError` — it **returns** a
`Result`:
```typescript
import { merge, IdentityMergeConflictError, isErr } from "@nicia-ai/typegraph/graph-merge";
const result = await merge(store, [branch]);
if (isErr(result)) {
if (result.error instanceof IdentityMergeConflictError) {
console.log(result.error.code); // "GRAPH_MERGE_IDENTITY_CONFLICT"
console.log(result.error.details);
}
throw result.error;
}
```
### `MergeConstraintConflictError`
Returned when `merge()`, `mergeIncremental()`, or `applyMergePlan()` resolves a
plan whose final graph violates a deterministic store constraint. The store
remains the owner of constraint enforcement: the merge translates its typed
refusal only at the commit boundary, after the transaction has rolled back.
```typescript
import {
isErr,
merge,
MergeConstraintConflictError,
} from "@nicia-ai/typegraph/graph-merge";
const result = await merge(store, branches);
if (isErr(result) && result.error instanceof MergeConstraintConflictError) {
console.log(result.error.code); // "GRAPH_MERGE_CONSTRAINT_CONFLICT"
console.log(result.error.category); // "constraint"
console.log(result.error.details.constraintCode); // e.g. "CARDINALITY_ERROR"
console.log(result.error.details.edgeKind); // copied from the store error
console.log(result.error.cause); // the original CardinalityError, etc.
}
```
Cardinality, uniqueness, endpoint, disjointness, and restricted-delete
refusals share this surface when they arise from node or edge application.
The planner normally co-buckets nodes with the same declared unique key, but a
late store-owned uniqueness refusal uses the same completeness boundary rather
than falling back to a system error.
Identity truth conflicts retain `IdentityMergeConflictError`; backend,
environment, and stale-plan failures retain their existing system errors.
Constraint failure is atomic: neither graph writes nor merge provenance records
survive.
### Merge plan and evidence errors
The reviewable merge lifecycle also returns errors in its `Result` arm. It does
not throw them:
```typescript
import {
applyMergePlan,
isErr,
planMerge,
StaleMergePlanError,
} from "@nicia-ai/typegraph/graph-merge";
const planned = await planMerge(store, branches, options);
if (isErr(planned)) throw planned.error;
const applied = await applyMergePlan(store, planned.data);
if (isErr(applied)) {
if (applied.error instanceof StaleMergePlanError) {
// The reviewed artifact no longer describes the target. Plan and review again.
}
throw applied.error;
}
```
| Error | Code | Meaning |
| --- | --- | --- |
| `MergePlanCapabilityError` | `GRAPH_MERGE_PLAN_CAPABILITY` | The target cannot supply a durable revision fence for a cross-time plan. Enable `revisionTracking` or `history`; the contiguous `merge()` wrappers retain their documented compatibility behavior. |
| `MergePlanningStaleError` | `GRAPH_MERGE_PLANNING_STALE` | The target revision changed between the planner's opening and closing observations. This is an expected retry-and-replan outcome under concurrency: no artifact is returned, so recapture the target and create a new plan before retrying. |
| `StaleMergePlanError` | `GRAPH_MERGE_PLAN_STALE` | The target moved after planning, the plan already succeeded, or another concurrent application won. No plan writes committed. |
| `InvalidMergePlanError` | `GRAPH_MERGE_PLAN_INVALID` | The value failed the versioned plan schema or a semantic invariant. |
| `UnsupportedMergePlanVersionError` | `GRAPH_MERGE_PLAN_VERSION_UNSUPPORTED` | `formatVersion` is not supported by this TypeGraph version. |
| `MergePlanDigestMismatchError` | `GRAPH_MERGE_PLAN_DIGEST_MISMATCH` | Canonical plan content differs from the recorded digest. |
| `MergePlanTargetMismatchError` | `GRAPH_MERGE_PLAN_TARGET_MISMATCH` | The plan names a different graph id from the supplied target. |
| `MergePlanSchemaMismatchError` | `GRAPH_MERGE_PLAN_SCHEMA_MISMATCH` | The plan was resolved under a different active schema version or hash. |
| `MergePlanOriginMismatchError` | `GRAPH_MERGE_PLAN_ORIGIN_MISMATCH` | The target has an independently-created revision clock, even if its numeric revision happens to match. |
| `CandidateSourceError` | `GRAPH_MERGE_CANDIDATE_SOURCE` | A built-in candidate source failed. `details` identifies its source id, entity kind, and operation context. |
| `MatchEvidenceError` | `GRAPH_MERGE_EVIDENCE` | Candidate evidence is malformed or a score is non-finite. `NaN` and infinity are refused, never serialized or silently dropped. |
Plan validation and the target/schema/origin/revision fence run before canonical
writes. The revision check is inside the same transaction as apply, so two
concurrent attempts cannot both commit. A stale plan is not repaired or adapted:
create a new plan and obtain approval for its new `digest`.
Plans may contain the complete proposed application data. Their digest detects
content changes and gives approval systems a stable identity, but it is not a
signature and does not authenticate storage, authorize a caller, or prove who
created the artifact. Protect plan data and enforce those trust decisions in the
application before calling `applyMergePlan()`.
### `EndpointError`
Thrown when an edge is created with invalid endpoint types.
```typescript
// If worksAt only allows Person -> Company:
try {
await store.edges.worksAt.create(company, person, {}); // Wrong direction
} catch (error) {
if (error instanceof EndpointError) {
console.log(error.category); // "constraint"
console.log(error.suggestion);
// "Check the edge definition to see which node types are allowed..."
}
}
```
### `EndpointPairError`
Thrown when a [source-dependent edge](/core-concepts#source-dependent-targets)
receives a source/target combination that matches no declared pair. It extends
`TypeGraphError` directly, so catching `EndpointError` alone does not catch it.
An invalid source kind continues to produce `EndpointError`.
```typescript
import { EndpointPairError } from "@nicia-ai/typegraph";
try {
// Dynamic callers are checked at runtime, too.
// assignedTo allows Employee -> Department and Student -> Course.
await store.getEdgeCollection("assignedTo").create(employee, course, {});
} catch (error) {
if (error instanceof EndpointPairError) {
console.log(error.code); // "ENDPOINT_PAIR_ERROR"
console.log(error.category); // "constraint"
console.log(error.details);
// {
// edgeKind: "assignedTo", endpoint: "pair",
// fromKind: "Employee", toKind: "Course",
// allowedPairs: [
// { from: "Employee", to: "Department" },
// { from: "Student", to: "Course" },
// ],
// }
}
}
```
Malformed target maps and graph registrations that widen built-in constraints
fail at configuration time with `ConfigurationError`.
### `CardinalityError`
Thrown when a cardinality constraint is violated.
```typescript
// If worksAt has cardinality: "one" (person can only work at one company):
await store.edges.worksAt.create(alice, acme, { role: "Engineer" });
try {
await store.edges.worksAt.create(alice, otherCompany, { role: "Consultant" });
} catch (error) {
if (error instanceof CardinalityError) {
console.log(error.category); // "constraint"
console.log(error.details);
// { edgeKind: "worksAt", fromKind: "Person", fromId: "", cardinality: "one", existingCount: 1 }
console.log(error.suggestion);
// "Remove the existing edge before creating a new one, or update the existing edge..."
}
}
```
### `UniquenessError`
Thrown when a uniqueness constraint is violated.
```typescript
// If email has a unique constraint:
await store.nodes.Person.create({ name: "Alice", email: "alice@example.com" });
try {
await store.nodes.Person.create({ name: "Bob", email: "alice@example.com" });
} catch (error) {
if (error instanceof UniquenessError) {
console.log(error.category); // "constraint"
console.log(error.details);
// { constraintName: "unique_email", kind: "Person", existingId: "", newId: "", fields: ["email"] }
console.log(error.suggestion);
// "Use a different value for the unique field, or update the existing record..."
}
}
```
### `EdgeMatchIdentityConflictError`
Thrown when a direct edge create collides with the edge kind's declared
`matchIdentity`. Use `getOrCreateByEndpoints()` when the intended behavior is
to return the existing identity owner.
## Not Found Errors
### `NodeNotFoundError`
Thrown when a referenced node does not exist.
```typescript
try {
await store.nodes.Person.update("nonexistent-id", { name: "New Name" });
} catch (error) {
if (error instanceof NodeNotFoundError) {
console.log(error.category); // "user"
console.log(error.details); // { kind: "Person", id: "nonexistent-id" }
console.log(error.suggestion);
// "Verify the node ID is correct and the node hasn't been deleted..."
}
}
```
### `EdgeNotFoundError`
Thrown when a referenced edge does not exist.
```typescript
try {
await store.edges.worksAt.update("nonexistent-edge", { role: "Manager" });
} catch (error) {
if (error instanceof EdgeNotFoundError) {
console.log(error.category); // "user"
console.log(error.details); // { kind: "worksAt", id: "nonexistent-edge" }
console.log(error.suggestion);
// "Verify the edge ID is correct and the edge hasn't been deleted..."
}
}
```
### `KindNotFoundError`
Thrown when referencing a node or edge type that doesn't exist in the graph definition.
```typescript
try {
await store.query().from("NonExistentType", "n").execute();
} catch (error) {
if (error instanceof KindNotFoundError) {
console.log(error.category); // "user"
console.log(error.details); // { kindName: "NonExistentType", entity: "node" }
console.log(error.suggestion);
// "Check the graph definition to see which node and edge types are available..."
}
}
```
### `EndpointNotFoundError`
Thrown when an edge references a node that doesn't exist.
```typescript
try {
await store.edges.worksAt.create(
{ kind: "Person", id: "nonexistent" },
company,
{ role: "Engineer" }
);
} catch (error) {
if (error instanceof EndpointNotFoundError) {
console.log(error.category); // "user"
console.log(error.details);
// { edgeKind: "worksAt", endpoint: "from", nodeKind: "Person", nodeId: "nonexistent" }
console.log(error.suggestion);
// "Create the referenced node first, or verify the node ID is correct..."
}
}
```
## Delete Errors
### `RestrictedDeleteError`
Thrown when delete is blocked due to existing edges (when `onDelete: "restrict"`).
```typescript
// If Person has edges and onDelete is "restrict":
try {
await store.nodes.Person.delete(alice.id);
} catch (error) {
if (error instanceof RestrictedDeleteError) {
console.log(error.category); // "constraint"
console.log(error.details);
// { nodeKind: "Person", nodeId: "", edgeCount: 3, edgeKinds: ["worksAt", "authored"] }
console.log(error.suggestion);
// "Delete all edges connected to this node first, or change the delete behavior..."
}
}
```
## Configuration Errors
### `ConfigurationError`
Thrown when the store, backend, or schema definition is misconfigured.
```typescript
// Using transactions on D1 (which doesn't support them):
try {
await store.transaction(async (tx) => {
// ...
});
} catch (error) {
if (error instanceof ConfigurationError) {
console.log(error.category); // "system"
console.log(error.suggestion);
// "Check the backend documentation for supported features..."
}
}
```
#### Definition-time unique-constraint refusals
`defineGraph()` validates every node kind's `unique` constraints when the graph
is defined, rather than leaving a broken `where` clause to surface as odd
behavior on the first write. Three states are refused with `ConfigurationError`:
- A `where` callback that **does not return a predicate** — `details` carries
`kind` and `constraintName`.
- A predicate naming a **field the kind's schema does not declare** — `details`
adds `field` and `declaredFields`.
- A `where` clause on a kind whose **schema is not an object schema** (it
exposes no `.shape`, so there is no declared-field set to check the clause
against) — `details` carries `kind` and `constraintName`. Refused rather than
left unvalidated, because skipping the check silently would disable this guard
for exactly the untyped callers it exists for. A plain `unique: [{ fields }]`
on such a schema is *not* refused: it names props by key and evaluates fine
against a non-object schema.
All three carry only the class-level code `CONFIGURATION_ERROR`; match them by
class, not by a `details.code`. The equivalent invariant on the graph-extension
document path does have a stable code, `UNKNOWN_UNIQUE_WHERE_FIELD`.
A constraint built **outside** `defineGraph` never passed this gate, so the
non-predicate case is refused at evaluation too: `checkWherePredicate` throws the
same `ConfigurationError` (with `constraintName` and `fields`) on the write path
instead of treating a broken clause as one that applies to every row. All three
readers of a `where` clause — definition-time validation, per-write evaluation,
and persistence-time capture — now agree, because they read it through one
shared function.
Because the check evaluates the clause, a `where` callback now runs once at
definition time in addition to its per-write evaluations — keep it pure. The
check applies to node kinds whose schema exposes an object shape; edge `unique`
constraints are not validated here. Statically typed callers were already unable
to name an undeclared field, so this bites untyped or generated definitions.
#### Definition-time `__proto__` property refusal
`defineNode()` / `defineEdge()` refuse a schema that declares a property named
`__proto__` with a `ConfigurationError` carrying `details.conflicts` and a
`nodeType` / `edgeType` key. The name is **unstorable**, not merely reserved:
Zod accepts it in a shape but drops it from every parse result — reporting
success even when the field is required — so a value written to it is silently
lost.
It is only reachable through a computed key. `z.object({ __proto__: … })`
written literally sets the shape object's own prototype instead of creating an
entry, while `z.object({ ["__proto__"]: z.string() })` yields a shape whose
`Object.keys` really does contain it.
The graph-extension document path refuses the identical declaration with the
stable issue code `RESERVED_PROPERTY_NAME`, at any nesting depth — so a nested
object field named `__proto__` is refused on the same grounds as a top-level
one. Before this, the two authoring paths disagreed about the same field: a
typed refusal on the document path, silent data loss on the typed one.
#### `CONSTRAINT_WRITE_FENCE_UNSUPPORTED`
A write guarded by a declared constraint runs its probe and its write under one
per-graph mutual exclusion. That fence is transaction-scoped on both dialects
(SQLite's `BEGIN IMMEDIATE`, PostgreSQL's `pg_advisory_xact_lock`), so a backend
reporting `capabilities.execution.interactiveTransactions: false` — Cloudflare D1, `drizzle-orm/neon-http`,
any SQLite backend built with `transactionMode: "none"` — cannot hold it, and the
write is refused rather than run unfenced. Durable Objects are unaffected.
`details.constraint` names which class needed the fence, because "this backend
cannot fence constrained writes" is unusable advice while "your
`cardinality: 'one'` edge cannot be enforced here" is actionable. The
`suggestion` carries the per-class way forward.
| `details.constraint` | The write it describes |
| --- | --- |
| `edgeCardinality` | Creating or resurrecting an edge whose `cardinality` is `one`, `unique`, or `oneActive`. |
| `edgeMatchKeyConvergence` | Endpoint convergence that requires the portable transaction-scoped path: an undeclared dynamic `matchOn`, constrained cardinality, update or temporal options, derived/custom backends, or schema-aware resurrection of a tombstoned winner. A schema-declared durable `matchIdentity` removes this fence from eligible live single-item and bulk create/found paths. |
| `nodeDisjointness` | Creating a node under a kind that participates in a `disjointWith` axiom. Probed only where a node comes into existence, so deletes and in-place updates are not refused. |
| `nodeUniquenessScope` | Creating **or updating** a node under a `scope: "kindWithSubClasses"` unique that actually expands past the node's own kind. A `scope: "kind"` unique is backed by the uniques primary key and needs no fence. |
`details.graphId` names the graph. Unconstrained writes on the same backend are
untouched — see
[Declared constraints require an interactive transaction](/backend-setup#declared-constraints-require-an-interactive-transaction)
for what still works there.
`CONSTRAINT_TRANSACTION_NOT_WRITE_FENCED` is the corresponding refusal for a
caller-adopted SQLite transaction whose `DEFERRED` snapshot became stale before
the constrained write could take the writer slot. Roll back that transaction
and retry it with `BEGIN IMMEDIATE`; TypeGraph-owned transactions already use
that mode. The refusal happens before the constraint probe, so the write is
fenced or refused rather than allowed to rely on a stale decision.
#### `BATCH_WRITE_UNSUPPORTED`
A backend whose `capabilities.execution.unitOfWork` is `"batch"` (Cloudflare
D1's `batch()`, Neon HTTP's `transaction(queries)`) fixes every statement
before the first one runs and commits them together with no session in
between. Every fused write on such a backend — a static batch and a
certified atomic program alike — asserts the active schema version inside
the very statement that writes, so a stale version writes nothing and the
store reports `StaleVersionError`. See
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares).
A write that needs more than that one guarded statement refuses, but
`BATCH_WRITE_UNSUPPORTED` is not itself a top-level error code: the
enforcing gate keeps its own class and code
(`CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, `UNSUPPORTED_BACKEND_CAPABILITY`,
`IDENTITY_REQUIRES_ATOMIC_BACKEND`, or a plain `ConfigurationError` for
`history` / `revisionTracking` / a schema commit) and nests
`{ code: "BATCH_WRITE_UNSUPPORTED", reason }` under `details.batchRefusal`,
naming what a closed batch cannot supply:
| `details.batchRefusal.reason` | What it needs | Raised by |
| --- | --- | --- |
| `interactive-callback` | Hold an interactive callback transaction open across several round trips. | `store.transaction(fn)` / `store.transactionWithReceipt(fn)` |
| `constraint-needs-probe` | Read a value it wrote earlier in the same write before deciding what to write next. | A declared constraint's probe-then-write (`CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, above) |
| `identity` | Read and write Operational Identity's closure across several round trips inside one held transaction. | `Store` construction, or `requireAtomicIdentityBackend`, when `graph.identity` is declared |
| `history` | Hold the per-graph write lock and clock open across a whole write cascade. | `history: true` or `revisionTracking: true` |
| `schema-commit` | Hold one transaction across its compare-and-swap read and its activating write. | `commitSchemaVersion` / `setActiveVersion` |
`SCHEMA_WRITE_FENCE_UNSUPPORTED` — the portable schema-version fence an
ineligible write falls back to (see [Schema Migrations](/schema-management))
— does not carry `batchRefusal`. It is reached from many fuse failures that
are not specific to a batch-tier backend (an ineligible write kind, a
tombstone-resurrection write a supplied id falls through to, a derived
backend, a provenance mismatch), so it states its plain limitation without
guessing which of the reasons above, if any, applies.
#### Write-fence declaration codes
`capabilities.writeFence` resolves one of four write-fence plans a lock site
consumes — see
[Write fence declaration](/backend-setup#write-fence-declaration-writefence).
`ConfigurationError` codes name the ways a backend's fence declaration, or
its resolved plan, turns out not to cover what a write needs:
| `details.code` | Raised when |
| --- | --- |
| `WRITE_FENCE_DECLARATION_INVALID` | The declared `writeFence` fails runtime validation: an unrecognized `mechanism` string, an unrecognized `drain` string under `mechanism: "advisory"`, or a `drain` key present on `mechanism: "engine-serialized"` / `"caller-serialized"` (`drain` applies only to `"advisory"`). `details.field` names `"mechanism"` or `"drain"`; for an unrecognized value, `details.accepted` lists the allowed strings. Raised by `resolveWriteFencePlan` before any plan is shaped — an invalid `drain` never falls through to behaving like `"quiescent"`. |
| `WRITE_FENCE_SQL_UNAVAILABLE` | The resolved declaration's `mechanism` is `"advisory"` but the backend's `fenceSql` is missing the member that `mechanism`/`drain` combination needs to spell (`advisoryLockExpression`, `isolationFactExpression`, or, under `drain: "table-lock"`, `lockTables`) — or, independently of any lock plan, a session isolation-level read (recorded capture's isolation guard) finds no `fenceSql` at all. Raised at backend construction for the lock-plan case; at the point of the read for the session-fact case. |
| `RECORDED_CLOCK_REQUIRES_WRITE_FENCE` | The store is constructed with `history: true` or `revisionTracking: true` — TypeGraph-owned recorded-clock allocation — against a backend whose write-fence plan resolves `unfenced`. |
| `WRITE_FENCE_UNAVAILABLE` | A resolved plan cannot satisfy what a specific operation needs: either the plan is `unfenced` outright, or it is a `lock` plan whose `drain` is `"none"` meeting an operation whose `requires` is `"drain"`. `details.operation` names the operation and `details.requires` names which kind of exclusion (`"keyed"` or `"drain"`) it needed; a `drain: "none"` refusal also names the drain in the message. `"engine-serialized"` and `"caller-serialized"` satisfy either `requires` value without consulting `drain`. |
| `CALLER_SERIALIZED_REFUSES_ADOPTION` | `adoptTransaction` was called on a backend whose resolved write-fence plan is `caller-serialized`. An externally owned transaction's lifetime cannot be held by the backend's in-process write-unit queue, so `store.withTransaction(externalTx)` is refused rather than let its writes silently interleave with the queue's own. `details.member` names `"adoptTransaction"`. |
`RECORDED_CLOCK_REQUIRES_WRITE_FENCE` refuses at
`createStore`, never mid-flush, and the message names the exact declaration line to add.
`WRITE_FENCE_UNAVAILABLE` is not a `createStore`-time check: `requireWriteFence` is called from
every individual lock site (the identity graph lock, the identity-enablement drain, identity DDL,
trusted import, contribution DDL, recorded-clock allocation, schema-fence sites, graph-merge
provenance), so it fires wherever one of those runs — inside a live transaction, mid-operation,
not only at `createStore`. `WRITE_FENCE_SQL_UNAVAILABLE` and `WRITE_FENCE_DECLARATION_INVALID`
both refuse earlier, at backend construction for a `createSqlBackend`-built backend (or, for the
session-fact half of `WRITE_FENCE_SQL_UNAVAILABLE`, at the read that needed it), since they are
about the declaration itself rather than what a specific store option or operation requires of it.
`CALLER_SERIALIZED_REFUSES_ADOPTION` fires wherever `adoptTransaction` is actually called, which is
never at `createStore` time. `IDENTITY_REQUIRES_WRITE_FENCE` is another write-fence-related code —
see the Operational Identity guard codes table above — but is not in this table because it guards
identity construction, not recorded-clock allocation.
### Caller-serialized queue codes
The in-process queue a `writeFence: { mechanism: "caller-serialized" }` declaration builds
(`src/backend/serialized-execution-queue.ts`) raises two more `ConfigurationError` codes, both
naming `details.subject` — the SQLite dialect string for SQLite's own per-connection queue, or
`"caller-serialized"` for the write-unit queue a `caller-serialized` declaration builds:
| `details.code` | Raised when |
| --- | --- |
| `SERIALIZED_QUEUE_REENTRANT_SUBMISSION` | A queued operation was awaited from inside a transaction already running on the same queue — the transaction holds the queue's execution slot until it completes, so the nested operation could never run. Use the transaction-scoped context (`tx.nodes` / `tx.edges` / `tx.backend`) instead of the root store or backend inside a `store.transaction` callback, or move the operation outside the transaction. |
| `CALLER_SERIALIZED_REQUIRES_ASYNC_CONTEXT` | The queue's reentrancy detection depends on `node:async_hooks`' `AsyncLocalStorage`, which is unavailable on this runtime (or had not finished loading). A `caller-serialized` write-fence declaration's in-process promise depends on that detection actually working, so every submission is refused rather than run without it. SQLite's own per-connection queue never raises this code: it runs without detection instead of refusing when the context is unavailable. |
These codes are not part of `RECORDED_CAPTURE_GUARD_CODES` — that set is
closed to the three codes documented under
[Recorded-capture guard codes](#recorded-capture-guard-codes) below, and
`isRecordedCaptureGuardError` does not recognize any write-fence code.
### Optimistic-retry unit codes
The retry owner every `"optimistic-retry"`-tier unit of work runs through
(`src/backend/capabilities/retried-unit.ts`) raises one more
`ConfigurationError` code, naming `details.operation` — the same operation
name `TransactionConflictError` reports for the same unit:
| `details.code` | Raised when |
| --- | --- |
| `OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` | The unit's target is on the `"optimistic-retry"` execution tier (see [Backend Capabilities](/backend-setup#backend-capabilities)), and detecting a unit of work nested inside another one depends on `node:async_hooks`' `AsyncLocalStorage`, which is unavailable on this runtime. Running without that detection would let a nested unit's own independent retry commit against reads an outer attempt took before it ever conflicted, so the unit is refused, before its attempt ever runs, rather than run without it. A target on any other execution tier is unaffected: no nested owner exists there, so this code is never raised for it. |
#### Backend capability declaration codes
Custom backend declarations and capability bundles use stable `details.code` values when the
declared surface disagrees with what TypeGraph can safely execute:
| `details.code` | Raised when |
| --- | --- |
| `CAPABILITY_DECLARATION_CONTRADICTION` | `recursiveTraversal.supported` and its `reason` contradict each other: unsupported without a reason, or supported with a dangling reason. |
| `RECURSIVE_TRAVERSAL_UNSUPPORTED` | A backend declares recursive traversal unsupported and a recursive query, subgraph read, or historical identity operation needs it. `details.operation` names the refusing path and `details.reason` echoes the backend declaration. |
| `CONSTRAINT_CLAIM_SURFACE_MISMATCH` | The `constraintClaims` declaration and the claim members implemented by the backend disagree in either direction. |
| `BUNDLE_PORT_SURFACE_MISMATCH` | A non-claim capability bundle resolves a required member as present, but the backend port used by the operation cannot reach it. Fallback-disposition members degrade through their documented fallback instead of throwing this code. |
| `RECORDED_DDL_CONSTRAINT_NAME_MISMATCH` | `recordedTableDdl` names a primary-key constraint for only one of the temporary or final recorded-table name sets. |
The recorded-time preview migration also throws `UnsupportedBackendCapabilityError` with
`details.capability: "recordedTableDdl"` when a legacy schema needs rewriting and the custom
backend does not provide its DDL callback. See
[Migrating Preview Recorded Time](/schema-management#migrating-preview-recorded-time) and
[Capability bundles](/backend-setup#capability-bundles) for the corresponding migration and
backend-author guidance.
#### Approximate retrieval with a mismatched metric
`.similarTo(vector, k, { approximate: true, metric })` is refused with a
`ConfigurationError` when `metric` differs from the field's declared metric. An
ANN structure is built for one metric — `vec0` bakes `distance_metric` into the
virtual table, libSQL's DiskANN index is built with `metric=…`, pgvector's index
carries a per-metric operator class — so retrieving by the declared metric and
re-scoring under the override would return the declared metric's neighbors
wearing the override's scores. The two options state something that cannot both
hold, so the option is refused rather than downgraded to an exact scan behind the
caller's back.
`details` carries `nodeKind`, `fieldPath`, `requestedMetric`, `declaredMetric`,
and `indexType`; there is no stable `details.code`, so match by class and
`details`. A slot declared `indexType: "none"` is **not** refused — there is no
ANN structure to be bound to a metric, and the opt-in compiles to the exact scan,
a degradation stated on the `approximate` option itself. A mismatched metric with
no `approximate` is not refused on the query builder either; `store.search.vector`
and `store.search.hybrid` refuse every mismatched override on their own broader
rule. See
[Approximate retrieval](/semantic-search#approximate-retrieval-for-similarto-opt-in).
#### Durable edge match identity guard codes
Durable edge match identity uses stable `ConfigurationError` detail codes:
| `details.code` | Meaning |
| --- | --- |
| `EDGE_MATCH_IDENTITY_VALUE_NOT_SCALAR` | A declared identity field cannot be represented as a portable JSON scalar, or an untyped runtime value violated that declaration. |
| `EDGE_MATCH_IDENTITY_KEY_TOO_LARGE` | One complete durable identity tuple exceeds the portable 2,000-byte index budget. Normal import records this against the individual edge; trusted import is atomic and refuses the whole stream. |
| `EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE` | The adapter declares durable identity support, but the database is missing its columns or unique arbiter. Initialize or migrate the schema before serving writes. |
| `EDGE_MATCH_IDENTITY_REQUIRES_ATOMIC_BACKEND` | Initial adoption needs an atomic empty-kind fence or materialization preflight that the custom backend does not implement. |
| `DURABLE_EDGE_MATCH_IDENTITY_COMMAND_UNSUPPORTED` | A custom backend declares durable identity support but refuses the authoritative convergence command. TypeGraph fails closed because the portable read-then-write fallback has no equivalent database arbiter. |
| `IMPORT_EDGE_BATCH_RETRY_REQUIRES_SAVEPOINT` | A durable import batch was refused without savepoint rollback protection, either because the backend is non-transactional or because its root/transaction statement-execution contract cannot serve savepoints. TypeGraph will not retry rows individually because that could double-attribute an already-written prefix. |
The last refusal deliberately differs from optional fused-command fallback:
the durable identity declaration delegates correctness to a database key, so a
backend that claims the feature but refuses its command cannot safely re-enter
the dynamic portable path.
#### Heterogeneous edge-read guard codes
`findEdgesByHeterogeneousEndpointSet` refuses mixed endpoint modes instead of
guessing how incident and exact-pair rows should be interpreted:
| `details.code` | Meaning |
| --- | --- |
| `EDGE_HETEROGENEOUS_READ_MIXED_ENDPOINT_MODES` | The request contains both incident-endpoint rows (without an opposite endpoint) and exact directed-pair rows. Supply an opposite endpoint for every row to request exact-pair matching. |
| `EDGE_HETEROGENEOUS_READ_BIND_BUDGET_EXCEEDED` | The endpoint set cannot fit within the backend's bind-parameter budget. Split the request into smaller calls. |
#### Operational Identity guard codes
Operational Identity lifecycle failures use stable `details.code` values on
`ConfigurationError`:
| `details.code` | Meaning |
| --- | --- |
| `IDENTITY_REQUIRES_ATOMIC_BACKEND` | The selected adapter cannot provide the interactive transaction required by identity writes. |
| `IDENTITY_REQUIRES_STATEMENT_EXECUTION` | The backend cannot execute the raw statements Operational Identity issues internally. |
| `IDENTITY_REQUIRES_WRITE_FENCE` | Operational Identity was constructed against a backend whose `capabilities.writeFence` resolves `unfenced` — declare the capability, matching the engine's real locking support. See [Write-fence declaration codes](#write-fence-declaration-codes). |
| `IDENTITY_NOT_ENABLED` | `store.identity`, `tx.identity`, `StoreView.identity`, or an identity-expanded query option was reached on a graph without `identity: { ... }` — normally caught at compile time; this is the runtime guard for a widened or `any`-typed handle. |
| `IDENTITY_STORAGE_MISSING` | An identity relation disappeared after enablement, or exists without this graph's fill. Restore ledgers, or recreate and rebuild the derived closure, before serving traffic. `details.reason: "unfilled"` marks the second case: the separation relation is present but holds no row for this graph while the ledger holds a live `different` assertion across two distinct identity classes — reopen the Store (the open runs the fill) or run `rebuildIdentityClosure(store)`. A Store handle opened while the relation did not exist keeps failing until it is reopened, which is deliberate: the alternative is a confident "not separated" the moment another graph's upgrade creates the shared relation. |
| `IDENTITY_UPGRADE_REQUIRES_ATOMIC_DDL` | The backend cannot publish the derived separation relation's upgrade — the `CREATE` and the fill — as one commit, on a graph that owes rows. `details.missingPorts` names what is absent: `schemaWriteTransaction` / `identityTableDdl` on the fenced path, or `executeSchemaDdl` on the schema-commit path. Refused rather than degraded, because a relation created empty and filled afterwards reads as "nothing is separated" in between. Both bundled Drizzle backends implement all three when transactions are enabled, so this is a custom-backend path. |
| `IDENTITY_ENABLEMENT_PENDING` | First enablement is pending because `autoMigrate` is disabled. |
| `IDENTITY_PROFILE_MIGRATION_PENDING` | A `sameIdAcrossKinds` change (a breaking `fold`↔`ignore` flip, or disabling identity) has not been applied — either it is breaking, or `autoMigrate` is disabled. |
| `IDENTITY_SCHEMA_MIGRATION_PENDING` | An identity-relevant ontology change is pending because `autoMigrate` is disabled. |
| `IDENTITY_SEPARATION_VIOLATION` | The derived separation relation refused a write that would place both endpoints of a current `different` assertion in one identity class. The database-level backstop beneath identity validation; reaching it means an earlier guard let a contradiction through. |
| `IDENTITY_TRANSACTION_NOT_WRITE_FENCED` | SQLite refused an identity write because the enclosing transaction was begun `DEFERRED` and another connection committed before it could take the writer slot. Only reachable through `store.withTransaction(externalTx)` / `store.withRecordedTransaction(externalTx)`, where the caller owns the `BEGIN` — TypeGraph's own transactions open `BEGIN IMMEDIATE` and hold the slot from the start. SQLite cannot upgrade a stale snapshot in place, so roll back and re-run the transaction, opening it with `BEGIN IMMEDIATE`. |
| `IDENTITY_SCHEMA_CONTRADICTION` | Existing nodes or assertions contradict the proposed identity profile or ontology, or the materialized closure disagrees with the assertions it was derived from. Run `rebuildIdentityClosure(store)` to recover from a closure mismatch. |
| `IDENTITY_IMPORT_REQUIRES_PROFILE` | An interchange document carries an `identity` section but the target graph does not have the profile enabled. |
| `IDENTITY_MERGE_REQUIRES_PROFILE` | A branch carries identity changes but the merge target graph does not have the profile enabled. |
| `IDENTITY_EXPORT_REQUIRES_TEMPORAL_FIELDS` | An identity-enabled export explicitly disabled temporal fields. Remove `includeTemporal` or set it to `true`; endpoint bounds are required to validate assertion windows on import. |
| `IDENTITY_IMPORT_ID_CONFLICT` | An imported assertion id already exists in the target ledger identifying different truth (relation, endpoints, or validity window). |
| `RECORDED_IDENTITY_SCHEMA_MISSING` | A `history: true` open of an identity-enabled graph could not find the recorded identity relation. Bundled backends provision it, so this is rare there and more likely on a custom backend. |
When an unapplied migration's **only** breaking change is the identity one, the
specific pending code above wins over the generic `MigrationError` (which is
attached as `cause`); a diff that also breaks nodes, edges, ontology, or
indexes raises the generic `MigrationError` enumerating all of them.
Identity import also raises `ValidationError` with one of these
`details.issues[].code` values when an interchange document's `identity`
section fails shape or integrity checks. Each issue carries the offending
assertion's id structurally in `details.issues[].assertionId`, and
`importGraph`/`importGraphStream` record these failures as
`entityType: "identity"` entries in `result.errors` (a self-assertion —
`IDENTITY_SELF_ASSERTION` — included) rather than throwing:
| Issue `code` | Meaning |
| --- | --- |
| `IDENTITY_IMPORT_UNKNOWN_KIND` | An assertion endpoint names a node kind not in the target graph's registry. |
| `IDENTITY_IMPORT_PAIR_NOT_NORMALIZED` | An assertion's `a`/`b` endpoints are not in code-point order. |
| `IDENTITY_STATE_IMPORT_ENDED_ASSERTION` | A `state`-mode import (the default) contains an already-ended assertion; use `identityMode: "archival"` on export to carry ended assertions. |
| `IDENTITY_IMPORT_FUTURE_VALID_FROM` | An open (current) assertion's `validFrom` is in the future, in either import mode. |
| `IDENTITY_IMPORT_FUTURE_VALID_TO` | An ended assertion's `validTo` is in the future. |
| `IDENTITY_IMPORT_INVALID_WINDOW` | An assertion's `validTo` precedes its `validFrom`. |
| `IDENTITY_IMPORT_ENDED_BY_WITHOUT_END` | An assertion names an `endedBy` cause but carries no `validTo`; only an ended assertion has a cause. |
| `IDENTITY_IMPORT_ENDED_BY_NOT_ENDPOINT` | An assertion's `endedBy` names a node that is not one of its own endpoints; a deletion cascade only ends assertions that touch the deleted node. |
| `IDENTITY_SELF_ASSERTION` | An assertion's `a` and `b` name the same node. |
#### Merge provenance sidecar codes
`persistProvenance: true` writes to a *sidecar* graph beside the merge target,
and `openProvenanceStore` refuses any sidecar graph id it cannot prove it owns.
Both refusals are `ConfigurationError`s with a stable `details.code`, and both
carry `details.graphId` (the sidecar id) and `details.targetGraphId`:
| `details.code` | Meaning |
| --- | --- |
| `GRAPH_MERGE_PROVENANCE_ID_COLLISION` | The sidecar graph id is occupied by something this library did not write. `details.reason` names which state was found, and the suggestion is specific to it. |
| `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED` | The backend exposes no transactional schema fence (`schemaWriteTransaction`), so the id's emptiness check and its ownership-marker write cannot commit as one unit. Not a collision — the id may well be free. An already-owned sidecar still opens on such a backend, so read-only use of an existing sidecar stays available. |
The five `details.reason` values on `GRAPH_MERGE_PROVENANCE_ID_COLLISION`:
| `details.reason` | The state that was found |
| --- | --- |
| `application-graph` | The id holds rows (in any per-graph table) or a schema that is not the sidecar's, so it belongs to an application. When a pre-marker sidecar is classified, revision-change journal entries that record its own stored `Provenance` rows are not counted; every other journal entry is. Rename the colliding graph or point the merge elsewhere. |
| `empty-legacy-sidecar` | A pre-marker sidecar with no rows at all, which carries no evidence of authorship and is indistinguishable from an application graph of the same shape. |
| `unupgradeable-legacy-sidecar` | A pre-marker sidecar whose rows do not verify as provenance this library wrote for *this* target, so it cannot be upgraded to an owned sidecar. |
| `unowned-exact-schema-graph` | The current sidecar schema with no ownership marker. Because the marker is written *first*, this library cannot have produced this state; contents are not consulted, so an empty or provenance-shaped occupant is refused too. |
| `corrupt-ownership-marker` | A `ProvenanceOwner` row that is not a valid live claim for this target — soft-deleted, schema-invalid, naming a different target, or stored under a different row id. It is never overwritten or resurrected, because it may be an application's row. |
Under `persistProvenance: true` these arrive wrapped: the sidecar is opened and
claimed **before** the merge commits, and either code refuses the merge as an
`InvalidMergeOptionsError` (`details.option: "persistProvenance"`,
`details.provenanceErrorCode` echoing the code above, the `ConfigurationError`
as `cause`) with the target left unmodified. Only transient row-write failures
after the commit degrade to a `warnings` entry.
#### Interchange serialized-connection guard codes
Two long-lived interchange streams cannot share one serialized database
connection: an export snapshot holds a read transaction for the whole stream and
a streaming import writes a transaction per chunk on that same connection, so
the second one either nests a `BEGIN` or waits for a slot that never frees. The
lease is **exclusive** — one stream of any kind per connection — so all four
pairings refuse with a `ConfigurationError` rather than hanging:
| `details.code` | Raised when |
| --- | --- |
| `INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` | An export snapshot holds the connection, detected through the shared serialized resource the two backend wrappers were marked with. |
| `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT` | The same condition, reported by the object-identity detector: one SQLite backend is exporting into itself. Worth telling apart because the fix differs — pass a second backend rather than await whatever else is running. |
| `INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS` | A streaming import holds the connection, in either order of discovery. |
The code names *what holds the connection*; `details.requested` and
`details.heldBy` (each `"export-snapshot"` or `"import-stream"`) name which
pairing was actually refused, so a same-kind refusal is never reported as
something it is not. `details.graphId` names the graph the refused stream was
for.
`"import-stream"` is the kind of every long-lived import, not only
`importGraphStream`: `importGraph` holds the lease for the whole call, and
`trustedImportGraph` / `trustedImportGraphStream` hold it for the whole trusted
session — so those APIs throw this `ConfigurationError` as well as their own
`TrustedImportError`. Connections TypeGraph cannot observe are not refused: two
clients dialed at one server, or two SQLite handles on one file, are genuinely
independent. See
[Scaling branches and interchange](/graph-merge#scaling-branches-and-interchange)
for which drivers are recognized as serialized.
##### Declaring a connection the driver hides
Recognition is a duck-type over the client object, so a serialized driver
TypeGraph cannot identify (`expo-sqlite`, `op-sqlite`, `sqlite-proxy`,
`pg-proxy`, Bun `SQL`, a postgres-js client capped through a string it does not
coerce) is left unmarked and its stream pairs are not refused.
`createSqliteBackend` and `createPostgresBackend` accept a `serializedResource`
declaration for that gap — `{ mode: "shared", resource: client }` — and for the
reverse case, `{ mode: "independent" }`, when the detection is wrong for your
topology. See [Serialized connections](/backend-setup#serialized-connections).
The declaration is applied or refused, never quietly ignored:
| Declaration | Outcome |
| --- | --- |
| `{ mode: "shared", resource }` on a connection TypeGraph did not detect, or naming the client it did detect | The named object is the serialized resource; two backends naming the same object are one connection |
| `{ mode: "shared", resource }` naming a **different** object than the one detected | `ConfigurationError` (`code: "CONFIGURATION_ERROR"`) from the factory, with `details.reason: "serialized-resource-conflict"` and `details.declaredKind` / `details.detectedKind` naming what each side was |
| `{ mode: "independent" }` | Honored, whatever was detected — the documented escape hatch |
The conflict is refused rather than resolved because two wrappers over one
connection given two different sentinels would stop being seen as a pair, which
is precisely the refusal this guard exists to make.
The two `*Kind` details are constructor names (`"Database"`, `"BoundPool"`), not
the handles themselves: `details` is what `toLogString()` serializes, and a
driver handle there would print whatever that driver stores — a `pg.Pool` keeps
its `connectionString`, password included — into your logs.
`{ mode: "independent" }` lifts the shared-resource arm between two distinct
backend objects. It does **not** lift
`INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`: one SQLite backend exporting into
itself holds the one snapshot transaction its own import writes through, which
is a fact about a single handle rather than a claim about connection topology.
Pass a second backend for that case. That surviving refusal is SQLite-only, so
on PostgreSQL a backend declared independent exporting into itself is not
refused either — a client that hands out independent connections is exactly
what the declaration claims.
#### `ExportStreamCancelledError`
An export stream whose `signal` fires settles with `ExportStreamCancelledError`
(`code: "INTERCHANGE_EXPORT_STREAM_ABORTED"`) rather than a silent end of stream,
so a consumer never mistakes a cancelled export for a complete one. It is thrown
only *after* the export has given back everything it took, so receiving it means
the connection is already free. What that was depends on the backend: a
transactional one rolls back the snapshot and releases the connection's stream
lease; one without transactions held neither and simply abandons its remaining
reads, its delivered chunks never having been a single snapshot. The message
says which.
`details.graphId` names the exported graph and `cause` carries the signal's own
`reason` when the caller supplied one. A signal that is already aborted refuses
the export before any transaction is opened. See
[Cancelling an export](/interchange#cancelling-an-export).
#### `ExportStreamIdleTimeoutError`
An `exportGraphStream` configured with `idleTimeoutMs` settles with
`ExportStreamIdleTimeoutError` (`code: "INTERCHANGE_EXPORT_STREAM_IDLE_TIMEOUT"`)
when its consumer does not request another chunk within that bound. The timeout
measures only the interval after a chunk is yielded; time spent waiting for the
backend to produce the next chunk does not count. `details.graphId` identifies
the graph and
`details.idleTimeoutMs` carries the configured bound. As with explicit
cancellation, a transactional export rolls its snapshot back and releases its
stream lease before the error is delivered; a non-transactional export held
neither and abandons its remaining reads. See
[Cancelling an export](/interchange#cancelling-an-export).
#### Recorded-capture guard codes
`ConfigurationError` is intentionally open-shaped, but the guards that fire on a
`history: true` / `revisionTracking: true` store carry a **stable, branchable
`details.code`** so a portable caller does not have to substring-match the
message. The three codes are exported as a set, `RECORDED_CAPTURE_GUARD_CODES`,
and reachable through the `isRecordedCaptureGuardError` type guard:
| `details.code` | Raised when |
|----------------|-------------|
| `RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION` | `store.withTransaction(externalTx)` on a history-enabled store — it has no flush point before the caller commits. Use `store.withRecordedTransaction(externalTx, fn)`. (Also a compile error on an `AdapterHistoryStore`.) |
| `RECORDED_CAPTURE_RAW_SQL_DISABLED` | A raw SQL escape (`tx.sql`, `backend.executeStatement` / `executeDdl`) on a history-enabled store, where it would bypass recorded-time capture. |
| `REVISION_TRACKING_RAW_SQL_DISABLED` | The same raw SQL escape on a revision-tracked store, where it would bypass the revision anchor. |
Typed code cannot call `withTransaction` on an `AdapterHistoryStore`; use
`withRecordedTransaction` directly. The runtime code remains useful at
JavaScript and deliberately untyped boundaries. If one of those boundaries
throws, `isRecordedCaptureGuardError(error,
"RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION")` narrows both the error and
its `details.code` without message matching.
Pass a specific code to narrow to one guard, or omit it to match any. The guard
narrows `error` to a `ConfigurationError` whose `details.code` is the passed
`RecordedCaptureGuardCode` (or the full union when no code is given), so no
untyped `details` spelunking is needed.
This composes with
[`tx.sqlAvailability`](/queries/temporal/#raw-sql-under-history-capture): the
discriminant tells a caller *why* `tx.sql` is unusable ahead of time
(`"history"` / `"revisionTracking"` vs. `"unavailable"` for a backend with no
transactions), while the guard code identifies a guard that has already thrown.
Between them, "history capture forbids raw SQL here" and "this backend has no
transactions" (which carries **no** guard code) are cleanly distinguishable
without catching-and-string-matching.
#### Engine-native recorded-time codes
A backend can track recorded (system) time itself by declaring
`GraphBackend.recordedTime` instead of using TypeGraph's own recorded
relations and clock — see [Engine-native recorded
time](/queries/temporal#engine-native-recorded-time) and [Supplying
`recordedTime`](/backend-authoring#supplying-recordedtime). Which ownership
form a store reads under is derived from that member's presence, never
declared separately, so there is no `recordedTimeOwnership` option to set.
Every refusal specific to that form carries a stable `details.code`:
| `details.code` | Raised when |
|----------------|-------------|
| `ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE` | An engine profile declares `recordedTime` without also declaring `lineage` — engine-native history keeps no recorded relations of its own for TypeGraph to derive a graph-merge change delta from. Raised at backend construction, naming both members. |
| `RECORDED_TIME_UNAVAILABLE` | A caller reached `requireRecordedTime` and found `recordedTime` absent on the backend it asked — store construction under `history: true` and the shared `recordedNow()`/`revisionNow()`/receipt-stamping read, both reached only once ownership has already resolved to `"engine-native"`. |
| `ENGINE_NATIVE_REVISION_TRACKING_UNSUPPORTED` | A store is constructed with `revisionTracking: true` against an engine-native backend, whether or not `history: true` is also requested — there is no TypeGraph clock for `revisionTracking` to advance; the engine's own revision is available only under `history: true`. |
| `ENGINE_NATIVE_RECORDED_READ_UNSUPPORTED` | A store is constructed with an external `recordedRead` binding against an engine-native backend — there is no TypeGraph recorded relation for one to populate. |
| `ENGINE_NATIVE_RECORDED_IDENTITY_UNSUPPORTED` | `store.identityAtCoordinate` at a past recorded instant, or the query compiler's historical identity traversal, is reached under engine-native recorded time — identity history reads TypeGraph's own recorded relations directly, which an engine-native backend does not populate. |
| `ENGINE_NATIVE_MIGRATE_RECORDED_TIME_UNSUPPORTED` | `migrateLegacyRecordedTime` is called against an engine-native backend — the migration rewrites TypeGraph's own recorded relations, which an engine-native backend does not have. |
| `RECORDED_INSTANT_OWNERSHIP_MISMATCH` | `store.asOfRecorded(instant)` receives an instant minted under the OTHER recorded-time ownership form — an `r1:` (TypeGraph-owned) instant against an engine-native store, or an `e1:` (engine-native) instant against a TypeGraph-owned store. |
The profile refusal fires at backend construction; the two store-option
refusals, and `RECORDED_TIME_UNAVAILABLE`'s construction arm, fire at
`createStore`; the remaining codes, and `RECORDED_TIME_UNAVAILABLE`'s read
arm, fire at the specific call that cannot be honored. None of these codes
are members of `RECORDED_CAPTURE_GUARD_CODES` above — that set stays closed
to the three TypeGraph-capture guards.
### `SchemaMismatchError`
Thrown when the database schema doesn't match the expected graph definition.
```typescript
try {
const [store] = await createStoreWithSchema(graph, backend);
} catch (error) {
if (error instanceof SchemaMismatchError) {
console.log(error.category); // "system"
console.log(error.details);
// { graphId: "my-graph", expectedHash: "", actualHash: "" }
console.log(error.suggestion);
// "Run migrations to update the database schema..."
}
}
```
### `MigrationError`
Thrown when schema migration fails due to breaking changes that require manual intervention.
The `details.reason` value `"edge-match-identity-rekey"` means a populated edge kind
changed or newly adopted its durable match identity. Existing rows cannot be assigned
new identity keys without choosing how conflicts converge. Export the affected edges,
hard-delete them, apply the schema migration, then reimport them so TypeGraph
materializes and arbitrates the new durable keys.
```typescript
try {
const [store] = await createStoreWithSchema(graph, backend);
} catch (error) {
if (error instanceof MigrationError) {
console.log(error.category); // "system"
console.log(error.details);
// { graphId: "my-graph", fromVersion: 3, toVersion: 4, reason: "Removed required field 'email' from Person" }
console.log(error.suggestion);
// "Review the breaking changes and perform manual migration if needed..."
}
}
```
### `BaseSchemaMigrationError`
Thrown by zero-DDL verified and graph-template entry points when the
deployment-wide physical TypeGraph schema has not been adopted to the version
required by the running library. This is separate from `MigrationError`, which
describes one graph's serialized schema evolution.
```typescript
try {
const [store] = await createVerifiedStore(graph, backend);
} catch (error) {
if (error instanceof BaseSchemaMigrationError) {
console.log(error.details);
// {
// installedVersion: undefined,
// requiredVersion: 1,
// reason: "missing"
// }
}
}
```
`reason` is `"missing"`, `"stale"`, or `"newer"`. For missing or stale
storage, run `createStoreWithSchema()` or `createAdapterStoreWithSchema()` once
under a DDL-capable role, or apply the published external base-schema migration.
A newer marker requires a TypeGraph release that supports that version.
## Query Errors
### `UnsupportedPredicateError`
Thrown when using a query predicate that isn't supported by the current backend.
```typescript
// Using vector similarity on a backend without vector support:
try {
await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.similarTo(queryVector, 10))
.select((ctx) => ctx.d)
.execute();
} catch (error) {
if (error instanceof UnsupportedPredicateError) {
console.log(error.category); // "system"
console.log(error.suggestion);
// "Use a backend that supports this predicate, or rewrite the query..."
}
}
```
## Transaction Errors
### `TransactionClosedError`
Thrown when a statement reaches a transaction-scoped backend after its
transaction boundary has already returned.
A transaction pins one database connection, which carries one statement at a
time. When `store.transaction(...)` resolves or rejects, the driver emits
`COMMIT` or `ROLLBACK` on that connection and hands it back to the pool. Any
statement still in flight then has nowhere safe to go — it would execute inside
somebody else's transaction — so TypeGraph refuses it.
The usual source is a callback that lets work escape it. `Promise.all` rejects
on its first rejection while its siblings keep running:
```typescript
await store.transaction(async (tx) => {
// If `a` fails, `b`'s remaining statements are orphaned.
await Promise.all([tx.nodes.Doc.create(a), tx.nodes.Doc.create(b)]);
});
```
You will normally never see this error: `Promise.all` has already rejected with
the original failure and discards the orphan's. It surfaces only if you await
the orphaned promise yourself. To avoid orphaning writes at all, use
`Promise.allSettled` and inspect the results, or await the writes in sequence.
`adoptTransaction()` never closes its queue — only the caller knows when their
transaction ends — so this error cannot arise there. It remains the caller's
job to await every graph write before committing.
### `TransactionConflictError`
Thrown when a transaction was aborted by a serialization failure or deadlock
on every attempt available to it. `details.operation` names the transaction
that failed, `details.attempts` the number tried, and `cause` is the last
attempt's driver error — PostgreSQL's own protocol for both conditions is to
re-run the whole transaction from the top, which is what this error reports
as exhausted.
`store.transaction()` and `store.transactionWithReceipt()` raise it with
`attempts: 1` for a conflict on their single attempt; passing
`retry: { attempts }` (see
[Retrying on conflict](/schemas-stores/#retrying-on-conflict)) raises it only
once every attempt has conflicted. Graph-merge's commit paths raise
`MergeError` on the same exhaustion, with a `TransactionConflictError` as its
`cause`.
```typescript
try {
await store.transaction(fn, { retry: { attempts: 3 } });
} catch (error) {
if (error instanceof TransactionConflictError) {
console.log(error.details.attempts); // 3
console.log(error.cause); // the last driver error
}
}
```
**The serialization covers TypeGraph's own statements, not `tx.sql`.** The raw
Drizzle handle you get for writing your own relational tables in the same
transaction shares the one pinned connection but bypasses the queue. Running a
raw statement concurrently with a graph write — or with another raw statement —
races two queries on that connection (the overlap `pg@9` removes), and the
boundary cannot drain a raw statement it never saw. Await each `tx.sql`
statement before the next write.
## Error Handling Patterns
### Using Error Utilities
TypeGraph provides utility functions for common error handling patterns:
```typescript
import {
isTypeGraphError,
isUserRecoverable,
isConstraintError,
isSystemError,
getErrorSuggestion,
} from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create(data);
} catch (error) {
if (!isTypeGraphError(error)) {
// Not a TypeGraph error, handle differently
throw error;
}
// Get suggestion regardless of error type
const suggestion = getErrorSuggestion(error);
if (isUserRecoverable(error)) {
// User can fix this by providing different input
return {
error: error.toUserMessage(),
suggestion,
};
}
if (isConstraintError(error)) {
// Business rule violation
return {
error: "This operation violates a constraint",
details: error.details,
};
}
if (isSystemError(error)) {
// Infrastructure/configuration issue
console.error(error.toLogString());
throw error;
}
}
```
### Catch Specific Errors
```typescript
import {
ValidationError,
NodeNotFoundError,
DisjointError,
} from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create(data);
} catch (error) {
if (error instanceof ValidationError) {
// Handle validation failure with contextual details
return {
error: "Invalid data",
issues: error.details.issues,
entity: error.details.kind,
};
}
if (error instanceof DisjointError) {
// Handle constraint violation
return { error: "ID already used by different type" };
}
throw error; // Re-throw unexpected errors
}
```
### Check Error Codes
```typescript
try {
await store.nodes.Person.update(id, data);
} catch (error) {
if (error instanceof TypeGraphError) {
switch (error.code) {
case "NODE_NOT_FOUND":
return { error: "Person not found" };
case "VALIDATION_ERROR":
return { error: "Invalid data", issues: error.details.issues };
default:
throw error;
}
}
throw error;
}
```
### Transaction Error Handling
```typescript
try {
await store.transaction(async (tx) => {
const person = await tx.nodes.Person.create({ name: "Alice" });
const company = await tx.nodes.Company.create({ name: "Acme" });
await tx.edges.worksAt.create(person, company, { role: "Engineer" });
});
} catch (error) {
// Transaction is automatically rolled back on any error
if (error instanceof ValidationError) {
console.log("Validation failed, transaction rolled back");
console.log("Failed on:", error.details.kind, error.details.operation);
}
throw error;
}
```
## Contextual Validation Utilities
For library authors or advanced use cases, validation utilities are available from the schema sub-export:
```typescript
import {
validateNodeProps,
validateEdgeProps,
wrapZodError,
createValidationError,
} from "@nicia-ai/typegraph/schema";
// Validate node properties with full context
const validated = validateNodeProps(PersonSchema, inputData, {
kind: "Person",
operation: "create",
});
// Wrap a Zod error with TypeGraph context
try {
schema.parse(data);
} catch (zodError) {
throw wrapZodError(zodError, {
entityType: "node",
kind: "Person",
operation: "update",
id: "person-123",
});
}
```
## Error Codes Reference
| Code | Error Class | Category | Description |
|------|-------------|----------|-------------|
| `VALIDATION_ERROR` | `ValidationError` | user | Schema validation failed |
| `DISJOINT_ERROR` | `DisjointError` | constraint | Disjointness constraint violated |
| `IDENTITY_CONTRADICTION` | `IdentityContradictionError` | constraint | Identity mutation would make the assertion ledger contradictory |
| `IDENTITY_VALIDITY_FUTURE_START` | `IdentityValidityWindowError` | user | Identity assertion starts after the operation clock |
| `IDENTITY_VALIDITY_FUTURE_END` | `IdentityValidityWindowError` | user | Identity assertion ends after the operation clock |
| `IDENTITY_VALIDITY_INVERTED` | `IdentityValidityWindowError` | user | Identity assertion ends before it starts |
| `IDENTITY_VALIDITY_OPEN_WINDOW_CONFLICT` | `IdentityValidityWindowError` | constraint | A different open window already represents the current semantic pair |
| `IDENTITY_ENDPOINT_VALIDITY` | `IdentityEndpointValidityError` | constraint | An endpoint does not cover the explicit assertion window |
| `GRAPH_MERGE_IDENTITY_CONFLICT` | `IdentityMergeConflictError` | system | Branches carry opposing identity truth |
| `GRAPH_MERGE_CONSTRAINT_CONFLICT` | `MergeConstraintConflictError` | constraint | The resolved merge would violate a store constraint |
| `ENDPOINT_ERROR` | `EndpointError` | constraint | Invalid edge endpoint types |
| `ENDPOINT_PAIR_ERROR` | `EndpointPairError` | constraint | Undeclared source/target combination |
| `CARDINALITY_ERROR` | `CardinalityError` | constraint | Cardinality constraint violated |
| `UNIQUENESS_VIOLATION` | `UniquenessError` | constraint | Uniqueness constraint violated |
| `EDGE_MATCH_IDENTITY_CONFLICT` | `EdgeMatchIdentityConflictError` | constraint | A direct edge write collided with its declared endpoint/property identity |
| `NODE_NOT_FOUND` | `NodeNotFoundError` | user | Referenced node doesn't exist |
| `EDGE_NOT_FOUND` | `EdgeNotFoundError` | user | Referenced edge doesn't exist |
| `KIND_NOT_FOUND` | `KindNotFoundError` | user | Unknown node/edge type |
| `ENDPOINT_NOT_FOUND` | `EndpointNotFoundError` | user | Edge endpoint node doesn't exist |
| `RESTRICTED_DELETE` | `RestrictedDeleteError` | constraint | Delete blocked by existing edges |
| `CONFIGURATION_ERROR` | `ConfigurationError` | system | Invalid configuration |
| `SCHEMA_MISMATCH` | `SchemaMismatchError` | system | Database schema mismatch |
| `MIGRATION_ERROR` | `MigrationError` | system | Migration failed |
| `BASE_SCHEMA_MIGRATION_REQUIRED` | `BaseSchemaMigrationError` | system | Deployment-wide base storage requires privileged adoption |
| `UNSUPPORTED_PREDICATE` | `UnsupportedPredicateError` | system | Predicate not supported |
| `UNSUPPORTED_BACKEND_CAPABILITY` | `UnsupportedBackendCapabilityError` | user | The backend does not advertise a capability the call needs. `details.capability` names it — for example `vector.searchFrontierTuning` for `efSearch` on any SQLite vector or hybrid search, where the engine has no per-search ANN frontier, with `details.reason` naming the limitation |
| `INTERCHANGE_EXPORT_STREAM_ABORTED` | `ExportStreamCancelledError` | user | An export stream's `signal` fired, after the export gave back everything it took. On a transactional backend that is the snapshot transaction and the connection's stream lease; on one without transactions the export held neither and simply abandoned its remaining reads. The message says which |
| `INTERCHANGE_EXPORT_STREAM_IDLE_TIMEOUT` | `ExportStreamIdleTimeoutError` | user | An export stream's consumer left a delivered chunk unacknowledged past its configured `idleTimeoutMs`; the export settled its snapshot and lease before reporting the timeout |
| `TRANSACTION_CONFLICT` | `TransactionConflictError` | system | A transaction was aborted by a serialization failure or deadlock on every attempt available to it. `details.attempts` is the number tried; `cause` is the last driver error |
# Fulltext Search
> BM25-style fulltext search with hybrid retrieval for RAG applications
TypeGraph supports fulltext search directly in your SQLite or PostgreSQL
database — no external search service required. Combine it with semantic
search to get **hybrid retrieval**: the gold-standard pattern for RAG
applications.
## Overview
Vector search is great at finding *semantically* similar content, but it
misses exact matches: proper nouns, SKU numbers, code identifiers, rare
technical terms. Fulltext search handles those. Running both and fusing
the results with Reciprocal Rank Fusion typically beats either approach
alone.
**Key capabilities:**
- Declare `searchable()` string fields in your Zod schema
- Native BM25 ranking (SQLite FTS5) and `ts_rank_cd` (PostgreSQL tsvector)
- Google-style query syntax: quoted phrases, `-excluded`, `OR`
- `n.$fulltext.matches()` predicate composes with metadata filters and graph traversal
- Hybrid search via `$fulltext.matches()` + `.similarTo()` in one query, fused with RRF
- Tunable RRF via `.fuseWith({ k, weights })` on the query builder, or `store.search.hybrid({ fusion })`
## Use Cases
### Hybrid RAG
Combine exact-match retrieval with semantic similarity:
```typescript
const hits = await store.search.hybrid("Document", {
limit: 10,
vector: {
fieldPath: "embedding",
queryEmbedding: await embed(question),
},
fulltext: { query: question },
});
const context = hits.map((h) => h.node.content).join("\n\n");
```
### Multi-tenant fulltext with metadata filters
The most important composition — `$fulltext.matches()` in the same query as
any other predicate:
```typescript
const results = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext.matches("climate change", 20).and(d.tenantId.eq(tenant.id))
)
.select((ctx) => ctx.d)
.execute();
```
### Authorised search via graph traversal
Only return documents the user is allowed to read:
```typescript
const results = await store
.query()
.from("User", "u")
.whereNode("u", (u) => u.id.eq(currentUserId))
.traverse("canRead", "e")
.to("Document", "d")
.whereNode("d", (d) => d.$fulltext.matches(userQuery, 10))
.select((ctx) => ctx.d)
.execute();
```
## Schema Design
### Declaring Searchable Fields
Use `searchable()` to mark string fields for fulltext indexing:
```typescript
import { defineNode, searchable } from "@nicia-ai/typegraph";
import { z } from "zod";
const Document = defineNode("Document", {
schema: z.object({
title: searchable({ language: "english" }),
body: searchable({ language: "english" }),
tenantId: z.string(),
published: z.boolean(),
}),
});
```
### How Indexing Works
TypeGraph stores one fulltext row per node. When you create or update a
node, the values of every `searchable()` field are concatenated and
indexed as a single document. This single-document-per-node design lets a
single query find matches that span multiple source fields — a title hit
plus a body hit both contribute to the same score.
- **PostgreSQL**: the `typegraph_node_fulltext` table carries a
`tsvector` column populated at INSERT time, with a GIN index.
- **SQLite**: the same shape is backed by an FTS5 virtual table with
BM25 ranking.
Sync is automatic — the fulltext index stays in sync with node data
through every `create`, `update`, `upsert`, and `delete` (soft and hard).
### Searchable Options
```typescript
searchable({
language: "english", // Postgres regconfig / SQLite FTS5 tokenizer
})
```
- **`language`**: Postgres uses this as the `regconfig` for stemming
(`english`, `spanish`, `french`, etc.). SQLite FTS5 tokenizer is fixed
at table creation time, so the language is stored but treated as
metadata.
### Adding `searchable()` to an Existing Graph
When you add `searchable()` to a field on a node kind that already has
rows in production, those pre-existing rows are not indexed until you
backfill the index:
```typescript
const stats = await store.search.rebuildFulltext();
console.log(
`Upserted ${stats.upserted}, cleared ${stats.cleared}, ` +
`skipped ${stats.skipped} across ${stats.kinds.length} kinds`,
);
if (stats.skippedIds && stats.skippedIds.length > 0) {
console.warn("Nodes with corrupt props were skipped:", stats.skippedIds);
}
// For systemic corruption, raise the cap to collect the full list:
const forensic = await store.search.rebuildFulltext(undefined, {
maxSkippedIds: 1_000_000,
});
```
`store.search.rebuildFulltext()` iterates nodes with keyset pagination on `id`
(stable under shared timestamps and light concurrent writes), transacts
per page, and cleans up stale fulltext rows for soft-deleted nodes.
Rebuild is a maintenance operation: concurrent hard-deletes between page
fetches can be missed by a single pass. Run during a maintenance window
for full consistency. Scope to a single kind with
`store.search.rebuildFulltext("Document")` to avoid scanning unrelated
data.
Also useful for:
- Recovering after a `DROP TABLE` / `TRUNCATE` of the fulltext table.
- Re-tokenizing after changing `language` on a `searchable()` field.
- Recovering from bulk inserts that bypassed the store layer.
### Checking Whether Search Is Ready
`store.search.rebuildFulltext()` fixes *content*. It cannot fix storage
that is missing, unattested, or provisioned at the wrong shape — and it
throws `StoreNotInitializedError` when it is, because the hot-path gate
refuses fulltext writes until both the deployment marker attests the shared
table and the graph-local activation marker admits this graph. To find out
which situation you are in without writing anything:
```typescript
const health = await store.probeContributions();
const fulltext = health.entries.find(
(entry) => entry.contribution === "fulltext",
);
if (fulltext?.state !== "ready") {
console.warn(`fulltext search is ${fulltext?.state}`, fulltext?.detail);
}
```
The probe writes nothing, so it is safe to call from a health check, on a
read path, or on a replica — which is the point: the alternative was to
run `store.repairContributions()`, a write with repair side effects, and
hope. When it reports `degraded`, escalate through the contribution health
ladder — probe, then `repairContributions()`, then
`rebuildContribution("fulltext")`, which is scoped to the calling graph: it
deletes and refills only that graph's rows in the shared fulltext table, and
drops and recreates the table itself only when no other graph has rows in
it. The three rungs, what each one writes, and when to stop are in
[Contribution health: probe, repair, rebuild](/troubleshooting#contribution-health-probe-repair-rebuild).
## Database Setup
### Initialization is required (boot via `createStoreWithSchema`)
Fulltext storage is **durably materialized once, at application boot**,
by `createStoreWithSchema`:
```typescript
// Run this once at startup — outside request handlers and transactions.
const [store] = await createStoreWithSchema(graph, backend);
```
`createStore(graph, backend)` is a synchronous, zero-I/O *attach*: it
does not create tables, repair DDL, or record that fulltext storage is
materialized. A fulltext read or write — a `searchable()` field write,
`store.search.fulltext()`, `store.search.hybrid()`,
`n.$fulltext.matches()`, `store.search.rebuildFulltext()`, or a
transaction that touches fulltext — against a database that was never
initialized throws `StoreNotInitializedError`. Use `createStore()` only
to attach to a database a prior `createStoreWithSchema` boot already
initialized. Graphs with no `searchable()` fields are unaffected.
### PostgreSQL
No extensions required. The built-in `tsvector` type and GIN indexes
work on every managed Postgres (RDS, Supabase, Neon, Cloud SQL, Aiven).
The fulltext table's DDL ships in `bootstrapTables()` and the migration
SQL. The first privileged `createStoreWithSchema` boot records one
deployment-scoped physical marker for that shared table plus a graph-local
activation marker. Later graphs reuse the physical attestation and need only
the DML activation write; fulltext operations require both markers (see above):
```typescript
import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Includes the fulltext table, tsvector column, GIN index, and pgvector
const migrationSQL = generatePostgresMigrationSQL();
```
### SQLite
No extensions required. FTS5 is compiled into the standard SQLite
distribution shipped with `better-sqlite3`, `libsql`, `bun:sqlite`, and
most other drivers. The FTS5 virtual table uses the
`porter unicode61 remove_diacritics 2` tokenizer.
## Querying
### `n.$fulltext.matches()` — The Query Predicate
`n.$fulltext.matches(query, k?, options?)` is a node-level fulltext
predicate. It's exposed on every `NodeAccessor`; at runtime it throws a
clear `UnsupportedPredicateError` if the node kind has no `searchable()`
fields, with a suggestion for how to fix the schema.
> **Visible in types, guarded at runtime.** `$fulltext` is present on every
> `NodeAccessor` at the TypeScript level for simplicity — a type-level
> brand would not survive modifiers like `.min(1).optional()`. The runtime
> check is the single source of truth: adding a `searchable()` field is
> what makes `.matches()` actually work. A query that type-checks can still
> throw `UnsupportedPredicateError` the first time it runs if no field
> was declared searchable.
```typescript
store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.$fulltext.matches("climate change"))
.select((ctx) => ctx.d)
.execute();
```
It compiles to a JOIN against the fulltext index, adds an ORDER BY on
relevance rank, and applies the top-k limit — all in a single SQL
statement that composes with every other query-builder feature.
**`k` vs `limit`**: `k` (the second positional arg) is the top-k cap
applied **inside the fulltext CTE** — how many candidates to pull from
the index before outer filtering and fusion. It defaults to `50`, which
is fine for single-predicate use. `.limit()` on the query controls the
**final result count**. When feeding into RRF (`store.search.hybrid` or
`.fuseWith()`), pass a larger `k` per predicate (e.g. 200) so there are
enough candidates for the fused ranking to be meaningful.
Traversal happens after candidate generation. A candidate can therefore fan
out into several match rows. Use query-level `.where((ctx) => ...)` to filter
those completed rows, then apply an explicit `.orderBy()` and `.limit()` for
the final result. The final order does not change which nodes entered the
top-k candidate set, and the final limit counts match rows rather than distinct
source nodes.
### Query Modes
```typescript
d.$fulltext.matches("climate -warming", 10, { mode: "websearch" })
// Google-style: quoted phrases, -excluded terms, OR operator
d.$fulltext.matches("climate change", 10, { mode: "phrase" })
// Exact phrase match
d.$fulltext.matches("climate change", 10, { mode: "plain" })
// All terms must appear (implicit AND), no special syntax
d.$fulltext.matches("climate & !warming", 10, { mode: "raw" })
// Dialect-native syntax (tsquery on Postgres, FTS5 on SQLite)
```
**When to use each:**
| Mode | Best For | Example |
|------|----------|---------|
| `websearch` (default) | User-facing search boxes | `"climate change" -hoax OR warming` |
| `phrase` | Proper nouns, exact quotes | `"New York Times"` |
| `plain` | Programmatic queries | `climate change` |
| `raw` | Advanced users who know the dialect syntax | `climate<->change` |
### Composing with Filters
Fulltext is just another predicate — combine with `.and()`:
```typescript
store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext
.matches("machine learning", 20, { mode: "websearch" })
.and(d.published.eq(true))
.and(d.publishedAt.gte("2024-01-01"))
.and(d.tenantId.eq(tenant))
)
.select((ctx) => ctx.d)
.execute();
```
### Composing with Graph Traversal
`$fulltext.matches()` works inside any traversal:
```typescript
// Find documents matching "climate" that were written by someone I follow
const results = await store
.query()
.from("Person", "me")
.whereNode("me", (p) => p.id.eq(currentUserId))
.traverse("follows", "f")
.to("Person", "author")
.traverse("authored", "a", { direction: "in" })
.to("Document", "d")
.whereNode("d", (d) => d.$fulltext.matches("climate", 10))
.select((ctx) => ({
title: ctx.d.title,
author: ctx.author.name,
}))
.execute();
```
### Hybrid Search (Query Builder)
Use `$fulltext.matches()` and `.similarTo()` in the same query and
TypeGraph automatically fuses the two signals with Reciprocal Rank Fusion
at the SQL layer:
```typescript
const results = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext
.matches("renewable energy", 50)
.and(d.embedding.similarTo(queryVector, 50))
.and(d.tenantId.eq(tenant))
)
.select((ctx) => ctx.d)
.limit(10)
.execute();
```
The compiled SQL builds two CTEs (one for the vector side, one for the
fulltext side), orders each by relevance, and the outer query sorts by
`1/(60 + rank_vector) + 1/(60 + rank_fulltext)`. One round-trip, fully
composable with any other predicate.
### Tuning RRF
Defaults (k=60, equal weights) suit most workloads. Bias toward fulltext
for exact-match queries, toward vectors for conceptual queries:
```typescript
store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext
.matches("renewable energy", 50)
.and(d.embedding.similarTo(queryVector, 50))
.and(d.tenantId.eq(tenant))
)
.fuseWith({ k: 60, weights: { vector: 1.0, fulltext: 1.5 } })
.limit(10)
.execute();
```
`.fuseWith()` throws at compile time if the query lacks either a
`.similarTo()` or a `$fulltext.matches()`. Validation rejects non-finite
or negative `k`/weights. The same shape is accepted by
`store.search.hybrid({ fusion })` and validated by the same function on
both paths.
### Hybrid Search (Store API)
For tunable RRF parameters, use `store.search.hybrid()`. On the built-in
backends this runs as a **single SQL statement** — both sources, RRF
fusion, and node hydration composed together — so a hybrid query costs one
round trip instead of three. The saving scales with per-statement cost:
decisive on serverless HTTP drivers, Cloudflare D1 / Durable Objects, and
remote databases; on a local low-latency connection the two paths are
within a few milliseconds of each other. (Kind expansions via
`includeSubClasses`, and custom backends without the composed statement,
transparently use a multi-statement path with identical results.)
```typescript
const results = await store.search.hybrid("Document", {
limit: 10,
vector: {
fieldPath: "embedding",
queryEmbedding: await embed(question),
metric: "cosine",
k: 50, // Candidates to retrieve from vector
},
fulltext: {
query: question,
k: 50, // Candidates to retrieve from fulltext
includeSnippets: true,
},
fusion: {
method: "rrf",
k: 60, // RRF constant (classic default)
weights: {
vector: 1.0,
fulltext: 1.5, // Weight fulltext higher for exact-match workloads
},
},
});
```
Each hit carries sub-scores from both halves so you can debug ranking:
```typescript
for (const hit of results) {
console.log(hit.node.title, hit.score);
console.log(" vector rank:", hit.vector?.rank);
console.log(" fulltext rank:", hit.fulltext?.rank);
console.log(" snippet:", hit.fulltext?.snippet);
}
```
### Fulltext-Only Store API
For quick fulltext lookups that don't need the query builder:
```typescript
const hits = await store.search.fulltext("Document", {
query: "quarterly earnings",
limit: 10,
mode: "websearch",
includeSnippets: true,
});
for (const hit of hits) {
console.log(hit.node.title, hit.score, hit.snippet);
}
```
#### Options reference
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `query` | `string` | — *(required)* | User-supplied query string. Parsed according to `mode`. |
| `limit` | `number` | — *(required)* | Max rows to return. Must be a positive integer. |
| `mode` | `"websearch" \| "phrase" \| "plain" \| "raw"` | `"websearch"` | Parser for `query`. See [Query Modes](#query-modes). |
| `language` | `string` | the kind's declared language | Override the stemming/tokenization language for this query. By default the query is parsed with the language the kind's `searchable()` fields declare — a plan-time constant, which is what lets PostgreSQL serve the match from the `tsv` GIN index (a per-row language reference makes the tsquery non-constant and forces a scan). Postgres only — SQLite FTS5's tokenizer is fixed at table-create time and a per-query override throws. |
| `minScore` | `number` | *(none)* | Drop hits whose backend-native score is below this threshold. Score units depend on the strategy. |
| `includeSnippets` | `boolean` | `false` | Return a highlighted `…` snippet per hit. Noticeably slower than plain search — request only for final-page results. |
| `where` | `(accessor) => Predicate` | *(none)* | Property predicate compiled into the search statement's candidate set — the engine ranks only matching rows, so a filter never shrinks results below `limit` when enough matches exist (libSQL DiskANN: bounded by its 4× over-fetch). Same accessor and semantics as `store.nodes..find({ where })`. |
| `offset` | `number` | `0` | Rank-relative pagination: skip the first `offset` ranked hits. |
| `includeSubClasses` | `boolean` | `false` | Also search `subClassOf` descendant kinds and merge their scores into one ranking. |
The same three options are available on `store.search.vector` and
`store.search.hybrid` (where `where` and `includeSubClasses` apply to both
halves). Search always follows current-read semantics: tombstoned nodes and
nodes outside their validity window never rank.
Returned hits are `FulltextSearchHit>` with `node`, `score`
(higher = more relevant), `rank` (1-based), and `snippet` (when
requested).
## Reciprocal Rank Fusion
RRF is the de facto standard for combining ranked lists from multiple
retrievers. The formula:
```text
score(doc) = Σ weight_source / (k + rank_source)
```
Where `k` is the RRF constant (classic default: 60), `rank_source` is
the document's 1-based ordinal rank in each source, and `weight_source`
lets you bias toward one retriever.
**Why it works:** RRF is rank-based, not score-based. It doesn't care
that vector distances are in `[0, 2]` while BM25 scores are
unbounded — it only cares about ordinal position. That makes it robust
to heterogeneous score distributions across retrievers.
**Tuning tips:**
- Over-fetch from each side (default: 4× the requested limit). More
candidates per source = better recall.
- Bump `weights.fulltext` higher when exact matches matter (names, IDs,
proper nouns). Bump `weights.vector` for conceptual queries.
- Leave `k = 60` alone unless benchmarks show otherwise.
## Best Practices
### All Searchable Fields Share One Index
TypeGraph indexes all `searchable()` fields on a node as one document
(see [How Indexing Works](#how-indexing-works)). There's a single
`n.$fulltext` accessor per node — `searchable()` declarations on
individual fields are what bring it into existence and what determine
which text gets indexed.
### Use `includeSnippets` Sparingly
Highlighting (`ts_headline` on Postgres, `snippet()` on SQLite) is
noticeably slower than plain search. Request it only for final-page
results, not for large over-fetch pools.
### Pair with a Reranker for Top Quality
RRF is a strong baseline, but production RAG systems typically add a
cross-encoder reranker (Cohere Rerank, `bge-reranker`, etc.) as a final
stage. TypeGraph gives you the candidate set — the reranker picks the
winning order:
```typescript
const candidates = await store.search.hybrid("Document", {
limit: 50, // Over-fetch for reranker
vector: { fieldPath: "embedding", queryEmbedding },
fulltext: { query },
});
const reranked = await cohere.rerank({
query,
documents: candidates.map((c) => c.node.content),
top_n: 10,
});
```
### Filter Before You Fuse
Applying predicates via `.and()` shrinks the candidate pool before the
fusion ORDER BY, which improves both latency and ranking quality — there
are fewer irrelevant candidates competing for top positions:
```typescript
// Fast: tenant filter applied inside each CTE
.whereNode("d", (d) =>
d.$fulltext.matches(query, 50)
.and(d.embedding.similarTo(queryVec, 50))
.and(d.tenantId.eq(tenant))
)
// Slow: tenant filter applied AFTER fusion
.whereNode("d", (d) =>
d.$fulltext.matches(query, 5000)
.and(d.embedding.similarTo(queryVec, 5000))
)
// ...then filter results in JS
```
## Limitations
### One Fulltext Predicate Per Query
A single query can contain at most one `$fulltext.matches()` predicate.
This mirrors the constraint on `.similarTo()` and keeps the RRF fusion
model well-defined. If you need to search multiple terms, combine them
into one query string using websearch mode:
```typescript
// Good
d.$fulltext.matches("climate change OR global warming", 20)
// Rejected (at query-build time, not by the type checker)
d.$fulltext.matches("climate", 10).and(d.$fulltext.matches("warming", 10))
```
This invariant is enforced when the query is compiled
(`UnsupportedPredicateError`), not by TypeScript — so a surprising
second `.matches()` call surfaces as a runtime error the first time
the query runs.
### No `.matches()` Under OR or NOT
Fulltext predicates must appear at top level or inside AND groups. They
rewrite query structure (adding a CTE and ORDER BY) in a way that isn't
compatible with disjunction or negation semantics.
### Tokenizer Is Fixed on SQLite
FTS5 tokenizer options are set at CREATE VIRTUAL TABLE time. TypeGraph
ships with `porter unicode61 remove_diacritics 2` — a solid default for
English and accented Latin-script languages. For CJK or other tokenizers,
create the fulltext table manually with your preferred options.
### No Per-Field Weighting
All searchable fields on a node contribute equally to the combined
document. Postgres `setweight()`-style per-field bias is a planned
extension; today, structure your fields to put the most important text
first or split high-weight content into a dedicated kind. This
limitation applies even when you [swap in a custom
`FulltextStrategy`](#custom-fulltext-strategies) — TypeGraph
concatenates searchable fields into one `content` string before handing
it to the strategy.
## Custom Fulltext Strategies
`createPostgresBackend(db, { fulltext })` and
`createSqliteBackend(db, { fulltext })` accept a `FulltextStrategy`
that owns the **entire** fulltext pipeline — DDL, INSERT/UPSERT
(single + batch), DELETE (single + batch), MATCH condition, rank
expression, and snippet generation. The same strategy flows through
`store.search.fulltext()`, `store.search.hybrid()`,
`$fulltext.matches()` in the query builder,
`store.search.rebuildFulltext()`, and `bootstrapTables()` DDL.
Use this when the built-in `tsvector` isn't the right fit — for
example, BM25 inside Postgres (ParadeDB / `pg_search`), trigram
similarity (`pg_trgm`), or fulltext optimized for CJK languages
(`pgroonga`).
Most SQLite users should leave the default `fts5Strategy` in place.
### Constraints on alternate strategies
Before implementing a strategy, know what the abstraction does **not**
let you change today:
- **Side table is mandatory.** Every strategy writes one row per
`(graph_id, node_kind, node_id)` to a dedicated fulltext table. A
strategy cannot skip the side table and index a column on the main
nodes table directly (e.g. a GIN trigram index on
`typegraph_nodes.props`). Strategies *can* choose the column layout,
index type, and any computed projection inside that side table.
- **Content is pre-concatenated.** TypeGraph joins every
`searchable()` field value with `\n` before the strategy sees it —
`UpsertFulltextParams.content` is a single string. Per-field
indexing (`setweight`, per-column BM25 boosts, pgroonga
per-column weights) is not plumbed through today; a richer
per-field payload is planned but not yet part of the public
strategy contract.
- **One language per row.** When a node has multiple `searchable()`
fields with different `language` values, the first field's
language wins and is recorded on the row. TypeGraph emits a
one-time warning per conflicting schema; true multilingual
indexing needs a dedicated node kind per language.
### Strategy skeleton
The `FulltextStrategy` contract and every type referenced by it are exported
from the Drizzle-free backend-authoring entrypoint.
Fields below are the minimum surface; see
`src/query/dialect/fulltext-strategy.ts` in the TypeGraph source for
`tsvectorStrategy` and `fts5Strategy` as full references.
```typescript
import {
sql,
type FulltextStrategy,
type SqlFragment,
} from "@nicia-ai/typegraph/backend";
/**
* Example: a trigram-based strategy on top of pg_trgm. Illustrative —
* not production code. pg_trgm supports plain-term matching only, so
* `supportedModes` advertises `"plain"` and rejects everything else at
* compile time.
*/
export const pgTrgmStrategy: FulltextStrategy = {
name: "pg_trgm",
supportedModes: ["plain"],
supportsSnippets: false, // no native highlight; emit NULL snippet
supportsPrefix: false, // trigram similarity, not prefix
supportsLanguageOverride: false,
languages: ["simple"],
matchCondition(table, query) {
return sql`${sql.identifier(table)}."content" % ${query}`;
},
rankExpression(table, query) {
return sql`similarity(${sql.identifier(table)}."content", ${query})`;
},
snippetExpression() {
// `supportsSnippets: false` — callers get NULL and skip the field.
return sql`NULL`;
},
// Declares the table(s) this strategy owns as authoritative
// TableContributions. pg_trgm brings its own table (not the typed
// Drizzle `tables.fulltext`), so it is emitted verbatim from
// `createDdl` and is invisible to drizzle-kit unless you export your
// own table object. `runtimeEnsure: true` because no
// drizzle-kit-managed setup can create it.
ownedTables(primaryTableName) {
return [
{
logicalName: "fulltext",
owner: "pg_trgm",
tableName: primaryTableName,
createDdl: [
`CREATE EXTENSION IF NOT EXISTS pg_trgm;`,
`CREATE TABLE IF NOT EXISTS "${primaryTableName}" (
"graph_id" TEXT NOT NULL,
"node_kind" TEXT NOT NULL,
"node_id" TEXT NOT NULL,
"content" TEXT NOT NULL,
"language" TEXT NOT NULL,
"updated_at" TIMESTAMPTZ NOT NULL,
PRIMARY KEY ("graph_id", "node_kind", "node_id")
);`,
`CREATE INDEX IF NOT EXISTS "${primaryTableName}_trgm_idx"
ON "${primaryTableName}" USING GIN ("content" gin_trgm_ops);`,
],
runtimeEnsure: true,
},
];
},
buildUpsert(table, params, timestamp) {
return [
sql`
INSERT INTO ${sql.identifier(table)}
("graph_id", "node_kind", "node_id", "content", "language", "updated_at")
VALUES (${params.graphId}, ${params.nodeKind}, ${params.nodeId},
${params.content}, ${params.language}, ${timestamp})
ON CONFLICT ("graph_id", "node_kind", "node_id")
DO UPDATE SET
"content" = EXCLUDED."content",
"language" = EXCLUDED."language",
"updated_at" = EXCLUDED."updated_at"
`,
];
},
buildBatchUpsert(table, params, timestamp) {
if (params.rows.length === 0) return [];
// Dedup last-write-wins by nodeId, then emit a single multi-VALUES INSERT.
// Postgres ON CONFLICT rejects repeated conflict keys inside one statement.
// (The shipped helpers in fulltext-strategy.ts show this pattern.)
return [/* … */];
},
buildDelete(table, params) {
return [
sql`
DELETE FROM ${sql.identifier(table)}
WHERE "graph_id" = ${params.graphId}
AND "node_kind" = ${params.nodeKind}
AND "node_id" = ${params.nodeId}
`,
];
},
buildBatchDelete(table, params) {
if (params.nodeIds.length === 0) return [];
return [/* DELETE … WHERE node_id IN (…) */];
},
};
```
Wire it in at backend construction:
```typescript
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const backend = createPostgresBackend(db, { fulltext: pgTrgmStrategy });
```
Capabilities (`phraseQueries`, `prefixQueries`, `highlighting`,
`languages`) are derived automatically from the strategy, so
`store.search.fulltext({ mode: "websearch" })` now throws
`ConfigurationError` before any SQL is generated — the strategy's
`supportedModes` is the source of truth.
## Troubleshooting
### `StoreNotInitializedError: fulltext storage … is not initialized`
The database was never booted through `createStoreWithSchema`, so the
deployment-scoped physical marker or this graph's activation marker is missing
(or it is `stale` — the strategy/DDL changed since it was recorded, or
`failed` — the last boot-time attempt errored). Bare `createStore()`
deliberately does **not** self-heal this on the hot path.
Fix: call `createStoreWithSchema(graph, backend)` once at application
startup — outside request handlers and adopted transactions — before any
fulltext operation. If you previously relied on fulltext tables being
created lazily on first write via `createStore()`, that path was removed;
move the initialization to an explicit boot step. A `stale` reason means
the recorded shape no longer matches the active strategy/DDL: migrate or
drop the fulltext table and re-run the boot, or restore the original
strategy.
`ContributionUnavailableError` with `state: "physical-storage-missing"` is
different: the physical fulltext table disappeared after initialization. Gated
fulltext searches and searchable node writes raise this typed error; compiled
query-builder predicates can still surface the engine's missing-relation error.
Run `store.rebuildContribution("fulltext")`; rerunning ordinary initialization
cannot reconstruct the missing indexed content. Use `probeContributions()` at
startup when an application must detect this out-of-band catalog damage before
the first dependent read or write.
### `Cannot call .$fulltext.matches() on alias "x"`
`$fulltext` is exposed on every node accessor at the type level, but
calling `.matches()` requires the node kind to have at least one
`searchable()` field — otherwise there's no indexed content to search.
The runtime guard throws a clear error pointing at the alias:
```text
Cannot call .$fulltext.matches() on alias "d" — its node kind has no
fields declared with searchable(). Add at least one:
`title: searchable({ language: "english" })`.
```
Fix by adding a searchable field to the schema:
```typescript
// Before:
title: z.string(),
// After:
title: searchable({ language: "english" }),
```
Refinements like `.min(1)` and `.trim()` are preserved — you can write
`searchable({ language: "english" }).min(1)` and the field is still
indexed.
### Empty fulltext results after bulk insert
TypeGraph syncs the fulltext index inline with each node write. If you
bulk-inserted via raw SQL that bypassed the store layer, the fulltext
table won't have entries. Re-run the inserts through
`store.nodes.X.create()` / `.bulkCreate()`, run
`store.search.rebuildFulltext()` to populate the index from existing rows,
or issue `backend.upsertFulltext()` / `backend.upsertFulltextBatch()`
calls directly.
### After adding `searchable()` to existing data
See [Adding `searchable()` to an Existing Graph](#adding-searchable-to-an-existing-graph)
above for the rebuild recipe and the caveats that apply to concurrent
workloads.
### `"Fulltext match predicates cannot be nested under OR or NOT"`
See [No `.matches()` Under OR or NOT](#no-matches-under-or-or-not)
above. Move the `$fulltext.matches()` to the top level or inside an
`.and()`.
### Hybrid results miss obvious matches
Increase the per-source `k` (over-fetch). The default is 4× the final
`limit`, which is tuned for small result pages. Large corpora benefit
from `k: 200` or higher on each side.
### Postgres: `text search configuration "xyz" does not exist`
The `language` you passed to `searchable({ language })` must be an
installed `regconfig` on your Postgres server. Every stock install ships
`simple`, `english`, `french`, `german`, `italian`, `portuguese`,
`russian`, `spanish`, and `swedish`; anything else requires an extension
(`zhparser` for Chinese, `pg_trgm` for trigram-based matching, or a
custom dictionary).
TypeGraph emits a `console.warn` at query time when you pass a language
outside the backend-advertised list, but a typo or missing extension
only fails when Postgres tries to build the `tsvector`. To diagnose:
```sql
SELECT cfgname FROM pg_ts_config;
```
Pick a name from that list, or install the extension that provides the
one you want. If you're running a managed Postgres (RDS, Supabase, Neon,
Cloud SQL, Aiven), check the provider's docs for which language extensions
are enabled — some require a restart or explicit enabling.
### Swapping to a custom fulltext strategy
See [Custom Fulltext Strategies](#custom-fulltext-strategies) for the
full interface, constraints, and a skeleton implementation.
## API Reference
- **Schema**: [`searchable()`](/queries/predicates#searchable)
- **Predicate**: [`n.$fulltext.matches()`](/queries/predicates#searchable)
- **Tunable fusion**: `QueryBuilder.fuseWith({ k, weights })`
- **Rebuild**: `store.search.rebuildFulltext(nodeKind?, { pageSize? })`
- **Store API**: `store.search.fulltext()` and `store.search.hybrid()` —
see the [Schemas & Stores reference](/schemas-stores).
See also:
- [Semantic Search](/semantic-search) — vector embeddings and
`.similarTo()`
- [Predicates reference](/queries/predicates) — complete predicate
catalog
- [Knowledge Graph for RAG](/examples/knowledge-graph-rag) — end-to-end
example combining fulltext, vector, and graph traversal
# Graph Extensions
> Extend a TypeGraph schema at runtime — durable, multi-process safe, with full Zod validation and unique-constraint enforcement.
Graph extensions let your application declare new node and edge
kinds **at runtime** — durable across restarts, with semantic parity to
compile-time `defineNode` / `defineEdge`. The motivating use case:
**agent-driven schema induction**, where an LLM proposes a typed schema
from a corpus, an operator approves it, and the live graph immediately
ingests under the new schema with no code change or restart.
:::note[See it end-to-end]
For a runnable scenario with an operator-approved agent in TypeScript,
see [Agent-Driven Schema](/examples/agent-driven-schema). For the same
loop driven by an open-weight LLM against public-record clinical data —
with a repair loop and smoke-test pattern — see
[`pdlug/typegraph-clinical-demo`](https://github.com/pdlug/typegraph-clinical-demo).
:::
This guide covers the core verbs:
| Verb | Purpose |
| ------------------------------------------------ | ----------------------------------------------------------------- |
| `defineGraphExtension` | Build a typed extension (pure value, no I/O) |
| `store.evolve(extension)` | Atomically commit a new schema version with the extension applied |
| `store.introspect()` | Snapshot the merged schema, persisted extension, version, and hash |
| `store.materializeIndexes()` | Run declared `CREATE INDEX` DDL against the live database |
| `store.deprecateKinds(...)` / `undeprecateKinds` | Soft-deprecate kinds for codegen / lint signaling |
| `store.removeKinds(...)` | Remove graph-extension-declared kinds from the active schema |
| `store.materializeRemovals()` | Delete rows queued by graph-extension-kind removal |
For the schema-management primitives that graph extensions ride on top
of, see [Schema Migrations](/schema-management) and [Evolving
Schemas](/schema-evolution).
## When to use graph extensions
Use them when **the kind set is not known at code time**:
- Agent / LLM proposes a new typed schema from observed data.
- Multi-tenant deployments where each tenant defines their own kinds.
- ETL pipelines that ingest sources with shifting structure.
- Plugins / extensions that contribute kinds at install time.
For everything else — kinds you can declare in TypeScript at deploy time
— use the compile-time DSL (`defineNode`, `defineEdge`, `defineGraph`).
The compile-time path is type-safe end-to-end; graph extensions trade
some of that type-safety for the ability to evolve without redeploying.
## A complete example
```ts
import { z } from "zod";
import {
createStoreWithSchema,
defineGraph,
defineNode,
defineGraphExtension,
} from "@nicia-ai/typegraph";
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
// 1. Boot with a compile-time kind.
const Document = defineNode("Document", {
schema: z.object({ title: z.string(), body: z.string() }),
});
const baseGraph = defineGraph({
id: "research_corpus",
nodes: { Document: { type: Document } },
edges: {},
});
const { backend } = createLocalSqliteBackend();
const [store] = await createStoreWithSchema(baseGraph, backend);
// 2. An agent proposes a new kind at runtime.
const proposal = defineGraphExtension({
nodes: {
Paper: {
description: "An academic paper inferred from the corpus",
properties: {
title: { type: "string", minLength: 1 },
doi: { type: "string", minLength: 1 },
year: { type: "number", int: true, min: 1900, max: 2100 },
},
unique: [{ name: "paper_doi_unique", fields: ["doi"] }],
},
},
indexes: [
{
entity: "node",
kind: "Paper",
name: "paper_by_doi",
fields: ["doi"],
unique: true,
},
],
});
// 3. Operator approves; commit atomically.
const evolved = await store.evolve(proposal);
// 4. Use the dynamic-collection accessor (the type system does not
// widen for extension kinds — see "Reaching extension kinds" below).
const papers = evolved.getNodeCollection("Paper")!;
await papers.create({
title: "Attention is all you need",
doi: "10.5555/3295222.3295349",
year: 2017,
});
```
A complete runnable version is in [`examples/16-graph-extensions.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/16-graph-extensions.ts).
## Instantiate a graph template
For many tenant graphs with the same final shape, register a verified schema
once and instantiate each target as schema version 1. The template document
stays in the database; instantiation sends only its id, the target graph id,
and a client-computed content hash.
```ts
const [source] = await createAdapterStoreWithSchema(baseGraph, adminBackend);
const template = await registerGraphTemplate(adminBackend, {
templateId: "research-corpus-v1",
reconciled: source.reconciledSchema,
});
const target = await instantiateGraph(adminBackend, {
template,
graphId: "tenant-123",
});
// The returned target snapshot can create a request-scoped Store without
// another schema read. `tenantGraph` is the same compile-time shape with id
// "tenant-123".
const store = createAdapterStore(tenantGraph, runtimeBackend, {
reconciled: target.reconciled,
});
```
Instantiation is idempotent for the same template and target. A target already
initialized with a different schema is refused. On PostgreSQL the clone
statement takes the same graph advisory-lock key as schema commits; SQLite's
schema and marker writes are serialized by its writer lock.
Templates clone the source graph's graph-local runtime-contribution activation
markers along with its schema row. Deployment-scoped physical attestations,
such as the shared fulltext table marker, already belong to the database and
are not duplicated per template target. A target can therefore be reopened with
`createVerifiedStore` or `createVerifiedAdapterStore` from a later serverless
isolate without another schema reconciliation or provisioning DDL step.
Registration and instantiation are DML-only. They assert the deployment-wide
base-schema marker and throw `BaseSchemaMigrationError` rather than attempting
DDL when adoption is missing, stale, or newer than the running library. Run the
normal privileged `createStoreWithSchema` boot or the published external
migration before handing runtime requests to a DML-only role.
Template eligibility follows the backend's capabilities. Embedding declarations
are allowed when the backend is configured with `vector: false`, because that
backend has no graph-scoped vector storage to clone. A vector-enabled backend
continues to refuse embedding-bearing schema-only templates.
## The graph extension
`defineGraphExtension` accepts a structured value describing the new
kinds. The extension is JSON-serializable — that's load-bearing for
durability (see [Restart parity](#restart-parity-the-load-bearing-invariant)).
### Document format versioning
Every document carries a `version` field (currently `1`). The validator
stamps the version automatically when you call
`defineGraphExtension`, so consumer code never has to set it
explicitly. Stored documents from before this field existed are
treated as `version: 1` (the legacy default).
The forward-compat policy:
- **Additive minor changes** (new optional property modifier, new
`format` value, new top-level slice within the same major) ride
forward without bumping `version`. The validator does not reject
unknown top-level keys, and the persistence-side zod is `.loose()`
on every nested object — an older runtime reading a newer extension
silently ignores unknown fields and continues working.
- **Breaking changes** bump `version` to a higher major. An older
runtime reading a higher-version extension fails with
`GRAPH_EXTENSION_VERSION_UNSUPPORTED` and an actionable error
pointing the operator at upgrading the library — there is no
automatic downgrade path. The current major is exported as
`CURRENT_GRAPH_EXTENSION_VERSION` for tooling that wants to
pre-flight check.
- **Legacy extensions** (committed before `version` existed) and
extensions that explicitly omit `version` are interpreted as
`LEGACY_GRAPH_EXTENSION_VERSION`, pinned permanently to `1`. This
is deliberately distinct from `CURRENT`: when a future v2 ships,
legacy v1 extensions continue parsing as v1 (so the version-mismatch
path can route them through migration) rather than being silently
re-classified as v2 by a default-equals-current rule.
```ts
import {
CURRENT_GRAPH_EXTENSION_VERSION,
LEGACY_GRAPH_EXTENSION_VERSION,
} from "@nicia-ai/typegraph";
console.log(CURRENT_GRAPH_EXTENSION_VERSION); // 1 (today; bumps with breaking changes)
console.log(LEGACY_GRAPH_EXTENSION_VERSION); // 1 (always; the pre-versioning default)
```
### Property types (the v1 subset)
The following types are supported. The set is deliberately small so that
LLM-induced schemas can be audited at a glance and so the persistence
layer never has to reconstruct opaque Zod refinements from JSON.
| Type | JSON shape |
| --------- | ------------------------------------------------------------------------- |
| `string` | `{ type: "string", minLength?, maxLength?, pattern?, format? }` |
| `number` | `{ type: "number", int?, min?, max? }` |
| `boolean` | `{ type: "boolean" }` |
| `enum` | `{ type: "enum", values: ["a", "b", ...] }` |
| `array` | `{ type: "array", items: }` |
| `object` | `{ type: "object", properties: { foo: , ... } }` (one nesting only) |
Supported string formats: `"datetime"`, `"uri"`, `"email"`, `"uuid"`,
`"date"`. These route to the corresponding Zod factories
(`z.iso.datetime()`, `z.url()`, `z.email()`, `z.uuid()`, `z.iso.date()`).
Modifiers available on every property:
- `optional?: true` — omits from the required set.
- `description?: string` — surfaces in tooling.
- `searchable?: SearchableModifier` — string only; routes through the
`searchable()` brand for fulltext indexing.
- `embedding?: { dimensions: number }` — array-of-number only; routes
through the `embedding()` brand for vector search.
### Unique constraints
Pass `unique: [{ name, fields, ... }]` per kind. The `name` is required
(used as the diffing identity key) and must be unique within the kind.
Supports `scope`, `collation`, and a restricted `where` clause limited
to `isNull` / `isNotNull` (the only operations round-trippable through
the persisted form).
### Relational indexes
Pass `indexes: [...]` at the document top level to declare relational
indexes for graph-extension or compile-time host kinds:
```ts
const proposal = defineGraphExtension({
nodes: {
Paper: {
properties: {
doi: { type: "string" },
title: { type: "string", searchable: { language: "english" } },
},
},
},
indexes: [
{
entity: "node",
kind: "Paper",
name: "paper_by_doi",
fields: ["doi"],
unique: true,
},
],
});
```
Index `name`s are unique across the merged graph. A graph-extension index that
reuses a compile-time index name, or a later extension that reuses an
earlier graph-extension index name for a different declaration, is rejected.
### Edges
Edges follow the same shape as nodes and declare endpoints using kind names.
An array-valued `to` permits every combination of `from` and `to` kinds.
A map-valued `to` restricts targets by source kind:
```ts
const proposal = defineGraphExtension({
edges: {
assignedTo: {
from: ["Employee", "Student"],
to: {
Employee: ["Department"],
Student: ["Course"],
},
properties: {},
},
},
});
const evolved = await store.evolve(proposal);
```
All endpoint kinds must resolve in the merged graph when `evolve()` runs.
Map keys must exactly cover `from`, and each target array must be nonempty.
The map is persisted and restored on restart, preserving the same runtime pair
validation as [compile-time declarations](/core-concepts#source-dependent-targets).
Adding allowed pairs broadens an extension edge. Removing a pair tightens it,
even if the overall source and target kind sets stay the same. Tightening
currently requires the **entire edge kind** to be empty, not just the removed pair.
### Ontology
Pass `ontology: [{ metaEdge, from, to }, ...]` to declare ontology
relations between kinds (subClassOf, partOf, etc.). The meta-edge name
must match a meta-edge known to the merged graph.
## `store.evolve(extension, options?)`
```ts
const evolved = await store.evolve(extension);
const evolved = await store.evolve(extension, { ref });
const evolved = await store.evolve(extension, { eager: {} });
```
`evolve` is the consumer-facing primitive that drives extension. It:
1. **Catches up to persisted state** — folds any persisted
extension and deprecation set into the local baseline so a
stale store doesn't trample another writer's progress.
2. **Merges** the new extension into the baseline graph. Re-declaring
an existing extension kind with the same shape is a no-op; with a
non-additive change against existing rows it throws
`IncompatibleChangeError` (code `INCOMPATIBLE_CHANGE`).
3. **Atomically commits** a new schema version via `commitSchemaVersion`
(CAS on the active version).
4. **Returns the resulting `Store`** carrying the extended graph. The type
parameter `G` does NOT widen — see [Reaching extension
kinds](#reaching-extension-kinds-from-the-type-system) below.
:::caution[Use the returned Store]
`Store` instances are immutable schema snapshots. After `evolve()` resolves,
use its returned Store for **every subsequent Store operation in the same
request**, not just operations involving newly added kinds. A Store captured
before the commit remains pinned to the previous schema version, so its next
managed write fails the schema-version fence.
:::
### Plan outside and apply inside a caller-owned transaction
`planEvolution(extension)` prepares a named schema change before the caller
opens its write transaction. It returns an immutable `"noop"` or `"change"`
plan with `graphId`, `baseline: { version, hash }`, and
`result: { version, hash }`. Change plans expose an ordered `requirements`
array whose entries name new-kind additions, empty-kind checks, vector slots,
and identity work. A `new-kind` entry describes a graph delta; it does not
indicate that a removal is queued or direct callers to run
`materializeRemovals()`. The plan is opaque and bound to the loaded TypeGraph
module: it cannot be serialized, cloned, or reconstructed. It can be passed between
compatible Stores for the same graph that use the same loaded module; apply
still checks the active graph and fenced baseline version/hash. The
default `{ source: "database" }` reloads the active schema. `{ source:
"cached" }` uses a previously loaded planning snapshot on the same Store; a
cached plan is not fresh database evidence. A stale baseline is refused during
apply, so retry by rolling back the whole caller transaction and replanning
outside it.
Use `withEvolvedTransaction(nativeTx, plan, callback, { waitBudgetMs })` when
the schema change, TypeGraph writes, and application SQL must share the
caller's commit. Enter this boundary before other TypeGraph callbacks on the
same native transaction. The callback receives an evolved transaction context,
including its reads, collections, and supported composition operations. It
does not receive a replacement root Store.
```ts
const cachedStore = store;
const ref = { current: cachedStore };
const plan = await cachedStore.planEvolution(proposal);
const writerStore = cachedStore.withBackend(writerBackend);
const provisional = await db.transaction(async (nativeTx) => {
const outcome = await writerStore.withEvolvedTransaction(
nativeTx,
plan,
async (tx) => {
const person = await tx.nodes.Person.create({ name: "Ada" });
return person.id;
},
plan.status === "change" ? { waitBudgetMs: 5000 } : undefined,
);
await nativeTx.insert(applicationEvents).values({
personId: outcome.result,
schemaVersion: outcome.receipt.schema.version,
});
return outcome;
});
const refreshed = await cachedStore.refreshSchema({
minVersion: provisional.receipt.schema.version,
ref,
});
```
The callback and receipt finish before the outer SQL transaction commits.
The receipt's schema version/hash and recorded anchor are provisional until
that commit succeeds. A callback failure must reject the outer transaction;
catching it and committing does not prove rollback. Callback contexts and
queries built from them expire when the callback returns, including after a
failure. Use `refreshSchema()` only after awaiting a successful outer commit.
When the cached Store already matches `minVersion`, refresh returns it
without a read; that shortcut does not check for a newer database version.
Otherwise refresh reads the active schema, accepts a newer committed version,
and refuses a missing or older one. It applies no extension or storage
provisioning.
The adopted apply path supports metadata-only changes, required-empty checks,
and transactional identity and vector provisioning. Configure the adapter with
`schemaProvisioning: "transactional"` on a privileged connection when the plan
names new vector slots or identity work. The adapter's default DML-only policy
refuses those requirements before the schema fence, DDL, callback, or mutation.
The privileged path rechecks storage on the caller's fenced session, provisions
the required relations and vector contribution markers there, and rolls them
back with the outer transaction. Missing bootstrap tables still refuse; run
the normal bootstrap before serving adopted evolution requests.
The apply path uses a bounded schema fence, baseline validation, and a
version CAS; it can issue multiple statements. `waitBudgetMs` bounds fence
acquisition for change plans; no-op plans refuse an explicit `waitBudgetMs`.
On timeout, roll back and retry the complete native transaction.
Do not pass `ref` or eager-index options to `withEvolvedTransaction`; they
cannot be honored before the outer commit and are refused.
Generic and concurrent eager indexes remain explicit maintenance after commit:
call `materializeIndexes()` on the refreshed Store when they are needed.
When a wiring pass produces a no-op, ordinary recorded adoption is sufficient
and never takes the exclusive evolution fence. Reconcile only if the Store is
behind the named snapshot; matching versions make this refresh a cached read:
```typescript
if (plan.status === "noop") {
const current = await store.refreshSchema({ minVersion: plan.baseline.version });
await db.transaction((nativeTx) =>
current.withRecordedTransaction(nativeTx, async (tx) => {
await tx.nodes.Person.create({ name: "Ada" });
}),
);
}
```
### The `ref` pattern
`Store` is immutable by construction — `evolve()` returns the Store for the
resulting schema, using a fresh instance when the schema advances. Long-lived
consumer code that holds the Store in a singleton needs a way to re-bind it
before the operation completes. Pass `options.ref: { current: store }` (a
`StoreRef>`):
```ts
const ref: StoreRef> = { current: store };
const evolved = await ref.current.evolve(extension, { ref });
// `ref.current === evolved`; use either reference from here on.
const papers = ref.current.getNodeCollectionOrThrow("Paper");
await papers.create({ title: "...", doi: "...", year: 2024 });
```
Capture `ref.current` once at request entry only when that request will not
change the schema. If it does call `evolve()`, switch to the returned Store or
dereference `ref.current` again after the call. `StoreRef` is structurally
just `{ current: T }`; the library doesn't provide a factory because the
consumer composes the handle themselves (it could be a Vue ref, MobX
observable, Zustand atom, etc.).
The ref covers schema changes made by calls that receive it. If another process
or isolate can advance the schema, probe the committed version before reusing a
cached Store and perform a verified open when it changes. See
[Per-request connections: cache the verified Store](/integration#per-request-connections-cache-the-verified-store)
for the complete `getCommittedSchemaVersion()` recipe.
### Eager materialization
Pass `eager: {}` to materialize indexes immediately after the schema
commit:
```ts
const evolved = await store.evolve(extension, { eager: {} });
```
Or pass options for finer control:
```ts
// Restrict to the extension kind whose index was declared in the
// proposal above.
const evolved = await store.evolve(extension, {
eager: { kinds: ["Paper"], stopOnError: true },
});
```
Omit `eager` to skip materialization and run `materializeIndexes()`
later.
Per-index failures throw `EagerMaterializationError` AFTER the new
`Store` is constructed and `ref.current` is updated, so the caller can
recover via the ref handle. The schema commit is **not** rolled back
if materialization fails — eager is convenience, not a transaction.
```ts
const ref = { current: store };
try {
await store.evolve(extension, { ref, eager: {} });
} catch (error) {
if (error instanceof EagerMaterializationError) {
// Schema is committed; ref.current is the new store.
log.warn(
{ failed: error.failedIndexNames },
"indexes did not materialize; will retry",
);
await ref.current.materializeIndexes();
} else {
throw error;
}
}
```
## Reaching extension kinds from the type system
TypeScript can't see kinds that don't exist at compile time. The
`Store` returned by `evolve()` keeps the same generic parameter as
the original — `evolved.nodes.Paper` would not type-check.
The escape hatch is `store.getNodeCollection(kind)` and
`store.getEdgeCollection(kind)`, which return a typed
`DynamicNodeCollection` / `DynamicEdgeCollection`:
```ts
const papers = evolved.getNodeCollection("Paper");
if (papers === undefined) {
throw new Error("Paper kind not registered on this store");
}
await papers.create({ title: "...", doi: "...", year: 2024 });
const all = await papers.find({});
```
The throwing variants `getNodeCollectionOrThrow(kind)` /
`getEdgeCollectionOrThrow(kind)` are the right call when the caller
already knows the kind has been evolved onto the store — they raise
`KindNotFoundError` with the offending `kindName`, `entity`, and host
`graphId` instead of returning `undefined`, so a typo fails loudly at
the call site rather than crashing later on `papers!.create(...)`.
`DynamicNodeCollection` exposes the same CRUD surface as
`store.nodes.X` — `create`, `getById`, `find`, `update`, `delete`,
etc. — but with `DynamicNode` element types since the specific Zod schema isn't
visible to TypeScript at the call site.
In TypeScript, nodes returned by a dynamic collection carry the nominal
`DynamicNode` type, preserving the requested kind literal. That proof lets
runtime kinds participate directly in Operational Identity without weakening
compile-time references to arbitrary string kinds:
```ts
const paper = await evolved
.getNodeCollectionOrThrow("Paper")
.create({ title: "Runtime schemas" });
await evolved.identity.assertSame(document, paper);
```
Identity reads can consequently return `IdentityNodeReference` values for
either compile-time or runtime kinds. See [Operational Identity](/identity).
For consumers that need the live Zod schema itself — MCP tool wrappers
that validate inputs before forwarding to `collection.create`, or
agent prompts that want richer JSON Schema than `introspect()`
exposes — `store.getNodePropsSchema(kind)` /
`getNodePropsSchemaOrThrow(kind)` (and the edge counterparts) return
the exact `z.ZodObject` the store uses internally. Identity holds:
`evolved.getNodePropsSchema("Paper")` is the same instance the store
parses against on `papers.create(...)`.
```ts
import { z } from "zod";
const schema = evolved.getNodePropsSchemaOrThrow("Paper");
const parsed = schema.parse(input); // same Zod issues as papers.create surfaces
const jsonSchema = z.toJSONSchema(schema); // for MCP tool descriptions
```
These accessors return only the props validator. Operation-level
checks — uniqueness, endpoint resolution, temporal validity, backend
constraints — still run only through `collection.create` / `update`.
See [Dynamic Props Schema Access](/schemas-stores#dynamic-props-schema-access)
for the full reference.
For codegen consumers, the kind set is reachable by iterating the
registry's `nodeKinds` and `edgeKinds` maps:
```ts
const allNodeKinds = [...store.registry.nodeKinds.keys()];
const allEdgeKinds = [...store.registry.edgeKinds.keys()];
const personType = store.registry.getNodeType("Person"); // NodeType | undefined
```
`KindRegistry` also exposes `hasNodeType(name)` / `hasEdgeType(name)`
for existence checks.
### Querying extension kinds
`store.query()` requires every `from` / `traverse` / `to` kind to be a
compile-time literal in `Store`. The string-keyed siblings
`fromDynamic` / `traverseDynamic` / `optionalTraverseDynamic` /
`toDynamic` admit kinds added via `evolve()` so an MCP server (or any
caller working from kind names in a string variable) can build typed
multi-hop traversals without `as any`:
```ts
const rows = await store.query()
.fromDynamic("Paper", "p")
.traverseDynamic("authoredBy", "a")
.toDynamic("Author", "u")
.whereNode("p", (p) => p.field("year").number().gte(2020))
.select((ctx) => ({ paper: ctx.p, author: ctx.u, edge: ctx.a }))
.execute();
```
Each method runtime-validates against the registry: a typo throws
`KindNotFoundError`, and a `toDynamic` target that isn't a valid
endpoint for the current edge / direction throws `EndpointError`.
Compile-time `from` / `traverse` / `to` are unchanged.
When the extension document is available in typed code, mint Store-bound
runtime-kind evidence from the exact persisted definition. The token narrows
collections and query aliases without asking callers to restate the schema in
Zod:
```ts
const tagKind = store.runtimeNodeKind("Tag", extension.nodes.Tag);
const taggedWithKind = store.runtimeEdgeKind(
"taggedWith",
extension.edges.taggedWith,
);
const tags = store.getNodeCollectionOrThrow(tagKind);
const rows = await store.query()
.fromDynamic(tagKind, "tag") // ctx.tag has the declared Tag fields
.traverseDynamic(taggedWithKind, "edge") // ctx.edge is narrowed too
.toDynamic("Document", "document")
.select((ctx) => ({ label: ctx.tag.label, weight: ctx.edge.weight }))
.execute();
```
Definitions, rather than hand-authored Zod schemas, are the type evidence:
TypeGraph compares the complete graph-extension declaration that TypeScript
infers against the definition persisted for that kind. This keeps refinements
such as enums, optionality, arrays, and numeric constraints on one authoritative
surface. Tokens are bound to the issuing Store and active schema hash; use a
fresh token after reopening or evolving a Store. Token lookups intentionally
use the throwing `getNodeCollectionOrThrow` / `getEdgeCollectionOrThrow`
variants because valid Store-issued evidence already proves that the kind
exists.
Predicate accessors on dynamic aliases use a `.field(name)`
discriminator:
- `BaseFieldAccessor` methods (`eq`, `isNull`, `in`, `notIn`) are
available directly on `field("name")`.
- Type-specific predicates sit behind a discriminator method that
asserts the field's type — `.string()` / `.number()` / `.date()` /
`.array()` / `.object()` / `.embedding()`. Each validates against the
registered Zod schema and throws `TypeError` on mismatch, so
`field("year").string()` against a number field is caught at
query-build time, not as a silent "method is undefined" later.
- `.field("missing")` throws when the property isn't on the schema.
#### Mixed typed and dynamic aliases
Typed and dynamic aliases interleave freely in one query. The
predicate accessor is resolved per alias — a typed alias keeps its
narrow `StringFieldAccessor` etc., while a dynamic alias gets `.field()`:
```ts
const rows = await store.query()
.from("Document", "d") // typed compile-time kind
.traverseDynamic("taggedWith", "e") // runtime edge
.toDynamic("Tag", "n") // runtime target
.whereNode("d", (d) => d.title.eq("the doc")) // typed: direct
.whereNode("n", (n) => n.field("label").string().eq("research")) // dynamic: discriminator
.select((ctx) => ({ doc: ctx.d, tag: ctx.n }))
.execute();
```
A typed `traverse("typedEdge", "e")` followed by `.toDynamic(target, "n")`
keeps the edge alias `e` typed — `e.role.eq(...)` works directly, no
discriminator needed. Only the dynamic-declared aliases use `.field()`.
#### Optional dynamic traversal
`optionalTraverseDynamic` is the LEFT-JOIN sibling — papers without
authors still surface, with the edge and target aliases as `undefined`:
```ts
const rows = await store.query()
.fromDynamic("Paper", "p")
.optionalTraverseDynamic("authoredBy", "a")
.toDynamic("Author", "u")
.select((ctx) => ({ paper: ctx.p, author: ctx.u, edge: ctx.a }))
.execute();
// row.author and row.edge are undefined for papers without an authoredBy edge.
```
### Search facade
The `store.search` facade — `fulltext`, `vector`, `hybrid`, and
`rebuildFulltext` — accepts any registered kind, compile-time or
runtime, with no type cast. The hit's `node` type narrows to the
concrete typed node only when the kind literal is statically known
in `Store`; extension kinds widen to the base `Node`. Misspelled
kind names throw `KindNotFoundError` at the call site instead of
returning empty results.
```ts
// Compile-time kind: hit.node.title is narrowed.
const compileTimeHits = await store.search.fulltext("Document", {
query: "climate",
limit: 10,
});
// Extension kind: same call shape, no cast. hit.node is the base
// `Node` shape since "Paper" isn't in the static `G`.
const runtimeHits = await store.search.fulltext("Paper", {
query: "attention transformer",
limit: 10,
});
```
## `store.introspect()`
`introspect()` returns a frozen snapshot of the merged schema and the
durable-state metadata the store has loaded so far. Its shape:
| Field | Type | Notes |
| ---------------------- | ------------------------------------- | ---------------------------------------------------------------------------------------------- |
| `graphId` | `string` | The graph's stable id. |
| `annotations` | `GraphAnnotations \| undefined` | Merged graph-scoped metadata from `defineGraph` and runtime extensions. |
| `kinds` | `readonly KindIntrospection[]` | Merged node kinds with `origin: "compile-time" \| "runtime"`, description, annotations, etc. |
| `edges` | `readonly EdgeIntrospection[]` | Merged edge kinds with the same origin discriminator and endpoint information. |
| `ontology` | `readonly OntologyIntrospection[]` | Ontology relations declared on either tier. |
| `deprecatedKinds` | `ReadonlySet` | Kinds flagged via `deprecateKinds(...)`. Informational, not a gate. |
| `extension` | `GraphExtension \| undefined` | The persisted graph-extension document, or `undefined` when no extensions have been committed. |
| `schemaVersion` | `number \| undefined` | Active schema version on the backend. `undefined` until the first commit. |
| `schemaHash` | `string \| undefined` | Hash of the active schema document. `undefined` under the same condition. |
```ts
const intro = store.introspect();
console.log(intro.schemaVersion); // e.g. 2
console.log(intro.annotations?.displayName);
console.log(intro.extension?.nodes?.Paper); // ExtensionNodeDef or undefined
console.log([...intro.deprecatedKinds]); // ["LegacyDocument"]
```
The `extension` field round-trips: passing it back through
`defineGraphExtension(intro.extension!)` and `evolve()` against an
empty graph reconstructs the same extension kinds.
For schema tooling that has an extension document but no Store, call
`introspectGraphExtension(extension)`. It compiles the extension through the
same TypeGraph compiler used by `evolve()` and returns `kinds` and `edges`
with JSON Schema `properties`, descriptions, annotations, and endpoint names.
The result describes only the supplied document; it has no graph ID, committed
schema version, or schema hash.
```ts
import { introspectGraphExtension } from "@nicia-ai/typegraph";
const declaration = introspectGraphExtension(extension);
console.log(declaration.kinds[0]?.properties);
```
Graph extensions may also carry graph-scoped annotations:
```ts
const extension = defineGraphExtension({
annotations: {
displayName: "Customer knowledge",
capabilities: { semanticSearch: true },
},
});
```
Annotation keys are shallow-merged. A later extension replaces the complete
value of each key it supplies; it does not recursively merge nested objects.
## Population statistics and stored-data validation
`await store.describe()` pairs the merged schema introspection with current
population statistics. It returns node and edge counts for every declared kind
plus present, explicit-null, and non-null counts (and non-null coverage) for
each directly addressable declared property:
```ts
const description = await store.describe();
const people = description.statistics.nodes.find(
(entry) => entry.kind === "Person",
);
console.log(people?.count);
console.log(
people?.properties.find((property) => property.path === "/email")?.coverage,
);
```
The schema coordinate includes the active schema version and hash when present
and a `schemaFence`. TypeGraph reads that coordinate before and after the
bounded, sequential SQL aggregate statements and refuses the result if it
changed. Node and edge properties are queried separately, and wide schemas are
split into fixed-width path batches. The database, rather than TypeGraph's
JavaScript process, computes all counts. Coverage follows ordinary nested JSON
Schema `properties`; TypeGraph intentionally does not invent population
semantics through `$ref`, unions, intersections, arrays, or conditionals.
`validateStore()` remains authoritative for those schemas. Concurrent writes
can affect different `describe()` path batches differently; the schema fence
detects schema changes, not data changes.
Use `validateStore()` to find rows that no longer satisfy a kind's current
declared Zod schema, for example after tightening a rule around existing data:
```ts
let cursor: string | undefined;
do {
const page = await store.validateStore({
entity: "node",
kind: "Person",
pageSize: 250,
...(cursor === undefined ? {} : { cursor }),
});
for (const failure of page.violations) {
console.log(failure.id, failure.path, failure.reason);
}
cursor = page.nextCursor;
} while (cursor !== undefined);
```
Undeclared properties are healthy semi-structured state and are never reported
as violations, including when the authored Zod object is strict. Each failure
names the record id, JSON-pointer path, top-level property when applicable,
Zod issue code, and reason.
`pageSize` is the number of records scanned, not a cap on violations: one record
can contribute several Zod issues. Each request performs a bounded SQL keyset
scan (`LIMIT pageSize + 1`) and reports `scannedCount`; it never materializes or
rescans the complete kind just to continue. Cursors bind the entity, kind,
schema fence, and last scanned id. A schema change throws
`StoreAnalysisCursorStaleError`. Data pages are deliberately live rather than a
claimed cross-request snapshot, so concurrent inserts, updates, and deletes can
affect later pages. Each page still reads the schema coordinate before and after
its data statement and refuses a concurrent schema flip.
Both analysis methods are current-only. They are absent from `StoreView`;
recorded/as-of population analysis is deferred until it can be backed by an
equally explicit temporal contract. Root-store calls use sequential SQL
statements plus schema bracketing, so they also work on non-interactive
transactional adapters.
Transaction callbacks expose the same methods through their pinned session.
Choose `repeatable_read` or `serializable` when every `describe()` aggregate or
every `validateStore()` page must observe one database snapshot, and consume
all validation pages before the callback returns:
```ts
await store.transaction(
async (tx) => {
const description = await tx.describe();
let cursor: string | undefined;
do {
const page = await tx.validateStore({
entity: "node",
kind: "Person",
...(cursor === undefined ? {} : { cursor }),
});
cursor = page.nextCursor;
} while (cursor !== undefined);
return description;
},
{ isolationLevel: "repeatable_read", accessMode: "read_only" },
);
```
Calling root `store.describe()` from inside a transaction callback still does
not join that transaction; use `tx.describe()` or `tx.validateStore()` for the
bound-session behavior.
## `store.materializeIndexes(options?)`
```ts
const result = await store.materializeIndexes();
// Restrict to specific compile-time or extension kinds.
const result = await store.materializeIndexes({ kinds: ["Paper"] });
const result = await store.materializeIndexes({ stopOnError: true });
```
`materializeIndexes` runs `CREATE INDEX` DDL for the indexes declared
on the merged graph and tracks per-deployment status in
`typegraph_index_materializations`. It's a separate verb from
`evolve()` because:
- DDL is **per-database**, not per-graph (two replicas of the same
`schema_doc` are still two databases — DDL has to run on each).
- Postgres uses `CREATE INDEX CONCURRENTLY` so live tables never take
an `AccessExclusiveLock`. CIC cannot run inside a transaction, which
is why `materializeIndexes` runs at the top-level backend, never
inside `transaction()`.
- Best-effort by default: per-index failures land in the result with
the captured `Error` and the loop continues. Pass
`stopOnError: true` to halt on the first failure.
The returned `MaterializeIndexesResult` has one entry per declared
index with `status: "created" | "alreadyMaterialized" | "failed" | "skipped"`.
The `skipped` status surfaces when the backend recognizes the
declaration but has no separate ANN index to build for it — e.g.
sqlite-vec (KNN lives in the `vec0` virtual table), SQLite without a
vector engine, or `embedding(dims, { indexType: "none" })` opting out
of automatic materialization.
Graph-extension-declared relational indexes use the same declaration shape as
compile-time `defineNodeIndex` / `defineEdgeIndex`, but in a
JSON-serializable form. They are persisted in `schema_doc.extension`,
re-derived on restart, and surface in `store.graph.indexes` with
`origin: "runtime"`.
### Vector indexes
Vector indexes are **auto-derived** from `embedding()` brands on both
compile-time and extension node kinds. Every top-level node field
declared with `embedding(dims, opts?)` produces one
`VectorIndexDeclaration` that flows through `materializeIndexes()`
like any relational index. No extra wiring required.
```ts
const Document = defineNode("Document", {
schema: z.object({
title: z.string(),
// Auto-derives a cosine HNSW vector index with pgvector
// defaults (m=16, ef_construction=64).
embedding: embedding(384),
}),
});
// Customize the auto-derived index by passing options at the brand.
const Image = defineNode("Image", {
schema: z.object({
embedding: embedding(512, { metric: "l2", m: 32, efConstruction: 100 }),
}),
});
// Opt out of automatic materialization while keeping the embedding.
const Manual = defineNode("Manual", {
schema: z.object({
embedding: embedding(384, { indexType: "none" }),
}),
});
```
Embeddings live in per-`(graphId, nodeKind, fieldPath)` typed tables
named `tg_vec___` (each carrying the field's
fixed dimension) — there is no single shared embeddings table. The
privileged migrator provisions each table plus a durable marker: at boot
via `createStoreWithSchema`, and for a field a runtime `evolve()`
introduces, by that `evolve()` call. The runtime hot path then asserts
the marker (never DDL), so a least-privilege role can read/write
embeddings.
On `materializeIndexes()`:
- Postgres with pgvector: emits `CREATE INDEX ... USING hnsw ...` (or
`ivfflat`) on the field's per-`(graphId, kind, field)` vector table
and reports `created`.
- SQLite with `sqlite-vec`: KNN lives in the `vec0` virtual table, so
there's no separate ANN index to build; declarations report
`skipped` (with a reason), not `failed`.
- libSQL / Turso: the DiskANN index is created via the strategy's own
DDL (`libsql_vector_idx` + `vector_top_k`).
- SQLite without a vector engine: declarations report `skipped` with a
reason indicating the backend lacks vector support.
The vector declaration's identity key within a single graph is
`(kind, fieldPath)` — v1 allows at most one vector index per
(kind, field) pair. The auto-derived deterministic declaration
name is `tg_vec_{kind}_{field}_{metric}` — clean and scannable for
inspection in `pg_indexes` and result entries. Changing the metric
requires a different declaration name and explicit
re-materialization.
Cross-graph disambiguation lives at the materialization boundary,
not in the declaration name. Vector status rows in
`typegraph_index_materializations` are keyed on the compound
`{graphId}::{declaration.name}` for both auto-derived and explicit
`VectorIndexDeclaration` entries — so two graphs reusing the same
declaration name (whether auto-derived from the same kind/field or
constructed explicitly via `defineGraph({ indexes: [...] })`) don't
collide in the status table. Each graph's `materializeIndexes()`
call creates its own physical pgvector index on that graph's
per-`(graphId, kind, field)` vector table and records its own status row.
### Fulltext indexes (out of scope for v1)
Fulltext indexes are NOT in the unified declaration channel for v1.
The fulltext table's canonical index (Postgres GIN on `tsv`, SQLite
FTS5 virtual table) is created with the table itself by
`bootstrapTables` per the active `FulltextStrategy`. Per-kind fulltext
indexes are an "advanced strategy" surface that doesn't fit the
relational-style declaration model and is reserved for future work.
### Caveats (Postgres)
- `IF NOT EXISTS` does not validate shape — only that something with
that name exists. Drift detection uses TypeGraph's recorded
signature, not PG metadata. Signature mismatch surfaces as
`failed` with a `different signature` message.
- Failed `CONCURRENTLY` builds leave invalid indexes
(`pg_index.indisvalid = false`). v1 surfaces this as a `failed`
result; the operator drops the invalid index manually before retry.
## `store.deprecateKinds(...)` / `undeprecateKinds(...)`
```ts
await store.deprecateKinds(["LegacyDocument"]);
console.log([...store.introspect().deprecatedKinds]); // ["LegacyDocument"]
await store.undeprecateKinds(["LegacyDocument"]);
```
Soft-deprecation surfaces in
`store.introspect().deprecatedKinds: ReadonlySet` for
introspection (codegen, UI tooling, lints) but does not gate reads,
writes, or queries. Bumps the schema version like any other change;
idempotent — re-deprecating an already-deprecated kind is a no-op.
Use cases:
- Codegen routes around deprecated kinds when generating new client
code.
- Lint rules flag new code that touches deprecated kinds.
- UI tooling hides deprecated kinds from picker menus.
## `store.removeKinds(...)` / `materializeRemovals()`
`removeKinds()` removes graph-extension-declared kinds from the active schema.
It is intentionally two-phase:
1. **Schema commit.** `removeKinds(names)` rewrites the persisted graph
extension without the named graph-extension kinds, cascades extension edges
and ontology relations that can no longer resolve, and commits a
new schema version with CAS.
2. **Data cleanup.** `materializeRemovals()` deletes rows for removed
node and edge kinds on the current deployment.
```ts
const withoutPaper = await evolved.removeKinds(["Paper"]);
await withoutPaper.materializeRemovals();
```
Pass `{ eager: {} }` to run cleanup inline after the schema commit:
```ts
const withoutPaper = await evolved.removeKinds(["Paper"], { eager: {} });
```
Removing an embedding field from a surviving kind orphans its
per-`(graphId, kind, field)` `tg_vec_*` table; `materializeRemovals()`
reclaims it and reports the count in
`MaterializeRemovalsResult.reclaimedVectorFields`.
For source-dependent edges, removing a node kind removes its source entry and
any target references to it. A source entry whose targets are exhausted is also
removed. The edge kind survives while another valid pair remains; it is cascaded
only when no pairs remain. For example, removing `Course` from the `assignedTo`
extension above preserves `Employee → Department`.
Removal only applies to graph-extension-declared kinds. Removing a compile-time
kind throws `RemoveCompileTimeKindError`; deploy new TypeScript code
for compile-time schema removal. Removing a graph-extension kind that is still
referenced by a compile-time edge or ontology relation throws
`KindHasReferentsError`, because TypeGraph cannot rewrite your
compiled graph for you.
## Restart parity (the load-bearing invariant)
The graph extension is the **durable source of truth**.
Every call to `evolve()` persists the merged document into
`schema_doc.extension`. On startup, `createStoreWithSchema()`
reads it back, runs the same compiler, and reconstructs identical
Zod-bearing `GraphDef`. Net: an extension kind defined via `evolve()` is
indistinguishable from a compile-time kind after restart.
Verify this in your own tests:
```ts
const [store] = await createStoreWithSchema(baseGraph, backend);
const evolved = await store.evolve(proposal);
await evolved.getNodeCollection("Paper")!.create({ title: "...", doi: "...", year: 2024 });
// Different process / different deployment / fresh store...
const [restored] = await createStoreWithSchema(baseGraph, backend);
expect(restored.registry.hasNodeType("Paper")).toBe(true);
const all = await restored.getNodeCollection("Paper")!.find({});
expect(all).toHaveLength(1);
```
## Multi-process safety
Concurrent writers compete on the `commitSchemaVersion` CAS. One wins;
the loser sees one of two errors with very different recovery semantics:
- **`StaleVersionError`** — the local view of the active version is out
of date. Routine race signal: refetch and retry.
- **`SchemaContentConflictError`** — a different writer wrote a row at
the same version with a different content hash. NOT a routine race.
Two writers tried to commit semantically different schemas at the
same version, which means one of them is operating on an inconsistent
view of the world. Surface to the operator; do not blindly retry.
Retry recipe (only catches `StaleVersionError`):
```ts
async function evolveWithRetry(
ref: StoreRef>,
extension: GraphExtension,
attempts = 3,
): Promise> {
for (let attempt = 0; attempt < attempts; attempt++) {
try {
return await ref.current.evolve(extension, { ref });
} catch (error) {
if (error instanceof StaleVersionError) {
// Refetch happens implicitly inside evolve()'s next call —
// catch-up auto-merges the persisted state into the local
// baseline, so the next attempt diffs against fresh state.
continue;
}
// SchemaContentConflictError, GraphExtensionValidationError,
// EagerMaterializationError, etc. all surface to the caller —
// they require operator intervention or different handling, not
// blind retry.
throw error;
}
}
throw new Error(`Failed to evolve after ${attempts} attempts`);
}
```
The internal `#catchUpToStored` step inside `evolve()` (and
`deprecateKinds`, `materializeIndexes`) folds the persisted graph-extension
document and deprecation set into the local baseline before computing
the next state, so a stale store applying an extension on top of an
out-of-date baseline doesn't trample another writer's progress.
## Trust boundary
When the graph extension originates from an **untrusted
source** — an LLM completion, user input, an external API — treat it
as untrusted data. Specifically:
- **Validation runs at the boundary.** `defineGraphExtension(doc)`
rejects any input that doesn't match the v1 subset
(`GraphExtensionValidationError` with per-issue paths). Don't skip
this step. If you want Result-style handling for untrusted JSON, call
`validateGraphExtension(raw, { strict: true })` and surface the
structured issues before calling `evolve()`.
- **Property types are deliberately small.** The supported set
excludes things like `bigint`, `Date`, custom Zod refinements, and
arbitrary functions. An LLM cannot inject executable code by
proposing an extension document.
- **Operator approval is your gate.** The library doesn't enforce
human-in-the-loop — your application does. Show the diff to a human
before calling `evolve()`.
- **Persisted documents are part of your data.** They're stored in
`schema_doc` along with every other schema artifact; back them up,
audit them, version-control them.
## Out of scope for v1
- **Fulltext index unification.** Vector indexes flow through the
unified channel (auto-derived from `embedding()` brands). Fulltext
is still per-strategy: the GIN / FTS5 index is created with the
fulltext table at `bootstrapTables` time. Per-kind fulltext indexes
are reserved for future work.
- **Multiple vector indexes per (kind, field).** v1 allows at most
one. To use a different metric for the same field, use a different
field name or wait for v2.
- **Hard-blocking reads/writes on deprecated kinds.** Deprecation is
informational. If you want strict enforcement, wrap collection
access yourself.
- **Auto drop+recreate on signature drift.** `materializeIndexes`
surfaces drift as a `failed` result; manual remediation is required
to avoid risky lock semantics.
## See also
- [Schema Migrations](/schema-management) — the lower-level primitives `evolve()` rides on.
- [Evolving Schemas](/schema-evolution) — recipes for compile-time schema changes.
- [Errors](/errors) — `EagerMaterializationError`, `GraphExtensionValidationError`, `StaleVersionError`, `SchemaContentConflictError`.
# Graph Merge
> Branch a TypeGraph store, let many writers edit it independently, and fold their work back into one canonical graph with deterministic entity resolution, conflict reporting, edge repointing, and provenance.
Graph Merge turns a TypeGraph store into something you can **fork, edit in
parallel, and reconcile** — the way you already fork, branch, and merge code.
Several writers (agents, importers, reviewers, background workers) each build
graph changes in isolation, and a single deterministic step folds them back into
one canonical graph: duplicate entities are resolved, edges are repointed onto
the survivors, disagreements are surfaced (never silently overwritten), and you
get a full report of what happened and who contributed it.
It ships as a core package subpath:
```typescript
import { branch, merge } from "@nicia-ai/typegraph/graph-merge";
```
Everything here is defined over ordinary TypeGraph stores, schemas, indexes,
backends, and ontology semantics — there is no separate service to run.
## What you can build
Graph Merge exists because "append everything" is the wrong default for graphs:
it produces duplicate entities and dangling relationships. With a real merge
primitive you can build:
- **Multi-agent knowledge-graph construction.** Run N extraction agents in
parallel, each on its own branch, then merge. The same real-world entity
discovered by three agents collapses to one canonical node; every agent's
edges follow it; disagreements come back as conflicts to adjudicate.
- **Parallel ETL / import reconciliation.** Ingest an EHR export, a claims
feed, and a lab feed as independent branches and reconcile them into one
patient-care graph — by exact identifier, blocking key, or fuzzy name match.
- **Master-data / entity dedup (CRM, FHIR, catalogs).** Use declared `unique`
constraints as definitional identity and similarity scoring for the rest.
- **Human-in-the-loop review queues.** `planMerge()` returns the exact proposed
write set, conflicts, and entity-resolution evidence without changing the
target. Persist that JSON artifact, review it in another process, and apply
the reviewed bytes later with `applyMergePlan()`.
- **Incremental ingestion against a live graph.** `mergeIncremental()` lets new
batches land on a target that has *advanced* since the branch was taken,
re-discovering already-committed entities instead of duplicating them.
- **Semantic deduplication.** Plug in an embedder for `vector` or `hybrid`
similarity to collapse near-duplicates that exact and trigram matching miss.
The throughline: **isolation while writing, determinism while merging, and a
report you can act on.**
## How it works
The mental model is a three-act lifecycle:
1. **`branch()`** stamps the base store's `base@V` and materializes an
isolated, independently-mutable working copy. With `revisionTracking: true`
(or `history: true`), `base@V` uses the store's durable revision anchor: a
per-graph random origin plus a monotonic clock. Validation therefore does
not fingerprint every live row or mistake a coincident revision in a
separately created store for the branch's base. Existing stores retain the
schema-and-content-fingerprint fallback. Writers edit the working copy with
the normal store API; the base is never touched.
2. Writers do whatever they want — create nodes/edges, modify inherited rows,
delete inherited rows.
3. **Plan, then apply.** `planMerge()` diffs every branch against the base and
runs a fixed planning pipeline. `applyMergePlan()` validates the serialized
artifact and its digest, checks its revision fence inside the write
transaction, then mechanically applies the already-resolved writes:
```text
stage (diff every branch)
→ generate candidates (exact unique · blocking key · similarity)
→ cluster (group nodes that are the same entity)
→ canonicalize (pick a survivor, union properties, resolve conflicts)
→ repoint + dedupe edges onto survivors
→ reconcile delete/modify and types
→ emit a revision-fenced JSON plan
→ validate + commit transactionally + build the report
```
`merge()` remains the one-call convenience wrapper over this same lifecycle;
it plans and immediately applies. If the target Store carries a reconciled
schema version, its commit acquires
and validates the normal schema-write fence before row DML. A raw target
remains outside that guarantee. PostgreSQL serialization failures are retried
automatically around the complete merge commit.
The pipeline is **deterministic by construction**: candidate sets are sorted,
clusters resolve by stable keys, and every conflict is decided on an explicit
`branchOrder` (or lexicographic branch id) — *never* wall-clock arrival. Merging
the same branches in any order yields the same committed graph and the same
normalized report. That property is what makes a merge safe to retry, cache, and
reason about.
## Quick start
Create a base store, fork one branch per writer, write to the branch stores,
then merge them back into the target.
```typescript
import { createStoreWithSchema } from "@nicia-ai/typegraph";
import { asBranchId, branch, isOk, merge, unwrap } from "@nicia-ai/typegraph/graph-merge";
const [base] = await createStoreWithSchema(graph, baseBackend, {
// Recommended for graphs that branch repeatedly or stay live while agents work.
revisionTracking: true,
});
// branch() is backend-agnostic: you supply a factory for each branch's backend.
const makeBranchBackend = async () => createFreshBackend();
const sourceA = unwrap(await branch(base, makeBranchBackend, { id: asBranchId("source-a") }));
const sourceB = unwrap(await branch(base, makeBranchBackend, { id: asBranchId("source-b") }));
await sourceA.store.nodes.Patient.create({ name: "Anna Rivera", birthDate: "1974-03-09", mrn: "MRN-001" });
await sourceB.store.nodes.Patient.create({ name: "Ana Rivera", birthDate: "1974-03-09", mrn: "MRN-001" });
const result = await merge(base, [sourceA, sourceB], {
resolve: {
Patient: {
block: (node) => node.mrn ?? node.birthDate,
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.78,
},
},
onPropertyConflict: "flag",
branchOrder: [sourceA.id, sourceB.id],
});
if (!isOk(result)) throw result.error;
console.log(result.data.resolutions); // the two patients collapsed to one
console.log(result.data.conflicts); // the "Anna" vs "Ana" spelling disagreement
```
`branch()` returns a `Result`; `unwrap` throws on failure (or branch on
`isOk`). The default working-copy strategy clones the base through TypeGraph's
streaming interchange, so each branch gets a fresh backend from your factory
without building a graph-sized export document in memory.
## Reviewable plan/apply lifecycle
Use the two-step API when approval must happen before accepted graph truth
changes. The target must have `revisionTracking: true` or `history: true` so the
plan can carry a durable, store-specific revision fence.
```typescript
import {
applyMergePlan,
applyMergePlanInTransaction,
isOk,
planMerge,
} from "@nicia-ai/typegraph/graph-merge";
const planned = await planMerge(base, [sourceA, sourceB], {
resolve: {
Patient: {
block: (node) => node.mrn ?? node.birthDate,
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.78,
},
},
});
if (!isOk(planned)) throw planned.error;
// Persist outside the target graph, or send to a separate review process.
const stored = JSON.stringify(planned.data);
const reviewed = JSON.parse(stored);
// Later, against the same unchanged target:
const applied = await applyMergePlan(base, reviewed);
if (!isOk(applied)) throw applied.error;
console.log(applied.data.merged); // actual committed effects
```
Planning does not mutate the target. The public plan contains only JSON-safe,
deterministically ordered data: its target/schema/revision fence, resolved write
set, review information, match evidence, and a stable content digest. It never
contains a Store, backend, `Map`, `Set`, callback, or embedder. Applying it does
not re-run blocking, candidate generation, similarity scoring, embeddings,
canonical selection, or conflict callbacks. The final report copies the
reviewed entity-resolution evidence unchanged.
The envelope is deliberately explicit: `formatVersion` selects the wire schema;
`digest` identifies its canonical content; `mode`, `target`, and `anchors` state
what was observed; `proposed` summarizes the review; `writes` is the complete
mechanical write set; and `review` holds the conflicts, resolutions, evidence,
diagnostics, warnings, and other report material known before apply.
### Apply a plan with application writes
Use `applyMergePlanInTransaction(target, tx, artifact)` when the merge, a graph
receipt or anchor, and application SQL must share one caller-owned commit. Build
`tx` by passing the native transaction to the **same target Store's**
`withRecordedTransaction()` callback. Apply the plan before any other write to
the target graph in that transaction; after it returns, the callback may make
more graph writes and the caller may run more SQL on the native handle.
```typescript
await db.transaction(async (nativeTx) => {
const { result: report, receipt } = await target.withRecordedTransaction(
nativeTx,
async (tx) => {
const applied = await applyMergePlanInTransaction(target, tx, reviewed);
await tx.nodes.MergeReceipt.create({
planDigest: reviewed.digest.value,
mergedNodes: applied.merged.nodes,
});
return applied;
},
);
await nativeTx.insert(mergeRuns).values({
planDigest: reviewed.digest.value,
recordedAt: receipt.recorded,
mergedNodes: report.merged.nodes,
});
}); // await this outer commit before reporting success
```
The adopted applier returns `Promise` and throws a typed
`MergeError` on refusal or failure. It does not open, commit, roll back, or retry
a transaction. Let the exception reject the outer callback so all merge and
application writes roll back together. Never catch it inside the transaction
and then commit. When the driver reports a retryable transaction failure, retry
the entire outer transaction, including the application writes; do not add an
inner retry or nested transaction around the merge.
PostgreSQL requires the transaction's observed isolation to be `READ COMMITTED`.
SQLite acquires its serialized writer slot before checking the plan fence. The
plan must explicitly have `persistProvenance: false`: atomic sidecar provenance
persistence is refused on this path. `includeInReport` remains supported, so
the returned report can still contain the in-memory provenance index.
On a history store, `receipt.recorded` is allocated after the
`withRecordedTransaction()` capture callback returns. The caller can persist
that anchor with application SQL on `nativeTx` before the outer commit, as the
example above does. Await the outer commit before treating the report or receipt
as durable.
The plan's `proposed` summary describes **proposed changes**. It deliberately
does not call them “merged”: `MergeReport.merged` is reserved for the actual
effects returned after a successful transaction. Coalescing and idempotent
identity operations can make actual counts differ from the proposal.
### Merge after schema evolution in one caller transaction
Prepare the evolution first, then call
`planMergeForEvolution(target, evolutionPlan, branches, options?)` outside the
write transaction. This route resolves writes against the graph produced by
the evolution plan while checking the current target's durable data and
revision fence. The serialized merge plan names the resulting schema
version/hash. If the target schema or revision changes during planning, the
planner refuses the artifact; replan outside the transaction.
Branches forked from the original baseline can merge existing kinds. To
include a newly added kind, call
`branchForEvolution(target, evolutionPlan, makeBackend)` before the caller
transaction (on PostgreSQL, pass the working-copy manager's `makeBackend`; see
[PostgreSQL table-backed working copies](#postgresql-table-backed-working-copies)),
then add data on that isolated branch. The planner accepts
branches from either one matching baseline; a mixed set of old-schema and
resulting-schema forks is refused.
Pass `{ revisionJournal: false }` as the fourth `branchForEvolution()` argument
when its working copy does not need journal-backed changed-key lineage. The
branch remains revision-tracked, and merge planning uses the portable diff
when no other lineage source is available.
```typescript
const evolutionPlan = await target.planEvolution(extension);
const futureBranch = unwrap(
await branchForEvolution(target, evolutionPlan, makeIsolatedBackend),
);
try {
await futureBranch.store.getNodeCollectionOrThrow("Tag").create({ label: "New" });
const mergePlan = unwrap(
await planMergeForEvolution(target, evolutionPlan, [futureBranch]),
);
await db.transaction(async (nativeTx) => {
const { result: report, receipt } = await target.withEvolvedTransaction(
nativeTx,
evolutionPlan,
(tx) => applyMergePlanInTransaction(target, tx, mergePlan),
);
await nativeTx.insert(mergeRuns).values({
mergedNodes: report.merged.nodes,
schemaVersion: receipt.schema.version,
});
});
} finally {
await futureBranch.close();
}
```
Apply the merge before other graph writes in the evolved callback. The
applier uses the evolved graph and checks the plan's resulting schema and
revision fences on the same caller session. Passing a merge plan for the old
schema refuses before merge mutation. Evolution's schema CAS is not treated
as a prior callback entity write. Roll back the entire native transaction on
any refusal; the schema change, merge, recorded capture, and application SQL
then roll back together. The report and receipt are provisional until the
outer commit succeeds. An adapter configured with
`schemaProvisioning: "transactional"` can provision required identity or vector
storage on the same native session before the merge callback. The default
DML-only policy refuses such requirements before the schema fence or merge
mutation. Bootstrap base storage before adopting either route; run generic
eager index maintenance separately after the outer commit.
### Candidate write sets for a planned schema
`planCandidateWriteSetForEvolution()` is the branch-free counterpart for a
bounded candidate batch. First use
`captureCandidateWriteSetTargetForEvolution(target, evolutionPlan)` when
authoring the JSON document; it records the evolution plan's resulting schema
identity rather than the currently active one. The planner stages the candidate
against that resulting graph and returns the same resulting-schema merge
artifact accepted by `withEvolvedTransaction()`.
Candidate resolution still includes the committed target as an accepted source.
Existing unique matches and property conflicts are therefore visible in the
reviewed plan before the evolution transaction begins, rather than surfacing as
late write-time failures.
```typescript
const evolutionPlan = await target.planEvolution(extension);
const writeSet = {
formatVersion: 1 as const,
sourceId: "import-batch-42",
target: captureCandidateWriteSetTargetForEvolution(target, evolutionPlan),
nodes: [{
kind: "Tag",
id: "import-batch-42:tag-1",
properties: { label: "Research" },
validFrom: "2026-01-01T00:00:00.000Z",
}],
edges: [],
};
const mergePlan = unwrap(await planCandidateWriteSetForEvolution({
target,
evolutionPlan,
makeBackend: makeIsolatedBackend,
writeSet,
}));
await db.transaction(async (nativeTx) =>
target.withEvolvedTransaction(nativeTx, evolutionPlan, (tx) =>
applyMergePlanInTransaction(target, tx, mergePlan),
),
);
```
The schema change and accepted candidate writes share the caller's one
transaction and recorded revision. If another writer changes the target while
planning, `MergePlanningStaleError` is an expected concurrency result: discard
the candidate plan, recapture the target for a new evolution plan, and replan.
For a frozen ancestor and a live destination, use the named incremental planner:
```typescript
const planned = await planMergeIncremental({
forkPoint,
target,
branches,
options,
});
if (!isOk(planned)) throw planned.error;
const applied = await applyMergePlan(target, planned.data);
```
When `target` records history, a durable branch can use its sealed recorded
fork point without keeping a second frozen Store:
```typescript
const forkPoint = created.branch.recordedForkPoint;
if (forkPoint === undefined) throw new Error("History was not captured at fork");
const planned = await planMergeIncremental({
forkPoint,
target,
branches: [created.branch],
options: { onBasePropertyConflict: "flag" },
});
```
`recordedForkPoint` is available when the source captured history at fork time;
it contains both the recorded instant and the branch's `base@V` token. The
planner reads ancestor rows from the target's recorded relations, validates the
origin, schema, and revision anchor, and enumerates only changed keys when
lineage can prove a complete delta. A missing or incompatible anchor is refused
before planning. The direct `mergeIncremental()` wrapper accepts the same fork
point. Keep the durable descriptor with the branch: reopening restores the
recorded fork point from the sealed origin.
The same target revision must still be current when the reviewed plan is
applied. If it moved during planning, planning returns
`MergePlanningStaleError` and no artifact. This is an expected retry-and-replan
outcome under concurrency: recapture the target, create a new plan, and review
its new digest before retrying. If it moved afterwards, `applyMergePlan()`
returns `StaleMergePlanError` before plan writes. Re-plan, review the new
digest and proposal, then apply the new artifact; never edit an
old plan or retry it as though it still represented the target. A successful
plan is single-use: a second or concurrent application is stale.
Persisting a plan or approval in the target graph also advances this revision.
For exact-plan approval, use external storage or a separate graph ID; writes to
that graph do not advance this target's revision. This does not provide atomic
writes across graphs, and any intervening target write still requires a fresh
plan. For candidate batches whose review records belong in the target itself,
use the durable review protocol below.
`merge()` and `mergeIncremental()` remain convenient compatibility wrappers.
They invoke the same planner and applier contiguously and return the same
`MergeReport` shape as before, now with match evidence on each resolution. Use
the wrappers when no external approval boundary is needed.
:::caution[Sensitive plans and trust]
A plan contains the complete resolved writes and may therefore contain personal,
regulated, or otherwise sensitive application data. Protect it like the source
graph: encrypt it where appropriate, restrict access, and avoid logging it. The
digest identifies the exact canonical artifact and detects accidental or
unrecorded changes. It is **not** a signature, proof of origin, authentication,
or authorization. Authenticate untrusted storage and authorize the caller before
passing a plan to `applyMergePlan()`.
:::
### Durable candidate review in the target graph
`planCandidateWriteSetReview()` separates immutable review evidence from a
revision-bound execution plan. Its `MergeReviewArtifact` retains the original
candidate write set, reviewed plan, normalized merge options, explicit policy
identity/context, and target baseline. You can persist this artifact and later
approval records in the target before calling
`revalidateCandidateWriteSetReview()` to compute a fresh execution plan.
Both review versions support candidate write sets only. They do not rebase arbitrary artifacts
from `planMerge()` or `planMergeIncremental()`.
Candidate planning on revision-tracked graphs reads existing candidate ids and
edge endpoints by key, then seeds only those rows in the transient working
copy. On identity-enabled graphs, it also follows live same-id peers and
current identity assertions from those references to a fixed point. The
planner reads peers of a candidate edge with `one` cardinality by source,
peers of a `unique` edge by its endpoint pair, and the active peer of a
`oneActive` edge by source. The active-only read checks an open `validTo` even
when `validFrom` is in the future, and does not return ended history. These reads
let the transient copy enforce the same cardinality rule as a complete clone.
On graphs with ontology relations, it also reads live nodes sharing each
candidate reference's id across kinds, so disjointness sees the same peers as a
complete clone. Ontology subtype relationships remain graph metadata.
The candidate diff and its target baseline are bounded to that dependency set and
any committed rows recalled by configured unique or index sources. Planning
still fences the target revision before and after these reads. With edge
match-identity constraints, a backend offering `findEdgesByMatchIdentity`
seeds the exact durable owners named by the candidate. A missing keyed read,
an owner excluded from the clone projection, or a target without revision
tracking uses the complete clone path. A custom backend lacking the optional
`findActiveEdgesBySourceV1` read also uses that path for `oneActive` graphs.
On the complete clone path, when the copy and target really share one serialized connection, clone export
is materialized before import, but its snapshot still holds the connection's
exclusive stream lease while it is collected. Concurrent review calls on that
resource can therefore return a merge error caused by a `ConfigurationError`
with `details.code: INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` or
`INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS`. Await the whole review call before
starting another on the same serialized resource. A `pg.Pool` with more than one
connection is not one serialized resource; do not declare its pool object as
`{ mode: "shared" }` just because the working copies use that pool. See
[Serialized connections](/backend-setup#serialized-connections).
The following continues the [candidate write set example](#constraint-aware-ingestion-branches).
`Artifact`, `Decision`, and `evidence` are application-defined node/edge kinds;
`proposal` is an existing node. The target enables `history` or `revisionTracking`.
```typescript
import {
applyMergePlan,
planCandidateWriteSetReview,
revalidateCandidateWriteSetReview,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const policy = {
id: "acceptance-policy-v1",
context: { requiredApprovals: 1, resolverVersion: "2026-09-01" },
};
const review = unwrap(
await planCandidateWriteSetReview({
target: store,
makeBackend,
writeSet,
policy,
}),
);
const artifact = await store.nodes.Artifact.create(
{ content: JSON.stringify(review) },
{ id: review.digest.value },
);
await store.edges.evidence.create(proposal, artifact, { note: "review" });
// After an authenticated reviewer approves under the application's policy:
const decision = await store.nodes.Decision.create({
approved: true,
reviewDigest: review.digest.value,
});
await store.edges.evidence.create(decision, artifact, { note: "approval" });
// Later: authenticate the stored artifact and decision, then check current
// authorization, approval validity, and policy before reusing that approval.
const persisted = await store.nodes.Artifact.getById(artifact.id);
if (persisted === undefined) throw new Error("Missing review artifact");
const checked = unwrap(
await revalidateCandidateWriteSetReview({
target: store,
makeBackend,
review: JSON.parse(persisted.content),
policy,
}),
);
if (checked.status !== "compatible") {
console.log(checked.differences);
throw new Error("Create a new review and obtain a new approval");
}
if (checked.reviewDigest.value !== decision.reviewDigest) {
throw new Error("Approval does not identify the validated review");
}
// Keep this fresh execution plan ephemeral: another target write makes it stale.
const applied = unwrap(await applyMergePlan(store, checked.plan));
// A separate commit AFTER successful apply; see the recovery boundary below.
await store.nodes.Artifact.create({
content: JSON.stringify({
reviewDigest: checked.reviewDigest,
approvalId: decision.id,
executionPlanDigest: checked.plan.digest,
executionTarget: checked.plan.target,
report: applied,
}),
});
```
All approval records and links must be committed before final revalidation.
Do not persist each replacement execution plan in the target: that repeats the
staleness cycle. Retain the original review, and use the returned `reviewDigest`
plus the fresh plan's `digest` and `target` fence to relate approval to execution.
Both review APIs return `Result<..., MergeError>`. Revalidation accepts the
persisted artifact as `unknown` and replans its retained candidate input once
target, policy, and baseline checks pass.
Supply current merge `options` and `policy` again; callbacks are never restored
from serialized data.
| Revalidation status | Meaning and next step |
| --- | --- |
| `compatible` | Includes a fresh `plan` and the original `reviewDigest`. Application policy may reuse approval; authorize the action and apply promptly. Compatibility itself grants no permission. |
| `changed` | `differences` identify changed policy/options, baseline entities/identity, or plan fields. Obtain a new review and approval. A `plan` is included only when fresh planning completed. |
| `incompatible` | The graph ID, schema identity, or revision origin differs. Approval cannot be reused for this target; resolve the mismatch and create a new review. |
Malformed/unsupported artifacts, mismatched digests, and missing required
evidence return `MergeReviewError` (`GRAPH_MERGE_REVIEW`). Existing typed planning
and constraint errors remain errors rather than compatibility statuses. A target
change during evidence capture/planning returns `MergePlanningStaleError`.
The V1 baseline is deliberately conservative:
- Every original node and edge row, including tombstones and validity metadata,
must remain unchanged. Editing an old audit record requires a new review even
when the candidate's resolved writes would be identical.
- Expected absences for candidate/write/guard references must remain absent.
Same-ID nodes of other kinds are also guarded, because they can change implicit
identity membership. Complete archival identity evidence must remain unchanged.
- Newly added rows can coexist with approval only when fresh planning produces
identical resolved writes, guards, conflicts, evidence, provenance, and other
plan content. Candidate-derived anchors and the execution digest/fence are
regenerated. There is no exemption for an “audit” kind.
For an eligible revision-tracked graph, pass
`reviewScope: "candidate"` to `planCandidateWriteSetReview()` to emit V2
candidate-scoped evidence. V2 fingerprints the candidate's node and edge ids,
edge endpoints, resolved writes, and plan guards, including expected absences
across kinds. On Operational Identity graphs it also records the reachable
identity assertion and same-id peer closure, plus assertion-ID collision
evidence. Revalidation expands that retained identity scope, rereads the
referenced rows, and replans the candidate under a new target fence. An unrelated original row may change
without invalidating V2 when it cannot affect the fresh resolved plan; V1
would report that row change. Applications whose approval policy needs the
V1 whole-graph rule should omit `reviewScope`. The review artifact records
its version and scope, so revalidation applies the rule originally reviewed.
Candidate-scoped review refuses graphs outside those eligibility rules.
On a `oneActive` graph, a custom backend must expose
`findActiveEdgesBySourceV1` for candidate-scoped review; the complete-clone
candidate planner and V1 review remain available when it does not.
On an Operational Identity graph, a custom Store runtime must also expose
endpoint-scoped and assertion-ID-scoped identity reads. Without both reads,
ordinary candidate planning uses the complete working-copy clone and V1 review
remains available; an explicit V2 candidate-scoped review request is refused.
Applicable store constraints still run during atomic application. Compatibility
does not promise that apply will succeed: new rows may introduce constraint
conflicts, and any write between revalidation and apply causes
`StaleMergePlanError`. A failed application commits no partial candidate node,
edge, or identity writes. Revalidate again after a stale refusal; require reapproval if
the result changes.
`policy.id` identifies your policy implementation; `policy.context` explicitly
records every opaque dependency that can change its decision. Include callback
and resolver versions, model/prompt versions, external configuration or data
versions, and any application state used to authorize approval reuse. Use an
empty context only when no such dependencies exist. TypeGraph captures callback
presence and serializable options, but cannot discover callback code, closure
state, external reads, or hidden application policy dependencies.
The producer must supply complete evidence, and the application must authenticate
the entire stored review and its approval. Content addressing and SHA-256 detect
content changes; anyone able to replace evidence can recompute a digest. A valid
digest is neither proof that the baseline was complete nor authorization to
reuse approval. Enforce artifact immutability and access control in your storage
or application. The review contains candidate data and an entire reviewed plan,
so protect it with the same care as graph data.
V1 review capture and revalidation read and fingerprint the complete target
graph and archival identity ledger. The artifact stores one fingerprint per
original row plus expected absences. Budget graph-sized reads and artifact
storage for V1. V2 candidate-scoped review uses bounded point and identity
closure reads for its baseline on eligible graphs.
The execution receipt above is a separate commit. If its write fails or the
process stops after apply, the merge may already be committed without a receipt.
Retain the original review and approval, and reconcile committed history and
application operation identity before repairing the receipt. Do not treat a
missing receipt as permission to replay the candidate; applying its old execution
plan is stale, and replanning is not a duplicate-execution check. To commit the
receipt atomically with the merge, create it in an `afterApply` callback as
described in [Composing application checks and writes](#composing-application-checks-and-writes).
Review revalidation alone does not add that guarantee.
See the runnable [durable merge review example](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/27-durable-merge-review.ts)
for the complete schema and lifecycle.
## Composing application checks and writes
Pass execution callbacks to `applyMergePlan()` when a reviewed candidate and
related application records must commit together:
```typescript
const result = await applyMergePlan(target, reviewedPlan, {
beforeApply: async (reads) => {
const resource = await reads.nodes.Resource.getById(resourceId);
if (resource?.owner !== "unclaimed") {
throw new Error("Resource is already claimed");
}
},
afterApply: async (tx, applied) => {
await tx.nodes.Resource.update(resourceId, { owner: "accepted" });
const decision = await tx.nodes.Decision.create({
status: "accepted",
changedNodes: applied.merged.nodes,
});
await tx.edges.decides.create(decision.id, resourceId, {});
},
});
if (!isOk(result)) throw result.error;
// The plan and application writes have now committed together.
```
`Resource`, `Decision`, and `decides` stand for types registered in your graph.
Import `MergePlanApplyOptions`, `MergePlanReadContext`, and `MergePlanApplied`
from `@nicia-ai/typegraph/graph-merge` to type reusable helpers. Callbacks are
execution options: they are not stored in the artifact or covered by its digest.
The transaction acquires the schema fence before the graph write lock, validates
schema, revision origin, and revision, and then calls `beforeApply`. This context
exposes node and edge collection reads and, for identity-enabled graphs, identity
reads. Write methods, native SQL, and a root Store are absent. Application writes
before plan application are deliberately unsupported: an uncommitted write can
change the reviewed state without changing its durable revision yet.
After the precheck, TypeGraph performs the existing plan preflight, identity
checks, and writes. `afterApply` receives a `TransactionContext` whose reads see
those writes. Its `applied.merged` contains only the plan's provisional counts;
callback writes do not contribute to the final report's merge counts. Use the
supplied contexts for every graph operation. Calling the original Store inside a
callback does not enlist it in this transaction. Do not retain a context for later
work, and await every operation before returning.
Callbacks must resolve without a value. Throw or reject to abort; returning a
value, including an `Err`, is refused with `InvalidMergeOptionsError`. A callback
rejection, stale plan, merge failure, capture-flush failure, or commit failure
rolls back the combined graph operation. Errors are converted to the outer
`Result` after rollback; ordinary application errors are retained in the cause
chain. Existing typed merge errors and constraint translation remain intact.
Only the successful outer result confirms commit.
Transaction conflicts (PostgreSQL serialization failures or deadlocks) retry the
whole transaction up to **three attempts**, including both callbacks. Every
attempt checks the fence again; an intervening committed write makes the plan
stale rather than silently rebasing it. Keep callbacks safe to repeat. Do not send
messages, call external services with side effects, or publish an outcome inside
a callback. Perform those effects after successful completion, or write an
application outbox record through `tx` for later delivery. Returned contexts and
provisional outcomes are not durable notifications.
Protection covers the target graph's transactional state and participating
TypeGraph writers using its graph fence. It does not make an application policy a
declarative constraint: every writer changing that policy's state must enforce
it, for example through its own conditional operation. It does not cover other
graphs, arbitrary SQL, or external systems. SQLite uses its writer transaction;
Composed PostgreSQL applications use read-committed isolation and the graph
write lock with or without history. The lock statement records the effective
session isolation; incompatible or unknown isolation is refused before callbacks.
Standalone revision-tracking-only applications retain serializable isolation. Unsupported
transaction capabilities are refused before callbacks. Existing session-bound
fence and recorded-capture isolation checks still apply.
With history enabled, plan and application writes share the transaction's
recorded capture and flush, producing one per-graph recorded revision. Without
history, revision tracking likewise advances for the combined transaction.
Failure leaves no live changes or recorded revision from the failed attempt.
Existing valid-time bounds, including open bounds, retain their semantics.
Optional persisted merge provenance remains separate from recorded history:
provenance records are persisted only after successful graph commit, and a
persistence failure remains a report warning. Callbacks do not receive a
post-commit provenance result. Previously committed review records in the target
still invalidate a plan's revision fence; this API does not relax plan staleness.
For a durable candidate review, revalidate the stored review first, then pass
the compatible result's fresh `plan` and these callbacks to `applyMergePlan()`.
## Scaling branches and interchange
`revisionTracking: true` is the recommended mode for long-lived, repeatedly
branched graphs. It advances one durable revision anchor inside each successful
Store write transaction. The anchor combines a per-graph random origin with the
monotonic commit clock, so a branch can only match the store that created it —
not an independent database whose clock happens to share the same timestamp. A
branch and its merge precondition then read that constant-size anchor instead of
hashing every live node and edge. Stores created with `history: true` already
have the same guarantee through their recorded-time commit clock.
On PostgreSQL, the guarantee serializes writes to the same graph with a
transaction-scoped advisory lock. That is the correct trade-off for a live graph
whose branch merges must fail closed, but it can reduce throughput and increase
write latency for a high-concurrency, single-graph workload. Partition that
workload across graphs or leave revision tracking off when the content-fingerprint
fallback is acceptable.
Turning revision tracking off does **not** turn off all serialization.
*Constrained* writes now take the same per-graph mutual exclusion regardless of
`revisionTracking` or `history`, because their check-then-write is only sound if
nothing else writes the graph in between: edge cardinality (`one`, `unique`,
`oneActive`, and the `getOrCreateByEndpoints` create and resurrect legs),
node-kind disjointness on create, and a `kindWithSubClasses` uniqueness
constraint that actually expands to more than one kind — a scope covering a
single kind probes exactly the row the uniques table's primary key then
reserves, so that key is already its fence. Everything else — an unconstrained
create, a delete, a cardinality-`many` edge — pays nothing, so the cost is
proportional to the constraints you actually declared. On PostgreSQL that
exclusion is the same transaction-scoped advisory lock; on SQLite it is the
`BEGIN IMMEDIATE` writer slot the backend already takes. A backend running
without transactions (D1, `neon-http`, or `transactionMode: "none"`) has neither
and cannot be fenced.
This unlocks:
- Many concurrent agent, importer, or review branches without base-version
validation growing with the graph.
- Large graph copies, backup/export, and transfer pipelines that keep only one
interchange batch resident at a time via `exportGraphStream()` and
`importGraphStream()`.
- A safe fast path for a live base: a branch is rejected if any tracked base
write lands before its merge commits, rather than silently merging a stale
plan.
Streaming removes the graph-sized heap spike, but a physical working copy still
copies `O(graph)` rows and snapshot merge staging still compares branch state to
the base. Bundled backends page those comparisons across declared kinds, so
unused kinds do not each cost a database statement; custom backends without the
cross-kind read retain per-kind keyset pagination. Disposable candidate clones
also skip statistics refresh. Copy-on-write logical branches and delta-only
staging remain the next larger architectural step.
Revision tracking covers writes through the Store API. Direct backend writes and
raw graph-table writes through `tx.sql` bypass the anchor, so applications using
either escape hatch must avoid them for a branchable graph or retain the default
content-fingerprint validation. On transactional backends, streaming export holds
one read-only repeatable-read transaction across nodes, edges, and identity
assertions, so every chunk belongs to one committed snapshot. A snapshot stream
cannot be piped directly into a target that writes through the same serialized
connection: the same SQLite backend, distinct wrappers sharing one better-sqlite3
handle or one local (`file:`/`:memory:`) libSQL client, a bare `pg`/neon
`Client` (a checked-out `PoolClient` included), a `pg` `Pool` capped at one
connection (`{ max: 1 }`, and equally the uncoerced string forms `{ max: "1" }`
and the legacy `{ poolSize: "1" }` that `max: process.env.PG_MAX` produces), a
postgres-js client capped at one connection (`{ max: 1 }`, `?max=1` in the URL,
or `PGMAX=1`), distinct PGlite backend wrappers sharing one in-process
connection, or Cloudflare Durable Object storage, whose transaction frame is
ambient on the storage object — materialize it first or import it into an
independent backend. Pooled connections, HTTP drivers, remote libSQL, and
separate handles on one database are deliberately not treated as serialized:
each statement gets an independent connection there,
so refusing would refuse work that succeeds. The exclusion is one **exclusive** lease
per serialized connection, not a one-time check and not a cross-kind-only rule:
at most one long-lived interchange stream of any kind holds a given connection,
so all four pairings are refused — import behind export snapshot (even through a
user-wrapped stream that no longer identifies its source backend), export
snapshot behind streaming import, export behind export, and import behind
import. Whichever long-lived stream starts second gets a typed
`ConfigurationError` instead of both hanging; its `details.code` names the
condition holding the connection and `details.requested` / `details.heldBy` name
the pairing that was refused (see
[Interchange serialized-connection guard codes](/errors#interchange-serialized-connection-guard-codes)).
Every long-lived import claims that lease, not only the chunk-streaming one:
`importGraph` holds it for the whole call and `trustedImportGraph` /
`trustedImportGraphStream` for the whole trusted session, so those APIs can throw
this `ConfigurationError` too — new in 0.46 for trusted import, which previously
threw only `TrustedImportError`. TypeGraph's branch cloner detects
the shared-client case and materializes its snapshot before importing it.
Non-transactional backends can export identity-disabled graphs without this
snapshot guarantee. Identity-enabled stores already require a transactional
backend at construction, so every identity export has the snapshot guarantee.
## Entity resolution
Resolution is configured **per node kind** in `resolve`. A kind that is omitted
merges *by id only*: its new nodes and edges are copied through, but no fuzzy
matching runs. Each configured kind composes up to three candidate sources, all
feeding one shared scorer:
| Source | What it matches | Configured by |
| ------------ | ---------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- |
| Exact unique | Two staged nodes sharing all of a declared `unique` constraint's values — a *definitional* match that bypasses scoring | the graph's `unique` constraints |
| Blocking key | Cheap pre-grouping so similarity only compares plausibly-related nodes | `block` (staged) / `blockIndex` (vs. committed base) |
| Similarity | Fuzzy scoring of candidate pairs against a `threshold` | `similarity` + `threshold` |
```typescript
resolve: {
Patient: {
block: (node) => node.mrn ?? node.birthDate, // cheap candidate grouping
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.78, // pairs scoring >= 0.78 merge
},
}
```
### Blocking: `block` vs `blockIndex`
Blocking bounds the otherwise-`O(n²)` pairwise comparison by only comparing
nodes that share a cheap key.
- **`block(node) => string | undefined`** is an arbitrary function over staged
nodes — a normalized email, a tenant id, a birth date, a `soundex(name)`.
Returning `undefined` puts the node in the shared *unblocked* bucket.
- **`blockIndex`** names a declared `defineNodeIndex` and is the **new-vs-base**
block key: it lets the merge query *already-committed* nodes that share a
staged node's index key and propose them as candidates. It powers incremental
ingestion (see [Snapshot vs incremental](#snapshot-vs-incremental)) and is
ignored on the snapshot `merge()` path.
```typescript
import { defineNodeIndex } from "@nicia-ai/typegraph";
const patientCohort = defineNodeIndex(Patient, { name: "patient_cohort_idx", fields: ["cohort"] });
const graph = defineGraph({ /* ... */ indexes: [patientCohort] });
// In resolve, recall committed patients in the same cohort:
resolve: { Patient: { blockIndex: "patient_cohort_idx", similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.85 } }
```
### Keyless windows
A node with no block key and no unique signature lands in the *unblocked*
bucket, which is otherwise compared all-vs-all. For large unblocked sets, set
`keyless` to switch to bounded single-pass **sorted-neighbourhood**: nodes are
sorted by their similarity text and each is compared only to its next `window`
neighbours — `O(n·window)` instead of `O(n²)`, still fully deterministic.
```typescript
resolve: {
Article: {
similarity: { kind: "fulltext", fields: ["title"] },
threshold: 0.8,
keyless: { window: 20 }, // compare each unblocked article to its 20 nearest neighbours
},
}
```
### Similarity strategies
Four strategies cover the spectrum from zero-dependency to embedding-powered:
| Strategy | Needs embedder? | Use case |
| ---------- | --------------- | ---------------------------------------------------------------------------------------------------------------- |
| `fulltext` | No | Portable in-memory Sørensen–Dice trigram score over one or more fields (e.g. `name`). The cross-backend default. |
| `custom` | No | Your own deterministic `score(a, b) => number` — domain rules, weighted field blends, edit distance. |
| `vector` | Yes | Cosine similarity over one field's embedding. Catches semantic near-duplicates. |
| `hybrid` | Yes | Blend `vector` and `fulltext` by `weights` (default 0.5 / 0.5). |
The `fulltext` scorer runs **in memory** over the staged candidate text — it
deliberately does not consult database fulltext indexes, because branch
candidates are staged working-copy rows, not indexed search results. That keeps
scoring deterministic and identical across SQLite and Postgres.
For `vector` / `hybrid`, supply an `embedder` (batched, async, deterministic —
the same text must always map to the same vector):
```typescript
const result = await merge(base, branches, {
embedder: async (texts) => texts.map((text) => embedModel(text)), // text[] -> Float32Array[]
resolve: {
Article: {
similarity: { kind: "hybrid", fields: ["title", "summary"], weights: { vector: 0.7, fulltext: 0.3 } },
threshold: 0.84,
},
},
});
```
A `vector`/`hybrid` strategy with no embedder configured fails with a typed
`SimilarityUnavailableError`, never a silent no-op.
## Conflicts
When merged contributors disagree on a property value, Graph Merge **resolves by
an explicit, deterministic policy and records what it did** — it never lets
arrival order decide.
### Property conflicts
| Policy | Behavior |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `flag` (default) | Commit the deterministic survivor value (or the committed base value, for base-vs-branch) and record a `PropertyConflict` for review. The graph still gets a value; the disagreement is surfaced rather than resolved toward another branch. |
| `lastWriteWins` | Pick the value from the highest-priority branch (earliest in `branchOrder`) — *logical* order, never wall-clock. |
| `provenanceWeighted` | Pick the value from the highest-weight branch (see `provenanceWeights`). Ties fall back to branch order. |
| function | Delegate: `(conflict) => JsonValue` lets application code decide per conflict. |
There are **two** property-conflict knobs, deliberately separate so a fuzzy
branch match can never silently overwrite committed data:
- `onPropertyConflict` — staged branch vs. staged branch.
- `onBasePropertyConflict` — committed base vs. a branch (new-vs-base merges).
Defaults to `flag` independently, and does **not** inherit `onPropertyConflict`.
`provenanceWeighted` reads per-branch trust weights you supply:
```typescript
const result = await merge(base, branches, {
onPropertyConflict: "provenanceWeighted",
provenanceWeights: new Map([
[authoritativeFeed.id, 1.0], // the system of record wins ties of value
[bestEffortAgent.id, 0.2],
]),
});
```
### Delete / modify conflicts
An inherited node or edge that one branch **deletes** while another **modifies**
is neither a pure delete nor a pure modify. `onDeleteModifyConflict` governs it
for both nodes and edges:
| Policy | Behavior |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `flag` (default) | The modification survives **and** an unresolved `DeleteModifyConflict` is recorded — a merge must never silently destroy the only branch still carrying data. |
| `deleteWins` | Honor the delete; discard the modification; record the conflict. |
| `modifyWins` | Resurrect the row; keep the modification; record the conflict. |
Independent edits to the *same* inherited row by different branches are
**three-way merged against the base**: a field only one branch changed takes
that change with no conflict; only fields multiple branches changed to differing
values become conflicts. This holds for node *and* edge properties, so disjoint
edits compose instead of clobbering each other.
## Edges follow their entities
When nodes collapse, their edges must too. After clustering, Graph Merge:
1. **Repoints** every edge endpoint onto its cluster's canonical survivor.
2. **Drops** any edge whose endpoint was finally deleted (recorded in `dropped`).
3. **Dedupes** edges that repointing brought together, as a pure set keyed by
`(from, type, to, props)` — so `x → a` and `x → b` both landing on `x → c*`
collapse to one edge.
4. **Reconciles** edges that collapse that way but disagree on properties, via
the same conflict policy as nodes — over the properties each side actually
*changed*, so an inherited row's untouched value never competes with (or
outvotes) a value some branch authored.
Steps 3 and 4 are scoped to collisions **repointing caused**: edges are grouped by
the endpoint pair they named *before* repointing, and one row per pair collapses. A
TypeGraph store is a multigraph — nothing enforces uniqueness on `(from, kind, to)`,
`create()` makes a parallel edge, and `getOrCreateByEndpoints()` is the opt-in
set-semantics accessor — so a branch that created a parallel edge merges as a
parallel edge, and a window claim lands on the row its author touched. A repointed
edge landing on endpoints that already have several parallel rows merges into one of
them; the rest keep their own properties and windows. What makes two staged edges
"the same row" is their **edge id**, not equal properties: one inherited edge
staged by several branches folds into a single write, while a branch-created edge
is a new row even when its properties happen to match an existing one's.
When such a collapse mixes an **inherited** edge with a branch-created one, the
inherited row is the one kept: a collapse rewrites the row it keeps and does not end
the rows folded into it, so writing onto the row the target already holds is what
keeps a committed edge from being left beside the row that replaced it. This mirrors
the node rule below, and it is also what the surviving edge id in `PropertyConflict`,
window resolutions, and provenance names. A collapse of branch-created edges alone
keeps the lexicographically-minimal edge id.
Inherited edges that a branch **deleted** are removed from the target, and
inherited edges **modified** by multiple branches go through the same base-aware
three-way merge as nodes — so an edge's `since` edited by one branch and `note`
edited by another keep *both* edits. The collapse in step 4 is base-aware for the
same reason: a staged copy of an inherited row carries that row's whole property
bag, and only the values it *changed* count as claims. The clearest case is a row
staged solely to carry an end-of-validity — it authored no property, so it
contributes no claim and raises no conflict, whatever its branch's rank.
## Ontology type reconciliation
With `reconcileTypes: "ontology"`, two staged nodes that share an id but carry
subtype-compatible kinds (via the graph's `subClassOf` closure) are collapsed to
the **most-specific** common type, recorded as a `TypeReconciliation`. A base
`Doctor` and a branch `SpecialistDoctor` reconcile to `SpecialistDoctor` instead
of being dropped as incompatible. The default `"off"` keeps identity strictly
`(kind, id)`.
```typescript
const graph = defineGraph({ /* ... */ ontology: [subClassOf(SpecialistDoctor, Doctor)] });
const result = await merge(base, branches, { reconcileTypes: "ontology" });
```
## Choosing the survivor
By default a cluster's canonical survivor is the member with the
lexicographically-minimal id. A committed member always wins instead, so its
committed identity and the edges already attached to it stay stable: on
new-vs-base merges that is a committed base member, and on incremental merges
it is also a node the live target committed after the fork point, such as one
an earlier branch's merge added. Override the staged-vs-staged choice with
`canonical`:
```typescript
const result = await merge(base, branches, {
canonical: (cluster) => preferGoldenSource(cluster.members), // pick which id survives
});
```
## Scaling & safety
Two guards keep a merge bounded and predictable on large or pathological inputs:
- **`maxComparisonsPerKind`** caps fuzzy comparisons per kind. On overflow,
`onComparisonCeiling` decides: `"error"` (default) fails with a typed error,
or `"mergeByIdOnly"` skips similarity for that kind (still honoring exact
unique matches) and records a warning. Tighten your `block` to shrink buckets
rather than raising the ceiling blindly.
- **`clusterMaxDiameter`** optionally splits over-broad clusters: if a cluster's
single-link diameter exceeds the bound, the weakest edges are dropped
deterministically until every sub-cluster fits. This stops a chain of
near-matches (`a~b~c~…`) from fusing genuinely-distinct entities.
```typescript
const result = await merge(base, branches, {
maxComparisonsPerKind: 50_000,
onComparisonCeiling: "mergeByIdOnly",
clusterMaxDiameter: 2,
});
```
## The merge report
`merge()` returns `Result`. The report is the
**application boundary** — show conflicts to an operator, write a review record,
persist provenance, or feed a downstream step.
```typescript
type MergeReport = {
merged: {
nodes: number;
edges: number;
identity: { asserted: number; retracted: number }; // ledger effects
};
resolutions: EntityResolution[]; // collapse membership + decisive match evidence
conflicts: PropertyConflict[]; // per-property disagreements + how they resolved
deleteModifyConflicts: DeleteModifyConflict[]; // node/edge delete-vs-modify cases
typeReconciliations: TypeReconciliation[]; // ontology kind collapses
// Node drops (deleted endpoints, incompatible members), edge drops, identity
// drops (identity:duplicate-assertion, identity:endpoints-collapsed,
// identity:retraction-target-mismatch, identity:deletion-overruled), and
// lower-bound deltas the commit cannot apply (window-not-applicable)
dropped: DroppedItem[];
// Inherited rows whose end-of-validity the merge resolved. Each entry carries
// validTo for a set/move or clearValidTo: true for a reopening.
validityEnds: ValidityEndResolution[];
baseAmbiguities: BaseAmbiguity[]; // new-vs-base matches that spanned >= 2 committed entities
provenance: ProvenanceIndex; // byBranch(id) -> { nodeIds, edgeIds }
warnings: string[]; // non-fatal advisories (ceiling skips, provenance-persist failures)
candidateDiagnostics?: CandidateDiagnostics; // bounded, opt-in scored comparisons
provenancePersisted?: { graphId: string; count: number }; // when persistProvenance ran
};
```
A typical operator loop: auto-apply when `conflicts` and
`deleteModifyConflicts` are empty; otherwise enqueue them for review alongside
`resolutions` so the reviewer sees what merged and why.
### Why two entities matched
Every multi-member `EntityResolution` has `decisiveEdges`: a deterministic
minimal connectivity witness. A resolution over N distinct `(kind, id)`
identities normally has N−1 edges. Endpoints retain both kind and id, so
same-id nodes of different kinds remain distinguishable during ontology
reconciliation.
A same-id ontology retype remains a `TypeReconciliation`, rather than creating
an id-merge resolution. Its optional `decisiveEdges` carries the accepted retype
witness without changing the meaning of the existing resolution collection.
Each edge records every candidate source that proposed the pair in stable order.
Definitional evidence names the trusted rule, such as a unique constraint, and
does not pretend the internal forced match was a perfect similarity score.
Scored evidence records the strategy descriptor, actual score, and threshold
used by the shared scorer:
```typescript
type MatchEvidence =
| {
a: { kind: string; id: string };
b: { kind: string; id: string };
sources: MatchSource[];
decision: "definitional";
}
| {
a: { kind: string; id: string };
b: { kind: string; id: string };
sources: MatchSource[];
decision: "scored";
strategy: MatchStrategy;
score: number;
threshold: number;
};
```
Built-in source metadata distinguishes block, unique, base-unique, base-index,
keyless, and ontology-retype proposals. Several sources proposing the same pair
are all retained after deduplication. Strategy metadata describes `fulltext`,
`vector`, `hybrid`, or `custom` configuration, never custom function source.
Default evidence excludes the raw compared values and rejected pairs because
those may contain PII and can make reports enormous.
Candidate diagnostics are explicit and bounded:
```typescript
const planned = await planMerge(base, branches, {
...options,
candidateDiagnostics: { limit: 1_000 },
});
```
When enabled, the report and reviewable plan include accepted and rejected
scored comparisons in canonical order. A definitional edge removed by the base
ambiguity or diameter guard is also retained with its exclusion reason, so the
final partition remains explainable. The collection also carries `total`,
`limit`, and `truncated`.
The limit is deterministic: the same candidate set produces the same retained
prefix regardless of branch, source, or backend enumeration order. Diagnostics
still omit raw compared values; join their `(kind, id)` references to application
data only in an appropriately protected evaluation environment.
## Provenance
Provenance answers *which branch contributed each merged node and edge*. A
contribution is anything a branch authored into the committed row — the
properties it staged, the modification that survived, or the end-of-validity the
merge applied.
- **Report-only (default, `provenance: true`)** — `report.provenance.byBranch(id)`
returns the `{ nodeIds, edgeIds }` that branch contributed. In-memory; it
evaporates after the call.
- **Durable (`persistProvenance: true`)** — one `{branch, sourceId} → canonical`
row per contribution is upserted into a *sidecar* graph on the target's
backend (its own namespaced tables; your domain schema is untouched). The
sidecar is opened and claimed **before** the merge commits, so a sidecar graph
id TypeGraph cannot claim refuses the whole merge and leaves the target
unmodified; only the row write itself is post-commit and best-effort, where a
transient failure surfaces as a `warnings` entry rather than a failed merge.
Re-running the same merge upserts (deterministic ids), never duplicates.
`openProvenanceStore` only ever opens a sidecar graph id it can prove it owns,
and ownership is **marker-first**: a durable `ProvenanceOwner` marker row is the
sidecar's first write of any kind, committed inside the schema fence *before*
the sidecar schema is registered. A never-seen id is free to claim only when it
holds no row in **any** per-graph table — nodes and edges, but equally
recorded-time history, the revision clock and origins, identity assertions and
their derived closure and separation, fulltext, and unique keys — because a
plain `createStore` writes rows without registering a schema, so an unregistered
id is not by itself evidence of a free namespace. Ownership is then the marker
alone, checked independently of the schema hash, because an application is free
to define the same `Provenance` shape at an unrelated id. Because the marker
comes first, the resumable interrupted state is **marker without schema** (or a
marker beside a pre-marker legacy schema): that resumes by registering or
migrating the schema. The opposite state — the exact current sidecar schema with
no marker — is one TypeGraph cannot produce, and is refused unconditionally
whatever the graph contains, empty and provenance-shaped included, since
contents an application could have written are not evidence of authorship.
**What a claim costs, on PostgreSQL.** One writer class takes neither the
per-graph fence nor the graph's active schema row: a schema-less raw
`createStore` writer, or a direct `backend.insertNode` / `insertEdge` call. At
READ COMMITTED its insert could commit between the claim's re-inspection and the
claim's own commit, leaving the marker on an id an application had just made its
own. To close that, the claim issues
`LOCK TABLE , IN SHARE ROW EXCLUSIVE MODE` inside the fence and
before the re-inspection. That mode excludes every `INSERT` / `UPDATE` /
`DELETE` on those two tables **for every graph on the database** — they are
shared tables — while still admitting readers. So while a claim runs, every node
and edge write database-wide waits.
The bound is what makes it acceptable: the lock is taken **only inside a claim**,
which happens when a sidecar is created, upgraded from the pre-marker schema, or
resumed after a crash — never on the common path, where an already-owned sidecar
opens with no fence at all. Its duration is the re-inspection's probes plus one
`INSERT`, with no caller code and no caller I/O inside it. The mode is
`SHARE ROW EXCLUSIVE` rather than plain `SHARE` because it must be
self-exclusive: two concurrent claims on different sidecar ids hold different
advisory locks, so under `SHARE` both would acquire it and then both request
`ROW EXCLUSIVE` for their own marker insert — a lock-upgrade deadlock PostgreSQL
resolves by aborting one of them. SQLite takes no such lock; `BEGIN IMMEDIATE`
already owns the engine's single writer slot.
Refusals carry the code `GRAPH_MERGE_PROVENANCE_ID_COLLISION` and one of five
`details.reason` values — `application-graph`, `empty-legacy-sidecar`,
`unupgradeable-legacy-sidecar`, `unowned-exact-schema-graph`, or
`corrupt-ownership-marker` — so the remediation matches what is actually there
instead of generic advice; a backend with no transactional schema fence refuses
an unclaimed sidecar with `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED` (an
already-owned sidecar still opens there). Under `persistProvenance: true` both
of those arrive as a typed `InvalidMergeOptionsError` naming
`details.option: "persistProvenance"`, with the originating `ConfigurationError`
as its `cause` — see
[Merge provenance sidecar codes](/errors#merge-provenance-sidecar-codes).
Query persisted provenance back later:
```typescript
import { openProvenanceStore, readProvenance } from "@nicia-ai/typegraph/graph-merge";
const store = await openProvenanceStore(target);
const fromAgentA = await readProvenance(store, { branchId: "agent-a" }); // what did agent A contribute?
const whoMadeX = await readProvenance(store, { canonicalId: "patient-123" }); // who contributed node X?
```
Inspection tools that have a backend and graph id but not the target's
`GraphDef` can use the standalone overload:
```typescript
const store = await openProvenanceStore(backend, targetGraphId);
```
## Snapshot vs incremental
A branch is forked from a `base@V` — a token combining the base's schema hash
with the store's durable revision anchor when `revisionTracking: true` or
`history: true` is on, or a complete live-content fingerprint otherwise.
The revision anchor is namespaced by a durable per-graph origin, which
`Store.clear()` rotates. A lineage-capable untracked store whose backend
supports that origin relation also carries it beside its content fingerprint.
The two merge entry points differ in how they
treat that token.
The token is printable text, so it can be stored anywhere an application
keeps descriptors, plans, and fork points, including PostgreSQL `text` and
`jsonb` columns. Treat it as opaque: compare it whole and never parse it.
Tokens minted by releases before this format, which separated components
with a NUL character, are refused with a `BaseVersionMismatchError` whose
`details.reason` is `"legacy-token-format"`. Re-branch or re-plan from the
current target. Earlier `engine:` anchors and untracked content tokens without
the active schema version also require re-branching; they cannot match the
current target's token.
The token is printable text, so it can be stored anywhere an application
keeps descriptors, plans, and fork points, including PostgreSQL `text` and
`jsonb` columns. Treat it as opaque: compare it whole and never parse it.
Tokens minted by releases before this format, which separated components
with a NUL character, are refused with a `BaseVersionMismatchError` whose
`details.reason` is `"legacy-token-format"`. Re-branch or re-plan from the
current target.
**`merge()` is a snapshot merge.** Every branch must have forked from the
target's *current* `base@V`. If the target advanced since the branch was taken,
`merge()` returns a `BaseVersionMismatchError` rather than risk clobbering newer
data. This is the right model for "fork, do work, merge back" within one round.
**`mergeIncremental()` is a fork-point merge into a live target.** It merges
branches that forked from a frozen `forkPoint` into a `target` that may have
*moved on*. Additions are re-discovered against already-committed entities (via
`blockIndex` / unique constraints) so a re-seen entity updates the committed row
instead of duplicating it. Inherited node and edge modifications/deletions are
also propagated through the same three-way planner, with the live target kept
authoritative when it changed concurrently.
```typescript
import { mergeIncremental } from "@nicia-ai/typegraph/graph-merge";
const result = await mergeIncremental({
forkPoint, // the frozen ancestor the branches forked from
target, // the live committed graph (may have advanced)
branches,
options: {
resolve: { Patient: { blockIndex: "patient_cohort_idx", similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.85 } },
onBasePropertyConflict: "flag", // required: never overwrite a newer committed value
},
});
```
`mergeIncremental()` requires `onBasePropertyConflict: "flag"` — any other value
is refused with `InvalidMergeOptionsError` — so a stale branch value can never
overwrite a newer committed value during new-vs-base recall.
The `forkPoint` must stay **frozen for the duration of the call**: every branch
diff is computed against it, and the commit transaction re-reads its `base@V`
before applying anything, so a write landing on the fork point mid-merge is
refused with `BaseVersionMismatchError` instead of committing diffs against an
ancestor that no longer exists. Only the `target` may advance while the merge
runs.
If both the branch and the live target changed the same inherited row, the target
value/deletion wins and the conflict is reported. Both `merge()` and
`mergeIncremental()` commit **transactionally** and require a
transaction-capable target backend. Managed targets also acquire the
schema-version write fence; raw targets remain outside schema fencing. On
PostgreSQL, serialization failures from either the target-content guard or the
schema fence are retried automatically around the complete commit.
### Lineage and pruned diffs
A backend may declare a `lineage` capability: an opaque, whole-database
`revision(session)` it can report and compare, plus `changesSince(session,
revision, graphId)`, which names every node and edge of one graph that
changed (inserted, updated, deleted, or resurrected) after that revision — or
admits `{ kind: "unbounded" }` when it cannot bound the answer. Bundled SQLite
and PostgreSQL stores with `revisionTracking: true` also provide bounded
lineage through a DML journal when history capture is disabled. `lineageRevisionNow()`
mints a public anchor and `changesSince(anchor)` returns changed node and edge
keys. Use that anchor API rather than `revisionNow()`, which returns a clock
value without the graph's origin identity. The journal is installed when the
store is provisioned through `createStoreWithSchema()`, or explicitly with
`installRevisionChangesJournal(backend)` from `@nicia-ai/typegraph/schema`
under a schema owner role. Existing installations must first adopt base schema
version 4 through a privileged schema open or generated base-schema migration.
Runtime lineage checks the journal and its triggers without issuing DDL; a
revision-tracked store without history fails with `REVISION_JOURNAL_NOT_READY`
when the journal is not ready. Short-lived clones that do not need this bounded
lineage can set `revisionJournal: false`. Writes before the first anchor are
outside that anchor's range.
Node and edge inserts, updates, and deletes are recorded by database triggers.
Identity-only revisions and revisions whose write provenance is incomplete
produce `{ kind: "unbounded" }` rather than an incomplete key list. Custom
backends must provide their own lineage capability to get bounded results.
Each trigger is attached to a whole physical node, edge, or identity table; it
records every write to that table and uses `graph_id` to identify the affected
graph. On shared tables this captures writes from every graph, not only graphs
whose stores enabled the journal. Journal rows are retained per revision and
never cleaned up automatically; applications should avoid installing triggers
on shared tables unless cross-graph capture is intended, and should plan an
external retention policy that preserves every revision still used as a branch
anchor. `resolveLineage(store)`
selects backend lineage first, then captured history, then the first-party
revision journal. A lineage source is consulted only to avoid rework; it never
changes what a merge decides.
`revision()` reports `:`, never the bare clock value alone:
the durable, random per-graph revision-origin nonce
(`typegraph_revision_origins`) plus the recorded-time clock. Two
independently created stores that share a `graphId`, or the SAME store
across a `Store.clear()` boundary, can mint numerically comparable clock
values, and the origin is what keeps `changesSince` from mistaking one for
the other — a revision whose origin no longer matches the graph's LIVE
origin row is `unbounded`, regardless of what its numeric clock value is.
The recorded-relations derivation's delta is trustworthy only when EVERY
writer to the graph goes through a store that captures history — a precondition
it can partially, but not fully, enforce itself. `changesSince` proves
completeness directly rather than inferring it from a high-water mark: every
integer revision between the requested one and the graph's current clock
must carry direct evidence — a `recorded_from` or a non-sentinel
`recorded_to` — in one of the three recorded relations (nodes, edges,
identity assertions). This catches an incomplete record wherever the hole
falls, including a `revisionTracking`-only `Store` (no `history`) that
advanced the shared clock without inserting a row and was later FOLLOWED by
a capturing commit — a later capturing commit cannot retroactively supply
the missing evidence, so the gap is caught regardless of what comes after
it. What it CANNOT detect: a non-capturing writer bypassing every `Store`
entirely (a raw `GraphBackend` write, or an engine-side mutation outside
TypeGraph), which leaves no evidence to be short of. Route every writer
through a capturing `Store` if a `"keys"` delta from this source must be
exhaustive.
`session` is the connection the caller's decision is bound to — a
session-less bag could never be pinned to anything, so this one always
carries one. A caller planning outside any transaction (`branch()`'s
fork-revision capture, the pruning below) passes the root backend it holds;
a caller re-validating a content fingerprint inside an open commit transaction
reads through that transaction's own handle, so the fingerprint observes the
transaction's snapshot and establishes dependencies on the rows it covers.
**Untracked stores use a complete fingerprint.** An engine-wide revision and
node/edge-only `changesSince` result cannot fence an identity-only write. It
also cannot establish read dependencies on the graph state used in planning.
For this reason, a store without TypeGraph revision tracking fingerprints live
nodes, edges, and current identity assertions even if its backend exposes
`lineage`. Where supported, the token also carries the durable graph origin.
The commit transaction checks the origin and recomputes the fingerprint before
applying its writes. Previously minted `engine:` base tokens are retired; re-branch
from the current store rather than applying an old merge.
**`Store.clear()` rotates the revision origin.** For revision-tracked stores,
`clear()` deletes and re-mints the per-graph origin in the same transaction.
A branch forked before that clear cannot merge into the post-clear store even
when its revision clock has the same numeric value.
The origin row is also read fresh on every mint (`computeBaseVersion`,
`Store.revisionOriginNow()`), never cached on a `Store` instance. Two live
`Store` objects can legitimately observe the same graph — nothing requires
that only one `Store` ever exists per database — and only one of them runs
`clear()` at a time; a stale per-instance cache on the other would keep
minting anchors from the origin that existed before the clear, so a branch
it forks would fail every merge at commit until that `Store` happened to be
recreated. Reading fresh means a second `Store` over a graph another `Store`
just cleared sees the rotation immediately, with nothing to recreate.
**Pruning the diff.** `branch()` also records a `forkRevision` on the
returned `GraphBranch` — the fork's own `lineage.revision(session)`, read
right after the working copy is created and before any write reaches it,
with the working copy's own root backend as the session (this runs strictly
outside any transaction). For the recorded-relations source this is
origin-bearing like any other reading, so clearing and repopulating the
FORK itself to the same revision count `forkRevision` held is caught the
same way a cleared BASE store already is — there is no separate guard for
the fork side to add, because the token itself now carries the check. When
staging a branch for merge, its diff against the base is restricted to the
union of two deltas: what changed on the *fork* since `forkRevision`, and
what changed on the *base* since the anchor in its own `base@V` — instead of
enumerating every live row on both sides. A key absent from both deltas
cannot have changed since the fork point, so narrowing the read to their
union cannot miss anything the full diff would have found; it only fetches
fewer rows to compare. Pruning is a pure optimization with one rule:
whenever either side cannot supply a bounded delta, the merge falls back to
comparing every live row, exactly as it always has. That covers no
`forkRevision` (a hand-built branch, or one whose store resolved no
`lineage`); either side's `changesSince` answering `unbounded` or
REJECTING (a transient engine error never fails a merge the full diff would
have completed); and the base's own anchor failing to resolve against the
base store's lineage at all — an origin mismatch between a revision-anchored
`base` and the base store's live revision row, a revision anchor minted
before the base store's first tracked write, or an old engine anchor that
must be re-branched. Nothing about *what* a merge decides depends on
whether its diff was pruned.
## Working copies
`branch()` is backend-agnostic. The default `cloneWorkingCopyStrategy` exports
the base through TypeGraph's interchange and imports it into a fresh store on a
backend your factory provides — so it works identically across SQLite, Postgres,
and in-process PGlite, and needs no schema changes. The import is
fidelity-preserving: undeclared properties that `validateStore()` treats as
healthy semi-structured data are carried through. Stripping them would make a
later merge invent deletions against the original base.
```typescript
// Each branch gets its own in-memory SQLite backend:
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
const makeBackend = async () => createLocalSqliteBackend().backend;
const fork = unwrap(await branch(base, makeBackend, { id: asBranchId("worker-1") }));
```
For a custom isolation mechanism (e.g. a future copy-on-write namespace), pass a
`WorkingCopyStrategy` as the fourth argument to `branch()` — its single `create`
method receives the base store and the `BaseVersion` `branch()` already
stamped off it, and returns an independently-mutable store over the same
graph definition.
**A branch is a data fork.** `branch()` records the clone's committed schema
`(version, hash)` at fork time, and the merge refuses (typed, as
`BaseVersionMismatchError`) any branch whose store ran a schema operation
afterwards — `evolve()`, `migrateSchema()`, or `removeKinds()` — even a
round-trip migration that restores the original document hash. Those
operations mutate rows through their own preflights, and projecting the side
effects into a merge would detach them from the schema change that caused
them. Apply schema changes to the target first (or re-fork), then merge.
### PostgreSQL table-backed working copies
`createPostgresWorkingCopyManager` allocates a private set of TypeGraph tables
in the source PostgreSQL database. It derives the table inventory and base
schema marker from TypeGraph's PostgreSQL schema contributions, copies the
source graph with fenced `INSERT ... SELECT` statements, and records ownership
in `typegraph_working_copy_allocations`. The control backend, source backend,
and backends returned by `connect` must all reach the same database, and
`control` and `connect` must run as the same role
([One database role](#one-database-role)). TypeGraph checks the allocation's
private ownership token through each connection.
The control backend must execute DDL inside its PostgreSQL transactions;
its root `executeDdl` port is not required.
```typescript
import { drizzle } from "drizzle-orm/node-postgres";
import {
createPostgresBackend,
createPostgresTables,
} from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createPostgresWorkingCopyManager } from "@nicia-ai/typegraph/adapters/drizzle/postgres/working-copy";
import {
asBranchId,
branchDurable,
destroyDurableBranch,
reopenDurableBranch,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const control = createPostgresBackend(drizzle(pool));
const copies = createPostgresWorkingCopyManager({
control,
connect: (names, allocation) =>
Promise.resolve(
createPostgresBackend(drizzle(pool), {
tables: createPostgresTables(names),
...(allocation === undefined ?
{}
: { vector: allocation.vectorStrategy }),
}),
),
});
const { branch: copy, descriptor } = unwrap(
await branchDurable(sourceStore, copies.durable, {
id: asBranchId("candidate-42"),
allocationId: "candidate-allocation-42",
}),
);
await copy.close(); // Releases the connection; the tables remain.
const reopened = unwrap(
await reopenDurableBranch(graph, descriptor, copies.durable),
);
await reopened.close();
unwrap(await destroyDurableBranch(descriptor, copies.durable));
```
The same manager exposes `ephemeral` for `branch()`; closing that branch drops
its tables. `listUnsealedAllocations({ after, limit })` pages through durable
allocations awaiting seal and ephemeral allocations. These rows may still have
active owners; the ledger alone cannot identify a crashed process. After
confirming that no live branch or allocation uses a row, call
`abortAllocation(id)` to remove it. A durable branch's descriptor contains only
the allocation ID, not connection credentials. Pass `sourceTableNames` when the source backend uses
custom status table names; pass `reopenOptions` to restore process-local hooks
or query options on a later process. An external `recordedRead` binding is
refused because its relation is outside the owned table inventory. Reopen
options cannot replace the allocation's schema, recorded-read binding,
history mode, or revision-tracking mode. Pass `operations` to let
`durable.operations` commit host mutations atomically with immutable evidence;
see [Atomic operations and immutable evidence](#atomic-operations-and-immutable-evidence).
Each durable allocation also owns an evidence table (`op_evidence`) in
the allocation's schema, addressed through that schema rather than the
connection's `search_path`; destroy refuses to drop it while undelivered
evidence remains.
The source backend and every backend returned by `connect` must expose the
complete PostgreSQL `tableNames` inventory, including history, identity, and
status relations. The manager refuses missing or mismatched bindings with a
`BranchError` before cloning or opening a Store. For `ephemeral` and `durable`,
`connect` runs after the allocation tables are created, so custom callbacks may
inspect those tables; on binding failure, the manager removes the new tables and
ledger row. `makeBackend` connects earlier, before it provisions anything.
The table-backed strategy supports bundled tsvector fulltext, declared
PostgreSQL B-tree, GIN, and trigram graph indexes, and pgvector sidecars. It
builds each declared graph index on private tables under stable
allocation-scoped physical names while keeping logical index names and schema
hashes unchanged. `materializeIndexes()` can retry or repair indexes after
reopen; destroy removes their owned tables and indexes. When `connect` receives
an allocation vector strategy, pass it to `createPostgresBackend`; the strategy
assigns stable table and index names from the ledger-reserved physical prefix.
Allocation claims and all initial table and vector DDL commit together, so a
colliding or failed provision leaves no partly owned sidecars. Source vector sidecars
are copied under the same transaction locks as TypeGraph relations. The ledger
stores every relation name declared by each slot's `ownedTables()` contribution,
so destroy can remove them in reverse declaration order without a graph object.
Reopening requires the graph's vector slots and owned-relation inventory to
match the persisted allocation manifest. Older ledger rows that stored only
`tableName()` remain readable as single-relation slots. A declared vector slot
whose source sidecar is absent is refused because its contents cannot be
snapshotted exactly.
The `ephemeral` and `durable` copies have a fixed schema: `evolve`, kind
removal, and deprecation refuse before mutation. Use `makeBackend`, below, when
the working copy's schema must change. Custom fulltext strategies still need a host-level database
fork. The source and every copy connection, including durable reopen, must use
the bundled `tsvectorStrategy`: a custom strategy may own additional physical
tables whose rows cannot be copied safely from the generic contribution
inventory. A connection with fulltext disabled is refused for the same reason.
System index maintenance remains available. Source table locks cover the
entire TypeGraph relation set and vector sidecars while the SQL clone runs, so a
large clone briefly blocks writes to other graphs in the same database.
#### One database role
The manager supports one deployment shape: the `control` backend and every
session `connect` returns run as the **same PostgreSQL role**. TypeGraph reads
`current_user` on both sessions and refuses a difference with a
`ConfigurationError` whose `details.code` is `WORKING_COPY_ROLE_MISMATCH`, and
the refused allocation is not left behind.
The reason is ownership. A `control` session provisions and removes every
allocation, but the Store that opens on a connected backend issues its own DDL:
runtime-contribution markers, the revision journal and its triggers, system and
declared indexes, and vector tables an evolved graph introduces. Only a table's
owner (or a member of the owning role, or a superuser) can drop it, and the
comparison is by role name, so a `connect` role that is merely a member of
`control`'s role is refused rather than trusted. A different role would leave
the tables it creates behind on close and `abortAllocation`. The shared role
therefore needs `CREATE` on the schema.
`makeBackend` calls `connect` before it writes the ledger row or any DDL and
refuses a mismatch there, so nothing is allocated. `ephemeral` and `durable`
call `connect` after their allocation tables exist, so they refuse right after
it, before cloning or opening a Store, and remove the new allocation; a durable
reopen refuses the same way and leaves the sealed allocation untouched.
#### One schema per allocation
Every allocation lives in one schema: the `control` session's current schema when
the allocation is made, recorded in the ledger's `schema_name` column. No
`search_path` decides where an allocation's relations are created or dropped, so
a `connect` pool whose connections lead with different schemas cannot strand
tables that removal never finds.
- **Provisioning** fixes its transaction's search path to that schema before it
claims the ledger row, so the tables it creates land there whichever pooled
connection runs it, and the claim records the schema the statement itself
observed.
- **The connected backend** receives table names that carry the schema. A
backend built with `createPostgresTables(names)` over that object runs the DDL
it issues lazily (bundled tables a Store ensures on first use, fulltext and
contribution storage, schema-write transactions) with the schema leading its
search path, and the allocation's pgvector strategy names its tables and
indexes through the schema. `CREATE INDEX CONCURRENTLY` cannot run in a
transaction; it creates the index in the schema of the table it names, which
is already the allocation's. The backend's catalog probes (table, index, and
column lookups, including the recorded-time compatibility check a
`history: true` Store runs) read the allocation's schema, not the session's
current one. Extensions are database-global and create no
allocation relation, but their DDL still runs through the same DDL runner
wherever the write fence takes no lock: there, a backend built over a caller's
own transaction is subject to the same session check as any other lazy DDL
(below). Under a lock fence, a pooled backend installs the extension in its own
transaction, as before; a backend built over a caller's own transaction runs it
as a savepoint inside that transaction and makes no session check, because the
extension creates no allocation relation.
- **Refusals.** A connection whose backend was built over a *copy* of `names`
(which carries no schema) is refused with a `BranchError`. A `connect` driver
that cannot hold an interactive transaction (`drizzle-orm/neon-http`) is
refused with a `ConfigurationError`
(`ALLOCATION_SCHEMA_REQUIRES_INTERACTIVE_TRANSACTIONS`), because it cannot run
its DDL under a fixed schema. A backend built over a caller's own transaction
runs its lazy DDL and schema writes, and adopts that transaction for a schema
write, only when that session's current schema is the allocation's; otherwise
it is refused with a `ConfigurationError`
(`ALLOCATION_SCHEMA_SESSION_MISMATCH`). The caller owns that session's search
path, so it is checked rather than rewritten.
- **Removal** (`close`, `abort`, `destroy`, `abortAllocation`) searches the
catalog across every schema for relations named with the allocation's reserved
prefixes. It drops those in the recorded schema, schema-qualified in one
statement, and deletes the ledger row in the same transaction. If a drop fails
(a view that depends on an allocation table, for example) the transaction rolls
back, the row stays, and the allocation remains in `listUnsealedAllocations()`
for `abortAllocation()` once the dependency is gone. If any such relation sits
in a different schema, removal refuses with a `BranchError` that names the
schemas found and keeps the row, because deleting the row would discard the
only pointer to them. Three cases are worded differently. When the recorded
schema holds none of them, the schema was renamed or the tables moved (the
message says the relations are "not in its schema"; move the tables back or
correct the row's `schema_name` and remove again). When every relation found
elsewhere has a same-named relation in the recorded schema, it is a stale copy
left in another schema, such as a backup or restore schema (the message says
the allocation "also has relations" there; drop the copy and remove again,
since the copy blocks removal until it is gone). When some relations moved and
others stayed, for example one table moved to a backup schema while the rest
remain, the allocation is split and the relations elsewhere may be the only
copy (the message says the allocation "is split across schemas"; the
suggestion drops nothing, so move the relations back or correct the row's
`schema_name`). `details` carries `allocationId`, `schema`, `foundIn`, and
`schemas`, and `suggestion` names the recovery step. If the
allocation's relations exist nowhere (its tables were dropped entirely) there
is nothing to recover, and removal deletes the ledger row, so a crashed owner's
allocation cannot stay listed forever.
- **Ledger rows from before the schema was recorded** (written by 0.72.0) carry
no schema. They resolve through the session that removes them and reopen
without binding, and follow the same removal rule: relations found in a schema
other than the removing session's refuse removal and name that schema. `control` adds the column to an
existing ledger the first time it runs.
The connection must still be able to *resolve* the allocation's tables, so its
`search_path` must include the schema, typically `public`. A per-role `"$user"`
schema ahead of it is fine. The ledger itself lives where `control`'s session
creates it, so run `control` with one consistent `search_path`.
#### `makeBackend` for branches, candidate planning, and evolution previews
`copies.makeBackend` is a `MakeBackend`, so PostgreSQL callers no longer
hand-roll table prefixes, DDL, and cleanup. It fits every API that takes one:
`branch`, `ingestionBranch`, `planCandidateWriteSet`,
`planCandidateWriteSetReview` (including sparse staging), `branchForEvolution`,
and `planCandidateWriteSetForEvolution`.
```typescript
import { branch, branchForEvolution } from "@nicia-ai/typegraph/graph-merge";
const fork = unwrap(await branch(sourceStore, copies.makeBackend));
const preview = unwrap(
await branchForEvolution(sourceStore, evolutionPlan, copies.makeBackend),
);
```
Each call allocates a fresh allocation in the same ledger, in the `ephemeral`
state, and returns an **empty, schema-mutable** backend: the caller (or the
branch API) seeds it and may commit new kinds and fields, which the fixed-schema
`ephemeral` and `durable` copies refuse. Closing the backend drops the
allocation. While it is live it appears in `listUnsealedAllocations()`, and if
its owner crashes without closing it, `abortAllocation(id)` removes everything
it owns.
Because the graph is unknown when the backend is allocated:
- **Vector tables.** A graph that declares embeddings creates its per-field
pgvector tables after allocation, so the ledger manifest cannot list them.
Dropping an allocation therefore also removes every table in its schema whose
name starts with the allocation's reserved vector prefix. That prefix is
fixed-length and never truncated, so it cannot match another allocation's
tables. `connect` always receives the allocation vector strategy for
`makeBackend`; bind it with
`createPostgresBackend({ vector: allocation.vectorStrategy })`. A connection
that binds any other vector strategy is refused with a `BranchError`, because
it could create tables the allocation does not own. Pass `vector: false` to
opt out of vector support.
- **Graph indexes.** PostgreSQL index names are database-global, so a declared
index cannot reuse its logical name on a private table. `makeBackend` scopes
each declaration to the allocation (`gix_`) the first time the
Store's `materializeIndexes()` sees it, leaving logical names and schema
hashes unchanged and never touching the source's or another allocation's
indexes. A backend you derive from the returned one with `deriveBackend`
inherits the scoping; one you build by copying its members does not.
- **Fulltext.** The same bundled `tsvectorStrategy` requirement applies as for
the cloned copies.
`control` and `connect` must run as the same role
([One database role](#one-database-role)). Both must also use the allocation's
schema ([One schema per allocation](#one-schema-per-allocation)); a pooled
connection's own `search_path` does not decide where anything is created.
### Forked working copies
A second bundled strategy, `forkedWorkingCopyStrategy({ fork, connect })`,
targets a fork-capable host instead of a streamed-interchange clone: `fork`
asks the host itself to produce a complete, independent copy of the database
`baseStore` is on, and `connect` opens a backend on that copy.
```typescript
import {
asBranchId,
branch,
forkedWorkingCopyStrategy,
unwrap,
type ForkHandle,
} from "@nicia-ai/typegraph/graph-merge";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { decorateBackend } from "@nicia-ai/typegraph/backend";
import { drizzle } from "drizzle-orm/node-postgres";
import { Pool } from "pg";
// A host whose fork call returns a new connection string for the branch —
// this is the shape of the copy-on-write branching APIs some Postgres hosts
// offer (Neon and Supabase branches, for example), without either SDK.
type HostBranch = ForkHandle & Readonly<{ connectionString: string }>;
const strategy = forkedWorkingCopyStrategy({
fork: async () => {
const created = await hostBranchApi.createBranch(baseDatabaseId);
return {
connectionString: created.connectionString,
dispose: async () => hostBranchApi.deleteBranch(created.id),
};
},
connect: async (fork) => {
// `createPostgresBackend` takes a Drizzle database, not a pool — open
// one here. Its `close()` deliberately does not end a caller-owned pool
// (Drizzle leaves connection lifecycle to the caller), so compose the
// pool's own shutdown into this fork's `close` through the public
// `decorateBackend` (never a spread) — `branch()`'s composed close then
// ends the pool along with releasing the fork.
const pool = new Pool({ connectionString: fork.connectionString });
const backend = createPostgresBackend(drizzle(pool));
return decorateBackend(backend, {
close: async () => {
await backend.close();
await pool.end();
},
});
},
});
// `makeBackend` is ignored once an explicit strategy is supplied — pass a
// factory whose only job is to reject if it is ever called by mistake.
const rejectMakeBackend = () =>
Promise.reject(new Error("makeBackend must not be called"));
const worker = unwrap(
await branch(
base,
rejectMakeBackend,
{ id: asBranchId("worker-1") },
strategy,
),
);
// ... write on worker.store, plan and apply the merge ...
await worker.close();
```
`TFork` must extend `ForkHandle` (`{ dispose?: () => Promise }`).
`forkedWorkingCopyStrategy` supplies ephemeral copies only. Its base-version
comparison checks the graph's schema and revision or live-content anchor;
the host fork must preserve the full physical database, including TypeGraph
sidecars and extensions. A durable host strategy must persist its branch ID
and attest the sealed origin when reopening it.
For a hosted PostgreSQL branch such as [Neon](https://neon.com/docs/get-started-with-neon/workflow-primer),
`connect` must use that branch's connection string and compute endpoint for
every pooled checkout and transaction. Reusing the source pool can appear to
pass a base-version check while writing to the source. Doltgres can pin a
connection through a [database revision specifier](https://www.doltgres.com/docs/reference/version-control/branches/);
avoid session-level branch switching on a pool whose checkouts may retain
different branch state. Doltgres exposes native branch and merge commands,
but TypeGraph continues to use its own merge planner and apply path; native
merge and Doltgres backend support require separate conformance testing.
`create()` calls `fork(baseStore)`, then `connect(fork)`; the connected
backend's `close` is composed with the fork's `dispose` through `deriveBackend`
(never a spread), so `worker.close()` — the branch's public release call —
releases both the connection and the fork. A `connect` failure disposes the fork before
rethrowing, leaving nothing open and the base untouched.
A fork inherits the base's WHOLE construction option set — hooks, upsert
coalescing, the SQL schema (custom table names), the auto-refresh-statistics
threshold, query defaults, and an externally-bound recorded-read relation —
read once through `Store.workingCopyOptions`, plus `history`/
`revisionTracking`, matched to the base's own `historyEnabled`/
`revisionTrackingEnabled`. This is safe precisely because a fork is the SAME
physical database as the base: a custom `schema` names relations the fork
carries too, and an external `recordedRead` binding points at one. The clone
strategy inherits only `revisionTracking` — its fresh backend is a distinct,
empty database, so a schema naming the base's tables or a `recordedRead`
binding populated nowhere on the clone would misdirect it.
Because the fork's store reads and writes through the base's table names,
`connect()`'s backend must bind those SAME names. `create()` compares the
connected backend's own table bindings against the base's own resolved SQL
schema (`Store.revisionSchema` — the base's explicit `schema` option, or its
backend's own `tableNames` otherwise), and refuses with a `BranchError`,
closing the backend first, when they disagree: a backend bound to different
(often just the default) table names would read and write through tables the
fork's rows were never written to.
**A fork preserves what a clone drops, and that is why it is safe to merge.**
The clone strategy above streams the base through public interchange with
`includeDeleted: false`, so it omits every soft-deleted row entirely: the
interchange `meta` schema has no `deletedAt` field, so a tombstoned row would
otherwise round-trip as LIVE and read as a spurious resurrection on the
clone's diff. It also regenerates `created_at`/`updated_at` on import — safe
only because the merge's state diff always compares against the *original*
base store, never the clone. A fork is never rebuilt through
`exportGraphStream`/`importGraphStream`, so none of that applies: tombstones,
`created_at`/`updated_at`, and the `version` column carry over unchanged, and
— with `history: true` — the fork physically carries the base's recorded
relations, so `store.asOfRecorded()` answers from
that history. A clone-based branch never enables history, so the same call on
it refuses outright.
`create()` asserts `computeBaseVersion(forkStore) === base` right after
attaching the store, where `base` is the token `branch()` already stamped off
the ORIGINAL base store before invoking the strategy — cheap when the base has
revision tracking (an O(1) anchor compare), an O(graph) content fingerprint
otherwise, and computed exactly once either way. This proves base-token
equality at the instant the fork was taken, not byte-for-byte physical
identity: the untracked fingerprint deliberately omits tombstones,
`created_at`/`updated_at`, the `version` column, and recorded history (the "A
fork preserves what a clone drops" paragraph above) — providing those
unchanged is the FORK MECHANISM's job, not something this assertion re-verifies
on every branch. That is still the right fence: the merge's lost-update guard
reads `version` and the diff reads tombstones/timestamps straight off the
fork, so a `fork` that is not a true physical copy breaks them regardless of
what the content fingerprint agrees on. A mismatch closes the backend first
and refuses with a `BranchError` carrying `forkVersion`/`baseVersion` in
`error.details`; `branch()` catches it and returns that `BranchError` as the
`cause` of the outer `BranchError` it resolves with. Only a base-token
mismatch is refused here — a fork taken while the base was mid-write, or a
`fork` that returns a different graph; divergence confined to the physical
state the token omits (tombstones, timestamps, row versions, recorded
history) passes the fence, and keeping that state faithful remains the fork
mechanism's contract.
`create()` also refuses BEFORE ever attaching a store when `connect()`'s
backend aliases the base's own backend: the same backend object, one derived
from the other through `deriveBackend`, or two wrappers sharing one underlying
connection. Without this check, a `connect()` that mistakenly hands back the
base's own backend (a cached factory keyed by database name, say) would pass
every fence below trivially — every write on the "fork" would actually mutate
the base, and closing the working copy would close the base's own backend. The
refusal disposes only the fork (never the aliased backend, which the base
still owns) and throws a `BranchError` naming `connect()`. This cannot detect
every aliasing shape: a fresh backend built over the base's own connection
pool is indistinguishable from a real fork's connection when that pool audits
as independent (the normal case for a default-size `pg.Pool`) — a pooled
checkout genuinely is a different connection from the pool's perspective.
`ingestionBranch()` stays clone-based. Its strategy derives a working-copy
schema with node uniqueness deferred so an untrusted batch's repeated keys can
reach entity resolution before validation; a host-level fork carries the
base's schema exactly, uniqueness included, with no hook to relax it.
:::caution[Suspend hazard]
A fork-capable host that suspends idle compute to reclaim it between requests
drops that compute's in-process state, including anything memoized against a
particular connection or session. TypeGraph's own locking already assumes
this rather than trusting a lock survives idle time: the recorded-write lock
memo (`RecordedGraphLockMemo`, populated by
`memoizeAcquiredRecordedGraphWriteLock`) and the schema-fence lease
(`memoizeLeasedSchemaFence`) are both keyed weakly by the transaction-scoped
backend object, so they hold for exactly one transaction's lifetime and
re-acquire on the next one, and the write fence itself (see
[Write fence declaration](/backend-setup#write-fence-declaration-writefence))
is resolved and its lock taken fresh per transaction, never cached across
one. An ordinary sequence of separate `store` calls — each its own
transaction — therefore tolerates a suspend between any two of them.
What does NOT tolerate a suspend is a single `store.transaction` callback:
every read and write the callback issues, and the lock it holds, runs on one
native database transaction over one connection, so a suspend partway
through drops that connection out from under the callback and aborts
whatever was in flight. Keep a `store.transaction` callback's wall-clock
duration short and free of anything that could let the host suspend
underneath it — an external API call, a human approval step, a long queue
wait — and commit a long-running workflow across multiple `store.transaction`
calls instead of holding one open across such a wait.
:::
### Durable host-native branches
`branchDurable()` is the persistent counterpart to `branch()`. A
`DurableWorkingCopyStrategy` allocates a host branch, opens a Store on it, and
returns a non-secret JSON locator. TypeGraph seals the immutable fork origin
beside that allocation and returns a `DurableBranchDescriptor` that can cross a
queue, process, deployment, or machine boundary.
For a remote host, persist a chosen `{ id, allocationId }` before calling
`branchDurable(base, strategy, { id, allocationId })`. `create()` receives both
and must refuse an allocation ID that may already exist. If the host allocates
a branch but its response is lost, use host tooling to inspect the ID and
recover or remove the allocation before retrying. A failed create reports both
IDs for that reconciliation. The host must never allocate a second physical
copy for the same ID or return a sealed copy as though it were new.
```typescript
import {
applyDurableMergePlan,
branchDurable,
destroyDurableBranch,
planMerge,
reopenDurableBranch,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const created = unwrap(await branchDurable(base, durableStrategy));
await created.branch.store.nodes.Person.create({ name: "Ada" });
// Releases this process's connection and writer lease. The host branch stays.
await created.branch.close();
await queue.put(JSON.stringify(created.descriptor));
// A later process reconstructs the ordinary GraphBranch used by planning.
const descriptor = JSON.parse(await queue.get()) as typeof created.descriptor;
const reopened = unwrap(
await reopenDurableBranch(graph, descriptor, durableStrategy),
);
const plan = unwrap(await planMerge(base, [reopened]));
// Applies the complete TypeGraph plan inside the target transaction.
const report = unwrap(
await applyDurableMergePlan({
target: base,
branch: reopened,
descriptor,
strategy: durableStrategy,
plan,
}),
);
await reopened.close();
unwrap(await destroyDurableBranch(descriptor, durableStrategy));
```
Closing and destroying are deliberately separate. `GraphBranch.close()` closes
the backend and releases its access lease, but leaves the persistent allocation
reopenable. `destroyDurableBranch()` asks the strategy to attest the complete
origin and delete or archive that allocation atomically. A descriptor is
untrusted input: TypeGraph checks its allocation id, graph definition, branch
id, base token, schema anchor, and engine revision against the origin the host
sealed. The allocation id is independent of the caller's branch id, so
swapping or relabeling a locator cannot authorize deletion of another copy
even when two copies were given the same branch id.
Strategies write new locators using `version` and may list older supported
locator versions in `readableVersions`. Every method must understand each
listed version, including destroy and evidence access.
The strategy locator must be JSON-safe and **must not contain secrets**. Use a
branch id, database id, or other lookup key, then resolve credentials from
strategy-owned configuration. TypeGraph returns the locator to application code
so a connection URL, password, or bearer token placed there can escape through
ordinary descriptor storage. Framework cleanup errors deliberately omit the
locator and raw host cleanup error from diagnostic details.
#### Exact forks and access leases
After `strategy.create()` returns, TypeGraph recomputes `base@V` from the source.
A source write racing allocation therefore refuses and aborts the working copy
instead of sealing a branch from the wrong ancestor. TypeGraph then accepts an
exact matching working-copy token as the fast path. When a strategy creates an
equivalent persistent copy with an independent revision namespace, TypeGraph
instead verifies that its complete merge-visible graph state has no delta from
the source, fencing the source again after enumeration. The host remains
responsible for physical fidelity outside TypeGraph's graph semantics.
To enable lineage-pruned merge diffs, `create()` may return `forkRevision`
captured atomically with the physical fork. When it cannot prove that cut, omit
the revision and TypeGraph compares the complete graph state; reading a later
revision after the copy was opened could miss an intervening branch write.
Every `create()` and `reopen()` also returns a `DurableWorkingCopyAccess`:
- `engine-fenced` says the database provides sound cross-client isolation and
change fencing for the full Store planning/apply access pattern, across every
connection and process that could mutate the working copy.
- `exclusive` carries an allocation-wide writer lease. The strategy must acquire
it before returning and exclude every other process and backend instance.
TypeGraph closes the backend first, then releases the lease; a failed release
is retried by the next `close()` call.
Do not use `engine-fenced` merely because one backend object serializes its own
calls. A `caller-serialized` backend owns one in-memory queue per backend
instance, so two reopened pools or two processes still race. Such an engine must
use a host-wide `exclusive` lease, and a concurrent reopen must wait or refuse.
Merge planning also assumes the working copy is quiescent while it is diffed.
#### Native database branches
A strategy may allocate a working copy using a database-native branch, but
`applyDurableMergePlan()` always applies the approved TypeGraph plan through
the target Store transaction. The former native-merge callback was removed:
it could commit outside the transaction that checked the target revision.
A future native merge capability needs a host-native compare-and-swap on the
actual target, plus proof that the full physical diff equals the approved
TypeGraph writes, including schema, history, identity, and sidecars.
For a Doltgres strategy, pin each Store connection to the intended database
branch. [Doltgres revision specifiers](https://www.doltgres.com/docs/reference/version-control/branches/)
provide that connection-level selection. Its
[`DOLT_BRANCH()` and `DOLT_MERGE()` functions](https://www.doltgres.com/docs/reference/version-control/dolt-sql-functions/)
implicitly commit the current transaction, so a fence checked before those
functions cannot by itself protect their target.
#### Atomic operations and immutable evidence
A `DurableWorkingCopyStrategy` may also expose an optional `operations`
capability (`DurableOperationCapability`). It lets a durable host combine one
opaque graph mutation with its immutable operation evidence in a **single host
transaction**.
TypeGraph owns descriptor validation, sealed-origin attestation, request
canonicalization, and evidence validation; the host owns the database mechanics.
```typescript
import {
durableBranchHasUndeliveredEvidence,
getDurableOperation,
markDurableOperationDelivered,
operateDurableBranch,
scanDurableOperations,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const request = {
idempotencyKey: "statement-42",
// Host-defined, JSON-safe description of the graph change to apply.
mutation: { kind: "statement", op: "upsert", payload: { subject: "s-1" } },
// Host evidence, retained verbatim. TypeGraph never interprets either field.
metadata: { source: "etl", schemaVersion: 3 },
};
const outcome = unwrap(
await operateDurableBranch(descriptor, durableStrategy, request),
);
if (outcome.outcome === "unsupported") {
// The strategy applied no mutation and wrote no evidence; TypeGraph refuses
// rather than emulating atomicity with best effort or callbacks that run
// outside the evidence transaction.
throw new Error(`Missing capabilities: ${outcome.dimensions.join(", ")}`);
}
console.log(outcome.outcome); // "applied" | "replayed"
// Newly applied evidence is always false. A replay returns the current
// committed delivery state, which may already be true.
console.log(outcome.evidence.delivered);
```
Both `mutation` and `metadata` are **JSON-safe host values**. TypeGraph never
interprets their application fields; it canonicalizes `metadata` plus `mutation`
into the `operationDigest` and otherwise carries them through untouched. The
digest covers the complete request except the idempotency key, so reusing a key
with a different mutation *or* different metadata conflicts. Non-JSON content is
refused before any host call.
`metadata` is retained as evidence; `mutation` is the host's own description of
the graph change it must apply atomically with the evidence row. The strategy
attests the caller's `expectedOrigin` against the allocation the descriptor
names, exactly as reopen and destroy do. Every committed
operation returns `before`/`after` coordinates — the merge-visible `base`
fingerprint and, when the working copy resolves lineage, the engine `revision`.
TypeGraph validates that the returned evidence echoes the canonical request and
digest; a host cannot forge a different digest, echo a different request, or
return non-JSON metadata (`DurableOperationEvidenceError`).
**Idempotency.** The strategy treats `idempotencyKey` as its unique key:
- Identical key **and** digest: returns the previously committed evidence
(`outcome: "replayed"`) and re-applies nothing. Because delivery marking is
monotonic, a replay after delivery legitimately returns `delivered: true`.
- Identical key with a **different** digest: refuses with
`DurableOperationConflictError` and mutates nothing.
A first application (`outcome: "applied"`) must return `delivered: false`.
TypeGraph rejects `applied` evidence that is already delivered, so a host cannot
bypass downstream delivery or the destroy fence. It also validates the complete
host outcome envelope: malformed outcomes and empty, duplicate, or unknown
`unsupported` dimensions return `DurableOperationEvidenceError`.
**Evidence access and delivery.**
- `getDurableOperation(descriptor, strategy, idempotencyKey)` reads one
operation's evidence, or `undefined` when it was never committed.
- `scanDurableOperations(descriptor, strategy, { after?, limit? })` returns
`{ operations, cursor, hasMore }` in monotonic commit order, with ties broken
deterministically. Pass the opaque `cursor` back as `after` to resume, even
after `hasMore: false`; later commits must sort after that cursor. An empty
page echoes `after`, and only an empty initial scan omits `cursor`. `limit` defaults to
`DURABLE_OPERATION_SCAN_DEFAULT_LIMIT` (100) and may not exceed
`DURABLE_OPERATION_SCAN_MAX_LIMIT` (1000); a larger page is refused.
- `markDurableOperationDelivered(descriptor, strategy, idempotencyKey)` marks
one operation delivered, idempotently: marking an already-delivered operation
returns the same evidence and writes nothing, and an unknown key returns
`undefined`.
- `durableBranchHasUndeliveredEvidence(descriptor, strategy)` reports whether
any committed evidence is still undelivered — the queryable half of the
destroy fence below.
`operateDurableBranch()` is the only orchestrator that tolerates a missing
capability: a strategy with no `operations` returns the explicit `unsupported`
outcome (`dimensions: ["atomicMutation"]`) having executed no host call. `get`,
`scan`, `markDelivered`, and `hasUndelivered` instead refuse with a typed
`DurableOperationUnsupportedError`. TypeGraph never emulates the atomic
guarantee: a callback that runs inside the strategy's own evidence transaction
(as `apply` does in the bundled PostgreSQL manager below) is the host's atomic
mutation, while best effort or a callback outside that transaction is refused.
**Destroy fence.** A strategy with `operations` MUST refuse destruction while
undelivered evidence remains, throwing `DurableEvidenceUndeliveredError`;
`destroyDurableBranch()` preserves that typed refusal instead of flattening it
into a generic branch failure, so the caller can still recover the evidence.
Deliver (or archive) the outstanding evidence before destroying the branch.
Concurrent `operate` and `destroy` are serialized by the host's own transaction:
either the operation commits first (destroy then observes undelivered evidence
and refuses) or destroy commits first (the operation fails against the removed
allocation). No partial state is ever observable.
##### Bundled PostgreSQL manager
`createPostgresWorkingCopyManager` implements the capability when given an
`operations` option. `apply` is how the host's opaque mutation reaches the
graph; TypeGraph still never interprets `mutation`.
```typescript
const copies = createPostgresWorkingCopyManager({
control,
connect,
operations: {
graph,
// Runs inside the transaction that commits the evidence row. A throw rolls
// back both the mutation and the evidence.
apply: async (transaction, mutation) => {
await applyHostMutation(transaction, mutation);
},
},
});
const outcome = unwrap(
await operateDurableBranch(descriptor, copies.durable, request),
);
```
`operations.graph` is required because a capability member receives only the
descriptor, so the manager must reopen the allocation from the graph the host
names. Before any connection or transaction opens, every member checks that
graph against the sealed allocation's attested origin: its graph id and its
version-blind definition hash must equal the ones the branch was forked with, so
a graph that reuses the id with a different definition is refused. `apply`
receives the transaction-scoped context of the allocation's fixed-schema Store,
the same context `store.transaction` provides, so the allocation's fixed schema
applies. Without the option, `copies.durable.operations` is undefined and
`operateDurableBranch()` returns `unsupported` (`atomicMutation`).
Each durable allocation owns one evidence relation under its ledger-reserved
physical prefix, created in the provisioning transaction and dropped by destroy.
The ledger records whether an allocation has one (`operation_evidence`).
`operate` takes the allocation lock on the allocation's own transaction session,
attests the sealed origin, resolves idempotency, takes the graph write lock,
computes the `before` coordinates, calls `apply`, computes the `after`
coordinates once the transaction's revision bookkeeping has run, and inserts
undelivered evidence, all in one transaction. The allocation lock is a
transaction-scoped advisory lock keyed on the allocation id, in a namespace of
its own so it can never collide with a graph's write lock. It serializes
operations per allocation, so the evidence sequence that backs the opaque scan
cursor is commit order, and each operation's `before` equals the previous
operation's `after` whenever every writer to the allocation goes through
`operate` or takes the graph write lock. Ordinary writes take that lock on an
allocation that tracks history or revisions, so a direct write cannot commit
between `before` and `apply`; on an allocation that tracks neither, a direct
write is not fenced and the evidence's `before`/`after` pair may include it.
The graph write lock is graph-wide. While `apply` runs, tracked writes to the
source graph and to every sibling working copy of it wait on that lock, so keep
`apply` short and do not wait on other graph writers inside it.
Coordinates always carry `base`. They also carry `revision`, the engine
revision, when the allocation resolves lineage, which is when it tracks history
or revisions; both are read on the transaction's own session so they describe
one state. An allocation that tracks neither reports no `revision`, and its
`base` values are content fingerprints, which read the whole graph twice per
operation.
**Isolation is observed, not assumed.** `operate`, `markDelivered`, and destroy
each request READ COMMITTED, and the statement that takes the allocation lock
also reports the isolation level its session actually runs at. Any other level
is refused before anything is read or written, with a `ConfigurationError` whose
`details.code` is `WORKING_COPY_ISOLATION_UNSUPPORTED`, because the request is
honored only where a backend supports it and a role or server default of
REPEATABLE READ would otherwise give the fence and the idempotency lookup a
snapshot older than the lock wait. A `control` or `connect` wrapper must
therefore forward the transaction `isolationLevel` option. The refusal only
fires when a wrapper drops the requested option and the session's default is
not READ COMMITTED.
The same check runs everywhere the manager drops an allocation, not only in
destroy and `abortAllocation`: closing an ephemeral working-copy store, closing
a `makeBackend` backend, and the cleanup after a failed allocation. The first
two surface the refusal from `close()`. The cleanup swallows it so the
allocation's original failure reaches the caller, which leaves the allocation
behind. Every such orphan is discoverable with `listUnsealedAllocations` and is
removed by `abortAllocation` once `control` forwards the option.
**Destroy fence.** Destroy (and `abortAllocation`) takes the same allocation
lock. `destroyDurableBranch()` refuses with `DurableEvidenceUndeliveredError`
while undelivered evidence exists, even from a manager built without
`operations`; delivering the evidence requires a manager built with
`operations`. An in-flight `operate` and a destroy on one allocation serialize
on the lock: whichever commits first decides the other's outcome. The destroy
waits at most `cleanupLockTimeoutMs` (5000 ms by default); one that outwaits a
long `apply` fails with the database's lock timeout having committed nothing,
and can be retried after the operation settles. `get`, `scan`, and
`hasUndelivered` take no allocation lock, so they never wait behind an `apply`.
A destroy that commits after any member has attested the sealed row but before
that member holds the allocation (before its connection is attested, before
`operate` mints the revision origin, or before `get`, `scan`, or `hasUndelivered`
reads the evidence relation) fails the member with one `BranchError`
(`changed owner or was destroyed during the operation`), the same error
`operate` and `markDelivered` raise against a removed allocation. It is never a
raw missing-relation error or the "connection is not bound to the allocation
database" refusal, which is reserved for a connection that reaches a different
database than the one the ledger names.
The fence follows the manager's [removal rule](#one-schema-per-allocation). It
reads the evidence relation only in the allocation's own schema. Evidence
relations that sit in another schema refuse removal before the fence runs and
are kept. The fence runs only when the evidence relation is among the
relations removal drops. An allocation whose evidence relation is gone has no
evidence left to deliver and nothing to recover, so destroy removes its
remaining relations and its ledger row, exactly as it does for an allocation
whose relations exist nowhere.
**Mixed-version deployments.** Only managers on this version take the allocation
lock and honor the destroy fence. A manager from an earlier release that shares
the ledger destroys an allocation without consulting its evidence, so
undelivered evidence is lost with the allocation, and it does not drop the
evidence relation, so a later `allocate` with the same id refuses because
`op_evidence` exists without a ledger row. Upgrade every process that
shares a working-copy ledger before any of them creates or destroys a durable
allocation. To recover an orphaned evidence relation, read its undelivered rows
(`WHERE NOT delivered`) and deliver them, then drop the relation the refusal
names and retry. TypeGraph never drops it for you, because it may hold the only
copy of undelivered evidence.
An allocation provisioned by an earlier release has no evidence relation, and
its ledger row says so without any statement that changes the database.
`operate` returns `unsupported` with `dimensions: ["evidenceStore"]`. The only
statement it runs is one read-only ledger `SELECT` through `control`; it runs no
DDL, takes no lock, calls no `connect`, and applies and writes nothing. The read
members report no evidence: `get` and `markDelivered` return `undefined`, `scan`
returns an empty page (echoing `after`), and `hasUndelivered` returns `false`.
Re-fork the branch to gain evidence.
### Constraint-aware ingestion branches
For a bounded candidate batch, `planCandidateWriteSet()` hides the transient
branch lifecycle completely. It accepts a validated, versioned JSON document,
stages it through the same constraint-aware ingestion implementation, delegates
to incremental merge planning, and closes the working copy on every outcome.
The result is the ordinary `MergePlanArtifact`, so review and application use
the same APIs as every other merge plan.
On eligible revision-tracked graphs, planning
seeds existing candidate rows, edge
endpoints, cardinality peers, live same-id ontology peers, and any reachable
current identity component into the disposable working copy.
The resolver still queries the live target for declared unique and index peers,
and the plan retains its ordinary provenance, conflicts, digest, and commit-time
fences. Existing undeclared target properties survive staging; extra candidate
properties are refused. A custom backend without the active-only source read
uses the complete clone path for `oneActive` graphs. A custom backend without
the keyed match-identity owner read, or a candidate whose owner is excluded
from the clone projection, also uses that path. Other ineligible graphs use
the complete clone path so staging
still checks constraints that can depend on rows beyond the candidate's ids.
```typescript
import {
captureCandidateWriteSetTarget,
planCandidateWriteSet,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const writeSet = {
formatVersion: 1,
sourceId: "provider-a",
target: await captureCandidateWriteSetTarget(store),
nodes: [
{
kind: "Patient",
id: "provider-a:123",
properties: { name: "Ana", mrn: "123" },
validFrom: "2026-01-01T00:00:00.000Z",
},
],
edges: [],
} as const;
const plan = unwrap(
await planCandidateWriteSet({
target: store,
makeBackend,
writeSet: JSON.parse(JSON.stringify(writeSet)),
options: {
resolve: {
Patient: {
blockIndex: "patient_mrn_candidates",
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.9,
},
},
},
}),
);
```
`sourceId` is the stable attribution carried into conflicts, resolutions, and
provenance; node and edge ids remain the contribution source ids. The target
schema identity prevents a document authored against one graph contract from
being staged against another. `validFrom` is required (and may be `null`) so
replaying identical JSON cannot acquire a new import-time timestamp and change
the plan digest.
This adapter applies TypeGraph's existing entity/property merge semantics. Two
distinct records that both validate do not conflict merely because an
application interprets their subject, predicate, time, source, or value fields
as disagreement. Domain-specific acceptance and Statement semantics remain in
the consuming application.
Use `ingestionBranch()` when an untrusted ingestion batch may contain aliases
that deliberately repeat a canonical node's unique key. An ordinary `branch()`
keeps the complete graph schema and rejects the duplicate during staging,
before entity resolution can review and collapse it. An ingestion branch
materializes an honest working-copy schema with only node uniqueness deferred;
schema validation, edge endpoint checks, disjointness, and edge cardinality
still apply immediately.
```typescript
import { asNodeId } from "@nicia-ai/typegraph";
import {
applyMergePlan,
asBranchId,
ingestionBranch,
planMergeIncremental,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
import { importGraph } from "@nicia-ai/typegraph/interchange";
const incoming = unwrap(
await ingestionBranch(base, makeBackend, {
id: asBranchId("provider-a"),
}),
);
const imported = await importGraph(incoming, providerDocument, {
onConflict: "error",
onUnknownProperty: "error",
});
if (!imported.success) throw new Error("Provider import was rejected");
const alias = await incoming.nodes.Patient.getById(
asNodeId("incoming-patient"),
);
if (alias === undefined) throw new Error("Imported patient was not found");
// `canonicalPatient` is an existing Patient read from the base before forking.
// The repeated MRN and its identity evidence can be staged together.
await incoming.identity.assertSame(canonicalPatient, alias);
const plan = unwrap(
await planMergeIncremental({
forkPoint: base,
target: base,
branches: [incoming],
options: {
onBasePropertyConflict: "flag",
resolve: {
Patient: {
blockIndex: "patient_mrn_candidates",
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.9,
},
},
},
}),
);
const applied = unwrap(await applyMergePlan(base, plan));
await incoming.close();
```
Both `importGraph()` and `importGraphStream()` accept the returned handle, so an
interchange document can be staged without a hand-written collection copy loop.
Import remains the single owner of node-first ordering, validity windows, edge
endpoint order and reference validation.
On an identity-enabled graph, the handle also exposes an assertion-only
`IdentityAssertionWriteFacade` as `identity`: `assertSame`, `assertDifferent`,
`bulkAssertSame`, and `bulkAssertDifferent`. This lets a batch stage aliases
that repeat unique keys and the explicit identity evidence needed to reconcile
them before merge-time constraint validation. Assertion contradictions and
invalid endpoints are still refused while staging; only node uniqueness is
deferred.
The returned handle exposes those ingestion collections and identity assertion
writes, not the branch's underlying `Store`. Identity reads and retractions,
schema operations, transactions, and runtime internals remain unavailable, so
callers cannot bypass the deferred-constraint contract. As with `Store`, the
`identity` property is absent at the type level when the graph does not enable
Operational Identity.
The original graph definition remains the merge contract:
`applyMergePlan()` validates node uniqueness against the entire resolved write
set in the target transaction. Valid key handoffs and swaps are accepted as one
set. If reviewed resolution leaves two live owners of the same unique key, the
merge returns `MergeConstraintConflictError` and commits no graph or provenance
writes.
The derived schema is persisted on the working-copy backend, so the relaxed
contract is auditable and an explicit reattachment with an equivalent graph
definition verifies the same constraint behavior. `ingestionBranch()` does not
expose a general reopen/resume API. Deferral is not an in-memory flag and does
not disable database constraints ad hoc. Ingestion branches require a backend
with the batch uniqueness operations needed for atomic final validation.
Unsupported backends are refused rather than falling back to sequential checks.
## Valid-time windows
**A new row's window travels with the merge.** An explicitly open-left row stays
open-left through snapshot and incremental merges, including edge repointing.
Reviewable plans serialize that lower bound as `validFrom: null`; an omitted
plan field means the write states no lower-bound change. JSON export/import and
plan application preserve the distinction.
A branch-authored node or edge window — including a deliberately ended one on a resurrection — is written
as-is by the commit rather than reset to merge time. When the incremental
target itself also created the surviving row, the target's committed window
wins.
**An inherited row's end-of-validity is merged.** Both
`update(id, {}, { validTo })` and `update(id, {}, { clearValidTo: true })` on a
branch are ordinary writes. The merge carries the set, move, or reopening to the
target even when the row's properties are untouched:
```typescript
await fork.store.nodes.Patient.update(asNodeId("pat-1"), {}, { validTo: "2030-06-01T00:00:00.000Z" });
const report = unwrap(await merge(base, [fork]));
// base now holds pat-1 with valid_to = 2030-06-01, and:
report.validityEnds;
// [{ entity: "node", kind: "Patient", id: "pat-1",
// validTo: "2030-06-01T00:00:00.000Z", claimedBy: ["worker-1"] }]
await fork.store.nodes.Patient.update(asNodeId("pat-1"), {}, { clearValidTo: true });
// A merge now reopens pat-1 and reports:
// [{ entity: "node", kind: "Patient", id: "pat-1",
// clearValidTo: true, claimedBy: ["worker-1"] }]
```
An ending is treated as a **sibling of deletion**, not as a property, because it
makes the same kind of statement: *this stopped being true*. That single choice
explains the whole contract:
| Situation | Outcome |
| --------- | ------- |
| One branch ends the row | That end is written — including a *later* end, which extends the window. |
| One branch reopens the row | The end is cleared with `clearValidTo: true`. |
| Several branches end it differently | No conflict. The **earliest** end wins, and `report.validityEnds` names every claiming branch. |
| Sibling branches end and reopen it | The end wins as the stronger monotone claim; every claimant remains visible in `report.validityEnds`. |
| The incremental target already ended it | The target's end stands. A branch never re-windows a row the target itself windowed, and the row is left out of the merge's writes entirely — but the discarded claims are still reported, as an entry carrying `precedence: "target"` and the target's own instant. |
| One branch ends it, another deletes it | Deleted, with **no** `DeleteModifyConflict` — the stronger statement absorbs the weaker one. |
| A branch re-states the end the target holds | No write at all — nothing is staged, so there is no version bump or history row even with `coalesceUnchangedUpserts` off. |
| No branch touched the window | Untouched. A properties-only edit never passes a window, so the committed one stands. |
The earliest-end rule is fixed, not a policy knob: it is commutative and
associative, so the merge stays order-independent, and `onPropertyConflict`
never sees a property your schema does not have.
**The branch that authored the committed end is credited.** An ending is
authored state, so its author is a contributor to that row in
`report.provenance` and in the durable sidecar — even when moving the window is
the only thing that branch changed. Credit follows the *committed* end: when
several branches end a row differently, only the branches whose claim equals the
written instant are credited, while `validityEnds[].claimedBy` still names every
claimant, winning or not. An ending a deletion absorbed commits nothing, so it
credits nobody, and neither does an entry marked `precedence: "target"` — the
merge committed none of that end.
**Every claim the merge observed is visible in `validityEnds`, applied or not.**
An entry with no `precedence` is one the merge *decided*: `validTo` is the
instant it wrote, or `clearValidTo: true` says it reopened the row. An entry
with `precedence: "target"` is one it did **not** — the incremental target had
already changed that end, so the entry describes the target's set or clear,
`claimedBy` names the branch claims that were thrown away, and nothing was
written or credited for the row. A row no branch claimed at all produces no
entry, since there was nothing to discard.
`validityEnds` reports claims about rows inherited from the fork point. If the
fork point is empty, every branch row is branch-created and the array is always
empty. A demo or topology that needs to exercise this report must seed the row
before branching, then end that inherited row on one or more branches.
Because an ending is not a modification, `onDeleteModifyConflict` never sees
one: a row whose *only* change is its window loses to a concurrent deletion even
under `"prefer-modify"`, since there is no modification to prefer. A row with a
properties edit *and* an ending keeps the usual delete/modify behavior on the
properties, and its ending rides along only if that modification survives.
**What is still NOT merged, and why.** On a row that is live in both the base
and the branch, `validTo` is the only window field a branch can author *and* the
commit can apply. A row's lower bound is immutable outside resurrection —
`validFrom` is written only when a soft-deleted row is brought back — so that
lower-bound delta remains observable in a fork but unapplicable:
| Observed delta | Reachable how | Merged? |
| -------------- | ------------- | ------- |
| `validTo` set or moved | `update(id, {}, { validTo })` | **Yes** |
| `validTo` cleared back to open | `update(id, {}, { clearValidTo: true })` | **Yes** |
| `validFrom` changed | soft-delete + resurrect inside the fork | No |
Rather than silently ignore it, the merge reports the lower-bound change in
`report.dropped` with reason `"window-not-applicable"`. Reconciling a value the
commit would then drop is worse than not merging it: the report would claim a
change that never happened.
Delete+resurrect can also make an ended base row appear open because resurrection
creates a fresh window. When `validFrom` changed, that open end is part of the
same non-applicable resurrection artifact; it is not treated as a branch-authored
`clearValidTo`, and an incremental target artifact does not outrank another
branch's explicit end claim.
Full interval reconciliation (intersecting `[validFrom, validTo]` across
branches) is deliberately out of scope — it needs a write path that moves a live
row's lower bound, which contradicts the temporal model, and it would silently
discard a branch's extension.
## Forking one graph namespace
`forkGraphNamespace(sourceStore, privateBackend, operationKey)` copies one
history-enabled graph into an independently allocated PostgreSQL database. It
copies the graph's committed schema, current rows, tombstones, recorded-time
relations, revision clock and journal, identity relations, and TypeGraph
materialization records. It checks a repeatable-read source snapshot against a
pre-cut `base@V` token, compares every copied row before target commit, and
returns `{ store, proof, abort }`. One source transaction holds that snapshot
for the entire copy, from its first source read through the target copy and
digest checks. The source can accept writes after the snapshot cut, while the
long-lived snapshot remains open until copying finishes; `proof.sourceBase`
identifies the copied cut.
```typescript
import {
forkGraphNamespace,
prepareNamespaceForkTarget,
} from "@nicia-ai/typegraph/graph-merge";
// Run with the schema owner role before the runtime fork.
await prepareNamespaceForkTarget(sourceStore, privateBackend);
const fork = await forkGraphNamespace(sourceStore, privateBackend, "restore-42");
// Owner role again: builds IVFFlat indexes over the copied rows.
await fork.store.materializeIndexes();
const historical = await fork.store
.asOfRecorded(receipt.recorded)
.nodes.Item.getById(receipt.itemId);
// Publish the private database through your own placement registry only after
// checking the fork and any application-specific restore invariants.
// Before publication, await fork.abort() to discard an unchanged copy.
```
The caller provisions and owns `privateBackend`. It may contain other graph
namespaces, but it must contain no rows for the source graph. TypeGraph refuses
a connection to the source database, including an aliased backend object.
`prepareNamespaceForkTarget()` is the owner-side step, and the fork itself
issues no DDL. It installs the retry ledger, creates the graph's per-field
pgvector tables, and builds every index the source has materialized for the
graph with the DDL the source used. It writes no graph rows and no
materialization records, so it can run before the target is empty-checked,
and running it again is harmless. Indexes whose build never completed on the
source are neither built nor required. IVFFlat indexes are the exception:
IVFFlat clusters the rows present when it is built, so building one on an
empty table gives poor recall. They are not built by preparation and their
materialization records are not copied; run `fork.store.materializeIndexes()`
after the fork to build them over the copied rows. Every other index the fork
carried is already recorded, so that call only builds the IVFFlat ones. An
IVFFlat index left on the target by an aborted fork has no record, so the
next fork's `materializeIndexes()` drops and rebuilds it over the new rows. The target stays private
until the caller changes its own placement pointer;
TypeGraph does not publish it. `abort()` atomically removes the copied graph
and operation marker while preserving unrelated namespaces, and refuses if the
target has changed. A retry with the same operation key returns the same proof
after checking the target digest and base token; a different key cannot reuse
the populated target.
This first-party copy supports the bundled PostgreSQL table layout, bundled
`pgvector` embedding storage, and default `tsvector` fulltext storage.
Embeddings are copied, digested, and verified like every other graph relation,
and `abort()` removes them. A graph with embedding fields forks only between
backends with the same vector storage: pgvector on both sides, or
`vector: false` on both, where embeddings live only in node properties. A
vector-disabled source never wrote the vector tables a pgvector target would
search, so that pair is refused. The fork refuses custom table mappings, custom vector or fulltext
strategies, and contribution-owned tables it cannot copy and validate. The
current copy buffers one relation at a time and
inserts rows in bounded batches, so operators should size the private copy
process for its largest graph relation. It does not use interchange, whose payload lacks
recorded history and tombstones.
## Determinism
Graph Merge is built to be reproducible, which is what lets you retry, cache,
diff, and test a merge with confidence:
- Candidate sets are sorted before clustering; clusters resolve by stable keys.
- Conflict resolution consults only the captured `branchOrder` (or lexicographic
branch id) — never wall-clock.
- The committed graph and the normalized report are a pure function of the
*unordered* branch set.
Use `branchOrder` to make preference explicit wherever a policy needs ordering:
```typescript
const branchOrder = [systemOfRecord.id, agentA.id, agentB.id];
const result = await merge(base, [agentB, systemOfRecord, agentA], {
branchOrder,
onPropertyConflict: "lastWriteWins", // systemOfRecord wins, regardless of input order
});
```
## Errors
Most entry points return a `Result`; the error arm is a typed `TypeGraphError`
subclass you can branch on. `applyMergePlanInTransaction()` instead throws a
typed `MergeError` so a caller-owned transaction callback cannot resolve and
commit after a partially applied failure:
| Error | When |
| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `BranchError` | `branch()` or `ingestionBranch()` could not materialize a working copy. |
| `BaseVersionMismatchError` | A branch forked from a different `base@V` than the target now has (snapshot `merge()`). Also the typed replan error `mergeIncremental()`'s in-transaction guards raise, and the by-ID freshness check both commit modes run, when the target moved in the plan→commit window. |
| `IdentityMergeConflictError` | Code `GRAPH_MERGE_IDENTITY_CONFLICT`. Thrown by both `merge()` and `mergeIncremental()` for identity contradictions, assertion-ID collisions, and retract/reassert races. See the [identity guide](/identity/#interchange-and-branch-merge). |
| `MergeConstraintConflictError` | Code `GRAPH_MERGE_CONSTRAINT_CONFLICT`. The resolved plan would violate a deterministic store constraint, such as edge cardinality or node uniqueness. Its category is `constraint`, its `cause` is the original typed store error, and its details expose the original constraint fields. No graph or provenance writes commit. |
| `InvalidMergeOptionsError` | Code `GRAPH_MERGE_INVALID_OPTIONS`. The supplied option combination is invalid, `mergeIncremental()` was given the snapshot-only `target` option instead of silently ignoring it, or `mergeIncremental()`'s `onBasePropertyConflict` is not `"flag"`. |
| `SimilarityUnavailableError` | A `vector`/`hybrid` strategy was requested with no `embedder`. |
| `MergeConflictError` | A conflict could not be resolved under the configured policy. |
| `MergePlanCapabilityError` | Public planning was requested for a target without the durable revision guarantee needed across processes and time. Enable `revisionTracking` or `history`. |
| `MergePlanningStaleError` | The target moved while planning was reading it. This is an expected retry-and-replan outcome under concurrent writers: no plan was returned, so recapture the target and create a new plan before retrying. |
| `StaleMergePlanError` | The target revision changed after planning, or this plan was already applied. Review a newly-created plan. |
| `InvalidMergePlanError` | The input is not a valid plan artifact. More specific subclasses distinguish unsupported versions, digest changes, and target/schema/origin mismatches. |
| `CandidateSourceError` | A built-in candidate source failed; details identify its source id, entity kind, and operation. |
| `CandidateWriteSetError` | Code `GRAPH_MERGE_CANDIDATE_WRITE_SET`. Candidate JSON is malformed, targets another graph schema, cannot be staged, or violates the active graph contract. The accepted graph is unchanged. |
| `MergeReviewError` | Code `GRAPH_MERGE_REVIEW`. Durable review evidence is malformed, unsupported, incomplete, or inconsistent, or review options cannot be represented safely. |
| `DurableOperationError` | Code `GRAPH_MERGE_OPERATION`. System-category failure while calling a durable-operation host, including transport and strategy failures. |
| `DurableOperationRequestError` | Code `GRAPH_MERGE_OPERATION_REQUEST`. User-category refusal for an invalid durable-operation request, descriptor, or scan option. |
| `DurableOperationConflictError` | Code `GRAPH_MERGE_OPERATION_CONFLICT`. Constraint-category refusal when an idempotency key is reused with a different operation digest. The previously committed operation is returned untouched; nothing new is written. |
| `DurableOperationUnsupportedError` | Code `GRAPH_MERGE_OPERATION_UNSUPPORTED`. The strategy's `operations` capability lacks a requested member; TypeGraph refuses rather than emulating the atomic guarantee. |
| `DurableOperationEvidenceError` | Code `GRAPH_MERGE_OPERATION_EVIDENCE`. System-category failure because a host returned malformed or request-inconsistent operation evidence. |
| `DurableEvidenceUndeliveredError` | Code `GRAPH_MERGE_OPERATION_UNDELIVERED`. `destroyDurableBranch()` was refused because committed operation evidence is still undelivered. Deliver or archive it first; the typed refusal is preserved so the evidence stays recoverable. |
| `MatchEvidenceError` | Evidence could not be constructed safely, including a custom scorer returning `NaN` or infinity. |
| `MergeError` | Any other merge failure (e.g. comparison-ceiling `"error"`, a non-transactional target). `MERGE_ERROR_CODES` enumerates the codes. |
## Example
See [FHIR Graph Merge](/examples/fhir-graph-merge) for a complete runnable
snapshot merge that reconciles two independently-extracted patient-care branches,
and [Incremental Merge](/examples/incremental-merge) for live-target ingestion
against an advancing base with persisted, queryable provenance.
# Operational Identity
> Assert, retract, query, and historize identity between graph nodes
The TypeGraph Identity Profile records identity facts between **individual
nodes**. It is deliberately smaller than OWL: `same` is symmetric and
transitive, `different` is symmetric and class-lifted, and neither relation
substitutes properties or automatically expands every graph query.
## Enable the profile
Identity is graph-level and opt-in:
```typescript
const graph = defineGraph({
id: "knowledge",
nodes: { Person: { type: Person }, Author: { type: Author } },
edges: {},
identity: { sameIdAcrossKinds: "fold" },
});
```
The option is serialized with the schema. Enabled graph types expose the full
facade as `store.identity` and `tx.identity`, a read-only facade as
`StoreView.identity`, and the identity traversal option. These surfaces use
conditional **presence**: on an identity-disabled graph type, `identity` does
not exist on `Store`, `TransactionContext`, or the read-only views at all —
reaching for it is a compile error, not a `never`-typed property. A runtime
`ConfigurationError` with details code `IDENTITY_NOT_ENABLED` backs those
getters too, for widened or `any`-typed handles TypeScript can't check (a
JavaScript caller, or a store handle that lost its precise graph type).
Constraint-aware `IngestionBranch` handles follow the same conditional-presence
contract, but expose only `assertSame`, `assertDifferent`, and their bulk forms.
Reads and retractions stay unavailable so untrusted batches can stage identity
evidence without gaining the full operational surface. See
[Constraint-aware ingestion branches](/graph-merge/#constraint-aware-ingestion-branches)
for the staging and merge workflow.
At runtime, a disabled graph does no identity work: no identity locks, probes,
closure computation, or identity SQL run. That guarantee is scoped to runtime
behavior — a bundled backend still provisions the identity tables' schema (not
work) when it bootstraps a fresh database, independent of whether the specific
graph passed to `createStore`/`createStoreWithSchema` declares `identity`.
`sameIdAcrossKinds: "fold"` preserves TypeGraph's structural ID rule: live
nodes of different kinds with the same ID belong to one identity class. No
assertion row is manufactured for that implicit membership. Use
`sameIdAcrossKinds: "ignore"` to enable the assertion ledger without joining
equal IDs across kinds; only explicit `same` assertions then join classes.
## Write and read identity
```typescript
const alice = await store.nodes.Person.create(
{ name: "Alice" },
{ id: "person-alice" },
);
const author = await store.nodes.Author.create(
{ penName: "A. Example" },
{ id: "author-alice" },
);
const result = await store.identity.assertSame(alice, author);
// result.action is "created" or "existing"; result.assertion is durable truth
await store.identity.membersOf(alice);
// [{ kind: "Author", id: "author-alice" },
// { kind: "Person", id: "person-alice" }]
await store.identity.representativeOf(alice);
await store.identity.nodesOf(alice); // hydrated, kind-discriminated nodes
await store.identity.areSame(alice, author);
await store.identity.assertionsOf(alice);
await store.identity.explainSame(alice, author);
const ended = await store.identity.retractAssertion(result.assertion.id);
// ended?.validTo is the exact assertion end instant
```
The complete write surface is:
- `assertSame(a, b)` and `assertDifferent(a, b)`
- `bulkAssertSame(pairs)` and `bulkAssertDifferent(pairs)`
- `retractAssertion(id)`
- `retractSameAssertion(a, b)` and `retractDifferentAssertion(a, b)`
- `bulkRetractAssertions(ids)`
Bulk methods are eager and, on PostgreSQL, run under one graph identity lock
(see [Operational notes](#operational-notes) — SQLite serializes through its
single-writer lock instead). `bulkAssertSame` and `bulkAssertDifferent`
preserve input order and return exactly one result per input pair. Reasserting
a current semantic pair is idempotent; assertion results distinguish
`action: "created"` from `action: "existing"`. Retraction methods return the
ended assertion (or `undefined` for a missing current assertion).
`bulkRetractAssertions` does **not** share that one-result-per-input shape: it
dedupes the input ids and returns only the assertions that were actually open,
in dense, first-occurrence input order — so the result does not align
index-by-index with the input array. Self-assertions are rejected.
Assertion IDs use the exported private-symbol-branded
`IdentityAssertionId` type so unrelated strings cannot be passed accidentally.
When you hold a plain assertion-ID string that came from persistence or an
interchange document, re-enter the branded type with the `asIdentityAssertionId(value)`
caster rather than a `as` assertion.
Assertions may state an explicit half-open validity window. Scalar methods take
the window as their third argument; bulk methods carry one window per pair:
```typescript
await store.identity.assertSame(alice, legacyAlice, {
validFrom: "2020-01-01T00:00:00.000Z",
validTo: "2022-01-01T00:00:00.000Z",
});
await store.identity.bulkAssertDifferent([
{ a: alice, b: bob, validFrom: "2023-01-01T00:00:00.000Z" },
{ a: alice, b: carol }, // ordinary current assertion semantics
]);
```
A past-ended window affects historical reads only. An open window beginning in
the past affects both historical and current reads. Repeating the exact
relation, pair, and window is idempotent. A second open window for an already
current semantic pair is refused rather than silently collapsed onto a
different `validFrom`. Empty objects retain the ordinary unwindowed semantics,
including inside a mixed bulk call.
Runtime-evolved nodes carry a nominal dynamic-node type, so they flow through
the same identity surface without a cast:
```typescript
const evolved = await store.evolve(extension);
const person = await evolved.nodes.Person.create({ name: "Alice" });
const tag = await evolved
.getNodeCollectionOrThrow("Tag")
.create({ label: "author" });
await evolved.identity.assertSame(person, tag);
await evolved.identity.membersOf(tag);
```
Reference reads return `IdentityNodeReference` values covering both
compile-time graph kinds and registered runtime kinds. This widening is
necessary even when a read starts from `person`, because its class can contain
`tag`. Their IDs retain the appropriate nominal brand, and `nodesOf` hydrates
the class into static kind-discriminated members or `DynamicNode` values for
runtime members. A plain `{ kind: string, id: string }` does not prove that the
kind came through the evolved Store; pass the dynamic node or a nominal dynamic
reference returned by an identity read. Unknown and removed kinds still fail
at runtime with `KindNotFoundError`. A missing, deleted, or
coordinate-invisible input returns `undefined`, `[]`, or `false` according to
the method. A visible singleton returns itself from `membersOf` and
`representativeOf`, and `areSame(ref, ref)` is true. `areDifferent` lifts an
explicit different assertion across both identity classes and also reflects
ontology `disjointWith` constraints. Representatives are deterministic: the
code-point-smallest `(kind, id)` visible member wins.
`explainSame(a, b)` returns a shortest path of persisted `same` assertions and
implicit same-ID folds connecting two visible references. Each step names its
endpoints and either the assertion or `type: "same-id-fold"`. It returns `[]`
for one visible reference and `undefined` when the references are distinct or
not visible at the read coordinate. Use `store.asOf(instant).identity` for a
historical explanation.
Historical identity reads and identity-expanded traversals use the kinds
registered on the current Store. Assertions involving a removed kind remain
in recorded history but no longer connect active classes.
`classes({ limit, kinds?, cursor? })` lists visible classes, including
singletons, in representative order. A kind filter selects classes containing
at least one visible member of the requested kinds; each result still includes
all of that class's visible members. Pass `nextCursor` to the next call until
it is absent. The cursor is exclusive and applies to the same graph, read
coordinate, and kind filter. When `kinds` is omitted, the scan uses the
registered runtime kinds present when each page is requested; adding a runtime
kind during that scan changes the filter and invalidates its cursor. At current
coordinates, the database finds visible representatives
for the page and expands members only for those classes; discovering
representatives still examines the visible node set. Historical coordinates
reconstruct all visible classes before applying the page boundary. For paging
across writes, use a recorded-time coordinate when recorded history is enabled:
valid-time `asOf` reads still observe later changes to the live tables.
## Integrity and lifecycle
Ordinary unwindowed assertions require live endpoints. Explicit windows require
both endpoint rows to cover the assertion's whole half-open interval; an ended
or late-starting endpoint raises `IdentityEndpointValidityError`. Future bounds
and inverted windows raise `IdentityValidityWindowError`. Zero-width windows
are accepted as empty history. Contradictions are checked throughout every
overlapping segment, including transitive `same` paths; adjacent half-open
windows do not overlap.
`assertSame` fails when a current
`different` assertion spans the two classes or when any member kinds are
ontology-disjoint. `assertDifferent` fails when both endpoints are already in
one class. These checks, folding, node deletion, import, schema-transition
validation, and closure rebuild share one per-graph lock and one mutation
coordinator.
Soft-deleting a node ends its current assertions. Hard-deleting it removes
every current and ended assertion touching the node from the live assertion
ledger; when recorded history is enabled, earlier recorded coordinates remain
queryable. On every graph, a `create()` or `upsertById()` for a soft-deleted
same-`(kind, id)` row **resurrects** that row rather than erroring:
its properties are replaced and its validity window is reset, so `validFrom`
becomes the resurrection instant — unless the write carries an explicit
window, which is honored as given (this is how merge preserves
branch-authored windows). A resurrecting node write that supplies only a
historical `validTo` takes the same **born-already-ended** exception a create
takes: no lower bound is stored ("ended at T, start unknown") rather than a
start after its own end, so the row reads back at every `asOf` before that end
and `meta.validFrom` is `undefined`. One stated window reaches one stored shape
whichever node path resets it — `create()` on a fresh id, `create()` on a
tombstone, or a resurrecting `upsertById()`. (Edge resurrection instead keeps
its stored lower bound, so `getOrCreateByEndpoints` can resurrect an edge
directly into the ended state — but the end it names is held to that retained
bound, so reviving an edge into a window that closed before the edge began is
refused as a `ValidationError`, and means passing both bounds.) This graph-wide
rule does not depend on the
identity profile. Resurrection does not revive ended assertions, but folding
runs again over the resurrected node when configured. Kind removal
cascades assertion and closure rows for the removed kinds. Tightening ontology
disjointness is rejected when it would make a persisted class contradictory.
`rebuildIdentityClosure(store)` repairs the derived current closure from live
nodes and current assertions. It validates integrity and never advances the
content revision. Schema-managed rebuilds, including automatic startup repair
of derived identity relations, pin the schema version used by the rebuild.
If a concurrent migration advances that version first, repair refuses with
`StaleVersionError` without overwriting the newer closure. Reopen using the
current graph definition before retrying.
### The database-level backstop
The checks above are code deciding whether a write is legal, and code can be
wrong. Underneath them TypeGraph maintains a second derived relation — the
**separation relation** — that holds one row per pair of identity classes a
current `different` assertion keeps apart, keyed by the two class keys under a
`CHECK (class_key_low < class_key_high)` constraint.
Every transaction that fuses two identity classes relabels the affected
separation rows in the same statement batch. Fusing two classes that were
separated relabels both sides of their shared row to one key, the constraint
rejects it, and the transaction aborts — in the engine, with no application
code in the way. A write that reaches the ledger through a path that skipped
identity validation therefore still cannot commit a contradictory graph; it
fails with an `IdentitySeparationViolationError` naming the `different`
assertion it contradicts.
Nothing about the identity API changes. The relation is derived and
maintained wherever the closure is, `rebuildIdentityClosure(store)` recomputes
it from the ledger, and store-open validation checks it against that
recomputation the same way it checks the closure.
## Temporal identity
Integrity is **structural**; reads are **coordinate-visible**.
Current reads use a materialized closure and then filter members through the
same visibility predicate ordinary node reads use. `store.identity` and
`store.asOf(now).identity` therefore agree.
Non-current valid-time and recorded-time views reconstruct one fixed point over
both explicit `same` assertions and same-ID folding edges. A structurally
existing but coordinate-invisible bridge can conduct identity without being
returned as a member. Recorded assertions are captured in the same commit as
the truth-bearing write.
The assertion's validity window and the commit that recorded it are independent
coordinates. A retrospective assertion is therefore invisible before its
recorded-time commit even when its valid-time window reaches farther into the
past. Archival export includes the endpoint temporal bounds needed to validate
those windows on import, and graph merge carries branch-authored bounded
assertions without turning them into current truth.
Identity profile and ontology rules are schema-level interpretation, not a
third temporal dimension. Historical views apply the Store's pinned
`sameIdAcrossKinds` mode and ontology to the assertions and nodes visible at the
requested coordinate. Changing those schema rules can therefore reinterpret
older coordinates; it does not rewrite the recorded assertion ledger.
```typescript
const before = await store.recordedNow();
const historical = store.asOfRecorded(before!);
await historical.identity.membersOf(alice);
```
### Folds and time
Implicit same-id folds (`sameIdAcrossKinds: "fold"`) conduct based on a node's
**lifecycle** — whether it currently exists and is not soft-deleted — not its
valid-time window. A node created today with a backdated `validFrom` is
valid-time visible in the past (an ordinary node read at that past coordinate
returns it), but it does not conduct a fold there: the fold only takes effect
once the node actually exists. Symmetrically, a node with a future `validFrom`
does not suppress its folds today — it already exists and is live, so it
folds now even though it is not yet valid-time visible. Explicit `same` and
`different` assertions are unaffected by this: they carry their own validity
windows and conduct exactly when they are current. This keeps the fold
computation tied to write events rather than to valid-time windows, so the
materialized closure used by current reads and by `asOf(now)` reads is
identical — a fixed-point reconstruction of "current" never needs to
special-case valid-time skew on the folding edge itself.
## Identity-expanded traversal
Traversal expansion is per hop and defaults off:
```typescript
const results = await store
.query()
.from("Person", "person")
.traverse("authored", "edge", { includeIdentityMembers: true })
.to("Document", "document")
.select((ctx) => ({ edge: ctx.edge, document: ctx.document }))
.execute();
```
The hop considers coordinate-visible members of the source class, returns the
physical edge and target rows, preserves their provenance, and deduplicates a
physical edge within the step — with one legitimate exception: a self-inverse
edge (`inverseOf(edgeKind, edgeKind)`) traversed with `expand` between two
identity-folded peers can yield the same physical edge twice, once per
direction/target it matches through the fold. That is not a dedup bug; the
edge genuinely satisfies the traversal from both of its endpoints. Recursive
traversal supports the same option.
TypeGraph does not perform automatic graph-wide expansion and collection reads
such as `getById` have no identity option.
Both coordinates reach the candidate edge the same way — an ordinary indexed
equality on the class member, never a membership test evaluated per candidate
edge. How each one reaches the class differs, because what a class costs to
compute differs.
At the **current** coordinate the maintained closure already *is* the class
relation, so each traversal step seeks into it from its own frontier rows: the
frontier row's class through the closure's primary key, that class's members
through the class index, each member's node for its visibility. Cost is
proportional to the frontier and the size of its classes — never to how many
identity classes the graph holds. Measured on SQLite with *n* Person nodes, each
folded with a Company and an Alias peer sharing its id (a three-member class per
source), all *n* acting as source rows and every edge leaving the Company peer:
| source rows | fan-out | matching edges | before | after |
| --- | --- | --- | --- | --- |
| 250 | 1 | 250 | 67 ms | 6 ms |
| 1000 | 1 | 1000 | 1077 ms | 9 ms |
| 2000 | 1 | 2000 | 4616 ms | 19 ms |
| 1000 | 8 | 8000 | 8611 ms | 13 ms |
| 500 | 200 | 100,000 | 51,602 ms | 77 ms |
Growth is linear in graph size where it used to quadruple per doubling: the hop
no longer evaluates membership per candidate *(source row, edge)* pair. The
number to plan around is the last row — a hundred thousand matching edges over a
five-hundred-row frontier is where the old per-source rescan dominated.
A **historical** hop — one under `asOf`, `asOfRecorded`, or a non-current
`view()` — cannot use the materialized closure, because the closure represents
only the present. Its rows come from a reconstruction of identity classes out of
the assertion ledger, and under `sameIdAcrossKinds: "fold"` that reconstruction
also has to consider the structural same-id relation, which is proportional to
the number of live nodes in the graph. No frontier row narrows that fixed point,
so it is built once per statement into a materialized relation every traversal
step joins. Measured on the narrow-edge fixture that isolates the term (SQLite,
*n* Person nodes each folded with a Company peer, all *n* acting as source rows,
fan-out 1):
| *n* | before | after |
| --- | --- | --- |
| 250 | 122 ms | 7 ms |
| 500 | 486 ms | 7 ms |
| 1000 | 1984 ms | 14 ms |
| 2000 | 8261 ms | 28 ms |
Growth is linear in graph size where it used to quadruple per doubling.
The caveat that remains is the historical one, and it is worth planning around: a
past-coordinate hop rebuilds the whole graph's classes even when you asked about
one node, so its floor is a pass over the identity population regardless of how
narrow the frontier is. A **current** hop has no such floor — a single-start-row
hop over 50,000 folded triples measures 1 ms on SQLite against 387 ms when the
class relation was still built graph-wide, and nine unrelated 501-member classes
cost it nothing at all (0.5 ms on SQLite, 2.4 ms on PostgreSQL, against 564 ms
and 568 ms). Pick the coordinate you actually need: reading the present is the
cheaper question by a wide margin.
## Interchange and branch merge
Interchange format `2.0` optionally carries an identity section. State export
(the default) includes current assertions. Import into a populated target is
target-oriented: an existing current semantic pair keeps its target assertion
ID and `validFrom`. Working-copy branch cloning imports into an empty target and
preserves source IDs and `validFrom` exactly.
```typescript
const state = await exportGraph(store, { includeTemporal: true });
const archive = await exportGraph(store, {
identityMode: "archival",
includeDeleted: true,
});
```
Identity-enabled exports default `includeTemporal` to `true`, because importing
identity truth must prove that both endpoints existed throughout each assertion
window. Explicitly setting `includeTemporal: false` on an identity-enabled graph
is refused.
Archival mode also includes ended assertions. Those rows are restored after
shape validation and do not affect current closure. An ending a node deletion
caused carries that node as `endedBy`, so a round-trip preserves why each
assertion ended and not merely that it did; import rejects an `endedBy` on an
open assertion, or one naming a node that is not an endpoint of the assertion
it ends. Ended assertions can
reference soft-deleted nodes, and by default (`includeDeleted: false`) export
joins every assertion against its endpoints' live rows — an assertion with a
soft-deleted endpoint is silently **dropped from the export entirely**, not
carried with a dangling reference. Pair `identityMode: "archival"` with
`includeDeleted: true` to keep those assertions in the archive. Interchange
documents carry no `deletedAt` field, so a node exported only because of
`includeDeleted: true` re-imports as **live** — an `includeDeleted` archive
resurrects its soft-deleted nodes on import rather than restoring them as
deleted. Weigh that trade-off deliberately for a backup: without
`includeDeleted`, soft-deleted endpoints and the assertions that reference them
are silently absent; with it, those nodes come back alive. Recorded side
tables are not part of interchange.
Graph merge includes identity truth in staleness fingerprints and diffs.
Duplicate current assertions use the earliest `validFrom`, then the
code-point-smallest assertion ID — unless one candidate is already committed
on the target with the exact staged truth, which always wins: the applier is
idempotent per semantic pair, so a challenger could never actually be
written. A node deletion cascades into ending the assertions touching it, at
the node's own deletion instant, and records the deleted node on every row it
ends — so the diff reads which endings that deletion caused and stages each
one with its cause, however close in time the branch's own retractions
fell. When a delete/modify
conflict resolution keeps the node, an ending is dropped along with the
overruled deletion that caused it (reported as
`identity:deletion-overruled`), while a retraction a branch made itself
survives the deletion being overruled — including one the deleting branch
made before deleting the node, even in the deletion's own millisecond. A hard
delete removes the assertion rows outright, taking the recorded cause with
them and leaving nothing to separate cause from intent, so those endings count
as cascades. `merge()` detects identity conflicts at plan
time and returns them as a typed `IdentityMergeConflictError` — direct
opposing relations on one endpoint pair, transitive contradictions reached
through a chain of `same` assertions no single branch wrote, retract/reassert
races, and an assertion over a node another branch deleted. A branch that
retracts a pair and also reasserts it itself (convergent, not racing) merges
cleanly. This is mechanical truth propagation, not semantic entity
reconciliation. Plan time is the early surface, not the only one: any
identity refusal that still escapes to the applier inside the commit
transaction is translated into the same typed `IdentityMergeConflictError`,
with the original error preserved as its cause (identity environment and
storage-corruption codes pass through untranslated — they are not statements
about merge truth). See
[`IdentityMergeConflictError`](/errors/#identitymergeconflicterror)
for the exact `merge()` signature and how to catch it.
### Independent targets and assertion IDs
`mergeIncremental()` accepts a target that has moved on from the branches'
fork point, so a branch's assertion IDs can meet a ledger that assigned those
IDs independently. Snapshot `merge()` still requires its target to match the
branches' base@V exactly, but the same by-ID contract governs the divergence
a branch can create within its own lineage (hard-delete/recreate replacement)
and the plan→commit window. The contract is by ID, on complete truth:
- **One assertion ID, one complete truth.** A planned assertion whose ID the
target's ledger — ended rows included — already binds to a different
complete truth (relation, endpoints, validity) refuses at plan time as
`IdentityMergeConflictError`. An exact match is applied idempotently.
- **Retractions carry the truth they retract.** A branch retraction ends the
target's current row for its ID only when that row *is* the truth the
branch retracted. When the target reuses the ID for different truth, the
retraction is skipped and reported in `MergeReport.dropped` as
`identity:retraction-target-mismatch` — the branch's own assertion is
already absent from the target, and ending the target's unrelated row
would delete truth the branch never saw.
- **Truth replacement is a conflict, not a silent keep.** Within one lineage
a branch can legally rebind an assertion ID by hard-deleting an endpoint
(which physically removes the row) and importing the ID for different
truth. The diff stages that replacement as a retraction plus a new
assertion; because the target's ledger still holds the ID's prior truth in
an ended row, the plan-time one-ID-one-truth check refuses it typed rather
than silently keeping either side's truth.
- **The commit re-verifies IDs.** Both commit modes re-read every planned
assertion and retraction ID inside the commit transaction and refuse
plan→commit drift as `BaseVersionMismatchError` — retrying recomputes the
plan from current state. One deliberate exception: a planned retraction
whose row another writer already ended is accepted as a no-op, not drift.
`MergeReport.merged.identity` reports the rows the applier actually
created and ended; idempotent skips are excluded.
- **The commit proves the result, not the plan.** After its identity writes,
and still inside the same transaction, a merge re-derives the identity
classes it touched from the written state and refuses a contradiction there
as `IdentityMergeConflictError` — so a plan validated against state that has
since moved cannot leave a contradictory ledger behind. The whole merge rolls
back; there is no partial commit. If the derived classes disagree with the
materialized closure, the closure is rebuilt inside the same transaction and
the check re-runs, which repairs a lagging closure atomically with a merge
that is otherwise sound.
## Operational notes
On PostgreSQL, every identity-affecting node write on an identity-enabled graph
serializes on a per-graph advisory transaction lock: at most one writer per
graph proceeds at a time. This is a correctness guarantee for the assertion
ledger and closure, and it is also a throughput ceiling — concurrent writers to
the same graph queue behind the lock. Writes to other graphs, and all reads,
are unaffected.
First-time enablement is heavier than steady state. It takes a `SHARE` lock on
the shared nodes table, which briefly blocks writes for **every** graph in that
database, and it loads the whole graph to build the initial identity closure.
Plan enablement for a quiet window on large databases. `evolve()` on an
identity-enabled graph re-runs the same closure rebuild, so schema evolution
carries a comparable one-time cost proportional to graph size.
Changing `sameIdAcrossKinds` is a **breaking** schema change — a `fold`↔`ignore`
flip rewrites the materialized identity closure and changes every
`areSame`/`membersOf`/`includeIdentityMembers` answer against existing data —
so it requires the same explicit `migrateSchema()` opt-in as any other
breaking change; it never auto-migrates silently. Identity-relevant ontology
changes (`disjointWith`, `equivalentTo`/deprecated `sameAs`, or `subClassOf`)
are likewise persisted semantic migrations, not a local runtime toggle.
`createStoreWithSchema` and explicit `migrateSchema()` both rebuild and
validate the closure atomically with the schema commit that carries the
change. While the flip is unapplied, store construction refuses with
`ConfigurationError` details code `IDENTITY_PROFILE_MIGRATION_PENDING`
whenever the identity change is the only breaking one in the diff; a
migration that also breaks other schema surfaces raises the generic
`MigrationError` enumerating everything. First-time identity *enablement* (`autoMigrate: false` on a graph
newly declaring `identity: { ... }`) is a safe, additive change, and
`createStoreWithSchema` refuses to return a Store while it is pending with
`ConfigurationError` details code `IDENTITY_ENABLEMENT_PENDING`. The very
first schema commit of an identity-enabled graph is an enablement too: a
legacy database populated through an unmanaged `createStore` gets the same
atomic fold scan, contradiction validation, and closure build during
initialization — an empty database just makes them cheap no-ops.
## Migrating from type-level factories
The ontology factories `sameAs(A, B)` and `differentFrom(A, B)` are deprecated:
they relate **types**, not individual rows, and `differentFrom` never enforced
instance identity. To migrate:
1. Add `identity: { sameIdAcrossKinds: "fold" }` to the graph.
2. Open it with `createStoreWithSchema` so the capability is persisted and
existing cross-kind same-ID groups are validated and materialized.
3. Replace type-level facts with `store.identity` assertions between concrete
node references.
4. Use `equivalentTo` or `disjointWith` when the intended relation is genuinely
between kinds.
On PostgreSQL, first-time enablement waits for in-flight node writes before it
builds the initial identity closure. Quiesce or restart any store instances
that were opened with the identity-disabled schema before allowing writes to
resume; stale instances do not participate in identity locking.
Identity requires interactive atomic transactions. Bundled SQLite and
PostgreSQL drivers support it; Cloudflare D1 and `drizzle-orm/neon-http` reject
an enabled graph with `ConfigurationError` details code
`IDENTITY_REQUIRES_ATOMIC_BACKEND`. Identity-disabled graphs continue to work
on those drivers.
Durable entity handles, identity-group IDs, semantic reconciliation, automatic
OWL property substitution, and graph-wide identity expansion are reserved
future capabilities and are not implied by this profile.
# Integration Patterns
> Strategies for integrating TypeGraph into your application architecture
This guide covers common integration patterns for adding TypeGraph to existing
applications, from simple setups to production deployment strategies.
## Direct Drizzle Integration (Shared Database)
If you're already using Drizzle ORM, TypeGraph can share your existing database
connection. TypeGraph tables coexist alongside your application tables.
```typescript
import { drizzle } from "drizzle-orm/node-postgres";
import { Pool } from "pg";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStore } from "@nicia-ai/typegraph";
// Your existing Drizzle setup
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
const db = drizzle(pool);
// Add TypeGraph tables to your existing database
await pool.query(generatePostgresMigrationSQL());
// Create TypeGraph backend using the same connection
const backend = createPostgresBackend(db);
const store = createStore(graph, backend);
// For pure TypeGraph operations, use store.transaction()
await store.transaction(async (tx) => {
const person = await tx.nodes.Person.create({ name: "Alice" });
const company = await tx.nodes.Company.create({ name: "Acme" });
await tx.edges.worksAt.create(person, company, { role: "Engineer" });
});
```
### Mixed Drizzle + TypeGraph Transactions
When combining TypeGraph operations with direct Drizzle queries in the same atomic transaction,
create a temporary backend from the Drizzle transaction:
```typescript
await db.transaction(async (tx) => {
// Direct Drizzle operations
await tx.insert(auditLog).values({ action: "user_created" });
// TypeGraph operations in the same transaction
const txBackend = createPostgresBackend(tx);
const txStore = createStore(graph, txBackend);
await txStore.nodes.Person.create({ name: "Alice" });
});
```
This pattern is only needed when you must combine both in one atomic transaction.
**When to use:**
- You want a single database to manage
- Your graph data relates to existing tables
- You need cross-cutting transactions
**Considerations:**
- TypeGraph tables use the `typegraph_` prefix to avoid collisions
- Run TypeGraph migrations alongside your application migrations
- Connection pool is shared, so size accordingly
## Drizzle-Kit Managed Migrations (Recommended)
If you use `drizzle-kit` to manage migrations, you can import TypeGraph's table
definitions directly into your schema file. This lets drizzle-kit generate
migrations for all tables—both yours and TypeGraph's—in one place.
### Setup
**1. Import TypeGraph tables into your schema:**
```typescript
// schema.ts
import { sqliteTable, text, integer } from "drizzle-orm/sqlite-core";
// Import TypeGraph tables (these are standard Drizzle table definitions)
export * from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// Or for PostgreSQL:
// export * from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Your application tables
export const users = sqliteTable("users", {
id: text("id").primaryKey(),
name: text("name").notNull(),
email: text("email").notNull(),
});
```
**2. Generate migrations normally:**
```bash
npx drizzle-kit generate
```
Drizzle-kit will now see all tables—TypeGraph's and yours—and generate migrations
for them.
**3. Apply migrations:**
```bash
npx drizzle-kit migrate
# Or for Cloudflare D1:
wrangler d1 migrations apply your-database
```
**4. Create the backend:**
```typescript
import { drizzle } from "drizzle-orm/better-sqlite3";
import Database from "better-sqlite3";
import { createSqliteBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { createStore } from "@nicia-ai/typegraph";
const sqlite = new Database("app.db");
const db = drizzle(sqlite);
// Use the same tables that drizzle-kit manages
const backend = createSqliteBackend(db, { tables });
const store = createStore(graph, backend);
```
### Custom Table Names
To avoid conflicts or match your naming conventions, use the factory function:
```typescript
// schema.ts
import { createSqliteTables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// Create tables with custom names
export const typegraphTables = createSqliteTables({
nodes: "myapp_graph_nodes",
edges: "myapp_graph_edges",
uniques: "myapp_graph_uniques",
schemaVersions: "myapp_graph_schema_versions",
embeddings: "myapp_graph_embeddings",
fulltext: "myapp_graph_fulltext",
indexMaterializations: "myapp_graph_index_materializations",
kindRemovals: "myapp_graph_kind_removals",
reconciliationMarkers: "myapp_graph_reconciliation_markers",
});
// Export individual tables for drizzle-kit
export const { nodes: myappGraphNodes, edges: myappGraphEdges, uniques: myappGraphUniques, schemaVersions: myappGraphSchemaVersions, embeddings: myappGraphEmbeddings, indexMaterializations: myappGraphIndexMaterializations, kindRemovals: myappGraphKindRemovals, reconciliationMarkers: myappGraphReconciliationMarkers } = typegraphTables;
// SQLite fulltext is an FTS5 virtual table — drizzle-kit can't model
// virtual tables, so this name is exposed as a string. The backend
// creates the FTS5 table on first store boot via a focused
// `ensureFulltextTable()` ensure (idempotent CREATE VIRTUAL TABLE
// IF NOT EXISTS), so drizzle-kit-managed setups work without an
// extra manual step.
export const myappGraphFulltextTableName = typegraphTables.fulltextTableName;
```
For PostgreSQL with the default `tsvectorStrategy`, the factory
**does** return a typed Drizzle table — `tables.fulltext` — alongside
the others, so drizzle-kit-managed setups pick up the fulltext table
automatically:
```typescript
// schema.ts (PostgreSQL)
import { createPostgresTables } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
export const typegraphTables = createPostgresTables({
// …same names as above…
});
export const {
nodes: myappGraphNodes,
edges: myappGraphEdges,
// …
fulltext: myappGraphFulltext,
indexMaterializations: myappGraphIndexMaterializations,
// …
} = typegraphTables;
```
If you swap in an alternate Postgres fulltext strategy (pg_trgm,
ParadeDB / pg_search, pgroonga), the typed `tsvector`-shaped table
won't match what your strategy needs. Override `tables.fulltext` in
your schema barrel with your strategy's own Drizzle table, or skip
the typed export and rely on the backend's runtime
`ensureFulltextTable()` ensure to bootstrap your strategy's DDL.
Then pass the same tables to the backend:
```typescript
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { typegraphTables } from "./schema";
const backend = createSqliteBackend(db, { tables: typegraphTables });
```
### Adding TypeGraph Indexes
The table factory functions also accept `indexes`, which drizzle-kit will include in migrations:
```ts
// schema.ts
import { createSqliteTables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { defineNodeIndex } from "@nicia-ai/typegraph/indexes";
import { Person } from "./graph";
const personEmail = defineNodeIndex(Person, { fields: ["email"] });
export const typegraphTables = createSqliteTables({}, { indexes: [personEmail] });
```
For PostgreSQL, use `createPostgresTables` from `@nicia-ai/typegraph/adapters/drizzle/postgres`.
See [Indexes](/performance/indexes) for covering fields, partial indexes, and profiler integration.
Beyond accelerating queries, a declared index powers `store.nodes..bulkFindByIndex(indexName, items)` — a
batched lookup that returns, for each incoming record, the live nodes sharing its index key. This is the primitive for
**import reconciliation** and **dedup-candidate discovery**: probe an entire import batch against the graph in one
query to decide create-vs-merge per record (the key may be non-unique, so each record yields its own candidate list).
See the [batched index lookup reference](/performance/indexes#batched-index-lookup-bulkfindbyindex) and the runnable
[`examples/17-bulk-find-by-index.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/17-bulk-find-by-index.ts).
If you only need PostgreSQL adapter exports, import from `@nicia-ai/typegraph/adapters/drizzle/postgres`:
```typescript
import { createPostgresBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
```
### PostgreSQL with pgvector
For PostgreSQL with vector search, ensure the pgvector extension is enabled
before running migrations:
```sql
CREATE EXTENSION IF NOT EXISTS vector;
```
When multiple allocations share one PostgreSQL database, give each backend a
stable namespace so its pgvector tables and indexes remain physically isolated:
```typescript
import { createPgvectorStrategy } from "@nicia-ai/typegraph";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const backend = createPostgresBackend(pool, {
vector: createPgvectorStrategy("tenant-a"),
});
```
Keep the namespace stable for the lifetime of the allocation. The default
`pgvectorStrategy` continues to use the existing `tg_vec` / `tg_vecidx` names.
Then in your schema:
```typescript
// schema.ts
export * from "@nicia-ai/typegraph/adapters/drizzle/postgres";
export const users = pgTable("users", { ... });
```
**When to use:**
- You already use drizzle-kit for migrations
- You want a single migration workflow for all tables
- You need Cloudflare D1 or other platforms that require drizzle-kit migrations
**Advantages over raw SQL migrations:**
- Single source of truth for schema
- Type-safe schema in TypeScript
- Drizzle-kit handles migration diffs automatically
- Works with all drizzle-kit supported platforms
## Separate Database
Use a dedicated database when you want isolation between your application data
and graph data.
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Application database (your existing setup)
const appPool = new Pool({ connectionString: process.env.APP_DATABASE_URL });
const appDb = drizzle(appPool);
// Dedicated TypeGraph database
const graphPool = new Pool({ connectionString: process.env.GRAPH_DATABASE_URL });
const graphDb = drizzle(graphPool);
await graphPool.query(generatePostgresMigrationSQL());
const backend = createPostgresBackend(graphDb);
const store = createStore(graph, backend);
```
**When to use:**
- Your primary database doesn't support required features (e.g., pgvector)
- You want independent scaling for graph operations
- Compliance requires data separation
- You're adding graph capabilities to a legacy system
**Considerations:**
- No cross-database transactions (use eventual consistency patterns)
- Sync data between databases via application logic or events
- Separate backup/restore procedures
## In-Memory (Ephemeral Graphs)
Use in-memory SQLite for temporary graphs, caching, or computation.
```typescript
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
function createEphemeralStore(graph: GraphDef) {
const { backend } = createLocalSqliteBackend();
return createStore(graph, backend);
}
// Use case: Build a temporary graph for computation
async function computeRecommendations(userId: string): Promise {
const tempStore = createEphemeralStore(recommendationGraph);
// Load relevant data into temporary graph
const userData = await fetchUserData(userId);
await populateGraph(tempStore, userData);
// Run graph algorithms
const results = await tempStore
.query()
.from("User", "u")
.whereNode("u", (u) => u.id.eq(userId))
.traverse("similar", "s")
.to("Product", "p")
.select((ctx) => ctx.p)
.execute();
return results;
}
```
**When to use:**
- Temporary computation graphs
- Request-scoped graph state
- Graph-based caching with expiration
- Isolated test fixtures
**Considerations:**
- Data lost on process termination
- Memory usage scales with graph size
- No persistence—rebuild on restart
## Hybrid Overlay (Graph on Existing Data)
Add graph relationships on top of existing relational data without migrating
your data model. Your existing tables remain the source of truth; TypeGraph
stores only the relationships and graph-specific metadata.
Use the `externalRef()` helper to create type-safe references to external tables:
```typescript
import { createExternalRef, defineEdge, defineGraph, defineNode, embedding, externalRef } from "@nicia-ai/typegraph";
import { z } from "zod";
// Define nodes that reference your existing tables
const User = defineNode("User", {
schema: z.object({
// Type-safe reference to your existing users table
source: externalRef("users"),
// Denormalized fields for graph queries (optional)
displayName: z.string().optional(),
}),
});
const Document = defineNode("Document", {
schema: z.object({
source: externalRef("documents"),
embedding: embedding(1536).optional(),
}),
});
// Graph-only relationships not in your relational schema
const relatedTo = defineEdge("relatedTo", {
schema: z.object({
relationship: z.enum(["cites", "extends", "contradicts"]),
confidence: z.number().min(0).max(1),
}),
});
const authored = defineEdge("authored");
const graph = defineGraph({
id: "document_graph",
nodes: { User, Document },
edges: {
relatedTo: { type: relatedTo, from: [Document], to: [Document] },
authored: { type: authored, from: [User], to: [Document] },
},
});
```
The `externalRef()` helper validates that references include both the table name
and ID, catching errors at insert time:
```typescript
// Valid: includes table and id
await store.nodes.Document.create({
source: { table: "documents", id: "doc_123" },
});
// Error: wrong table name (caught by TypeScript and runtime validation)
await store.nodes.Document.create({
source: { table: "users", id: "doc_123" }, // Type error!
});
// Use createExternalRef() for a cleaner API
const docRef = createExternalRef("documents");
await store.nodes.Document.create({
source: docRef("doc_456"),
});
```
**Syncing with external data:**
```typescript
// Sync helper: Create or update graph node from app data
async function syncDocument(store: Store, appDocument: AppDocument) {
const existing = await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.source.get("id").eq(appDocument.id))
.select((ctx) => ctx.d)
.first();
if (existing) {
await store.nodes.Document.update(existing.id, {
embedding: await generateEmbedding(appDocument.content),
});
return existing;
}
return store.nodes.Document.create({
source: { table: "documents", id: appDocument.id },
embedding: await generateEmbedding(appDocument.content),
});
}
// Query combining graph traversal with app data hydration
async function findRelatedDocuments(documentId: string) {
// Get graph relationships
const related = await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.source.get("id").eq(documentId))
.traverse("relatedTo", "r")
.to("Document", "related")
.select((ctx) => ({
source: ctx.related.source,
relationship: ctx.r.relationship,
confidence: ctx.r.confidence,
}))
.execute();
// Hydrate with full data from app database
const externalIds = related.map((r) => r.source.id);
const fullDocuments = await appDb.select().from(documents).where(inArray(documents.id, externalIds));
return related.map((r) => ({
...r,
document: fullDocuments.find((d) => d.id === r.source.id),
}));
}
```
**When to use:**
- Adding graph capabilities to an existing application
- Semantic search over existing content
- Relationship discovery without schema changes
- Gradual migration from relational to graph thinking
**Considerations:**
- Maintain sync between app data and graph nodes
- Decide what to denormalize (tradeoff: query speed vs. sync complexity)
- The `table` field in `externalRef` enables referencing multiple external sources
## Background Embedding Workers
Decouple embedding generation from request handling using background jobs.
```typescript
// job-queue.ts - Define the embedding job
interface EmbeddingJob {
nodeType: string;
nodeId: string;
content: string;
}
// worker.ts - Process embedding jobs
import { createStore } from "@nicia-ai/typegraph";
async function processEmbeddingJob(job: EmbeddingJob) {
const { nodeType, nodeId, content } = job;
// Generate embedding (expensive operation)
const embedding = await openai.embeddings.create({
model: "text-embedding-ada-002",
input: content,
});
// Update the node
const collection = store.nodes[nodeType as keyof typeof store.nodes];
await collection.update(nodeId, {
embedding: embedding.data[0].embedding,
});
}
// api-handler.ts - Enqueue jobs on create/update
async function createDocument(data: DocumentInput) {
// Create node without embedding (fast)
const doc = await store.nodes.Document.create({
title: data.title,
content: data.content,
// embedding: undefined - will be populated by worker
});
// Enqueue embedding job (non-blocking)
await jobQueue.add("generate-embedding", {
nodeType: "Document",
nodeId: doc.id,
content: data.content,
});
return doc;
}
```
**Batch processing for bulk imports:**
```typescript
async function backfillEmbeddings(batchSize = 100) {
let processed = 0;
while (true) {
// Find nodes missing embeddings
const nodes = await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.isNull())
.select((ctx) => ({
id: ctx.d.id,
content: ctx.d.content,
}))
.limit(batchSize)
.execute();
if (nodes.length === 0) break;
// Batch embed
const embeddings = await openai.embeddings.create({
model: "text-embedding-ada-002",
input: nodes.map((n) => n.content),
});
// Batch update
await store.transaction(async (tx) => {
for (const [i, node] of nodes.entries()) {
await tx.nodes.Document.update(node.id, {
embedding: embeddings.data[i].embedding,
});
}
});
processed += nodes.length;
console.log(`Processed ${processed} documents`);
}
}
```
**When to use:**
- Embedding generation is slow (100-500ms per call)
- You want fast API response times
- Bulk importing existing content
- Retry logic for API failures
**Considerations:**
- Handle job failures and retries
- Consider rate limits on embedding APIs
- Queries on `embedding` should handle null values during population
## Testing
For test setup patterns, seed data strategies, and profiler-based index coverage checks,
see the dedicated [Testing](/testing) guide.
## Deployment Patterns
### Edge and Serverless
Deploy TypeGraph at the edge using SQLite-compatible runtimes.
> **Note:** Edge environments cannot use `@nicia-ai/typegraph/adapters/drizzle/sqlite/local`
> because it depends on `better-sqlite3`, a native Node.js addon. Instead, use
> `@nicia-ai/typegraph/adapters/drizzle/sqlite` which is driver-agnostic.
**Cloudflare Durable Objects (SQLite) — transactional:**
```typescript
import { drizzle } from "drizzle-orm/durable-sqlite";
import { createAdapterStoreWithSchema } from "@nicia-ai/typegraph";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
export class GraphObject {
constructor(private ctx: DurableObjectState) {}
async fetch(): Promise {
const db = drizzle(this.ctx.storage);
const backend = createSqliteBackend(db); // auto-detects "do-sqlite"
const [store] = await createAdapterStoreWithSchema(graph, backend);
// Atomic across TypeGraph + the caller's own relational tables:
await store.transaction(async (tx) => {
await tx.nodes.Document.update(documentId, props);
if (tx.sqlAvailability !== "available") {
throw new Error(`Native transaction unavailable: ${tx.sqlAvailability}`);
}
const sqlTx = tx.sql;
await sqlTx.insert(documentVersions).values(versionRow);
});
return new Response("ok");
}
}
```
Unlike D1, Durable Objects expose an interactive storage transaction runner,
so `store.transaction()` / `store.withTransaction()` are fully atomic
(`capabilities.execution.interactiveTransactions: true`). See
[Backend Setup](/backend-setup#cloudflare-durable-objects-sqlite) and the
[Cross-Store Transactions recipe](/recipes#cross-store-transactions-drizzle--typegraph).
**Cloudflare Workers with D1:**
```typescript
// worker.ts
import { drizzle } from "drizzle-orm/d1";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
export default {
async fetch(request: Request, env: Env): Promise {
const db = drizzle(env.DB);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
// Handle request with graph queries
const results = await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 5))
.select((ctx) => ctx.d)
.execute();
return Response.json(results);
},
};
```
**Turso (libSQL):**
```typescript
import { createClient } from "@libsql/client";
import { drizzle } from "drizzle-orm/libsql";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
const client = createClient({
url: process.env.TURSO_DATABASE_URL!,
authToken: process.env.TURSO_AUTH_TOKEN,
});
const db = drizzle(client);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
```
> For Turso and D1, use [drizzle-kit managed migrations](#drizzle-kit-managed-migrations-recommended)
> to set up the schema.
**Bun with built-in SQLite:**
Bun runs locally, so you can use the Node.js-compatible path with better-sqlite3, or
use bun:sqlite with drizzle-kit managed migrations:
```typescript
import { Database } from "bun:sqlite";
import { drizzle } from "drizzle-orm/bun-sqlite";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
const sqlite = new Database("app.db");
const db = drizzle(sqlite);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
```
> Use [drizzle-kit managed migrations](#drizzle-kit-managed-migrations-recommended)
> to set up the schema with bun:sqlite.
**When to use:**
- Low-latency requirements (data close to users)
- Serverless functions with graph queries
- Read-heavy workloads
**Considerations:**
- SQLite limitations (single-writer, no pgvector)
- Cold start times include DB initialization
- Vector search (cosine/L2): sqlite-vec on the local better-sqlite3 backend;
libSQL's built-in vectors on the libSQL / Turso backend
### Per-Request Connections (Cache the Verified Store)
Some serverless Postgres setups — Cloudflare Workers behind Hyperdrive, or any
platform that pools connections for you — want a **fresh connection per
request**. `createVerifiedAdapterStore` reconciles the committed schema and
checks index materialization at open time (a few `SELECT`s), so re-opening a
verified store on every request adds that cost to every graph-backed route.
Verify **once per isolate**, then build a zero-query store per request from the
cached reconciled schema:
```typescript
import { createAdapterStore, createVerifiedAdapterStore, getCommittedSchemaVersion, type GraphBackend, type ReconciledSchema } from "@nicia-ai/typegraph";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
type Cached = {
reconciled: ReconciledSchema;
version: number | undefined;
};
// Per-isolate cache, plus the in-flight reconciliation. Memoizing the promise
// collapses concurrent cold (or stale) requests onto ONE verify instead of each
// running its own — otherwise a burst of first requests reproduces the fan-out
// stampede this pattern exists to avoid.
let cached: Cached | undefined;
let inFlight: Promise | undefined;
function reconcileOnce(verifyBackend: GraphBackend): Promise {
inFlight ??= (async () => {
const [store, result] = await createVerifiedAdapterStore(graph, verifyBackend);
cached = {
reconciled: store.reconciledSchema,
version: result.status === "unchanged" ? result.version : store.reconciledSchema.version,
};
return cached;
})().finally(() => {
inFlight = undefined;
});
return inFlight;
}
export default {
async fetch(request: Request, env: Env): Promise {
const backend = createPostgresBackend(newPoolForThisRequest(env));
// Cold start: concurrent first requests all await the same reconciliation.
let snapshot = cached ?? (await reconcileOnce(backend));
// A one-row probe detects a schema commit from another isolate; a moved
// version refreshes through the same single-flight path.
const committed = await getCommittedSchemaVersion(backend, graph.id);
if (committed !== snapshot.version) snapshot = await reconcileOnce(backend);
// Zero database round-trips. Reads and writes still validate against
// runtime-committed kinds carried by the reconciled snapshot.
const store = createAdapterStore(graph, backend, { reconciled: snapshot.reconciled });
const results = await store
.query()
.from("Document", "d")
.select((ctx) => ctx.d)
.execute();
return Response.json(results);
},
};
```
`store.reconciledSchema` is an opaque snapshot of the reconciled graph
(compile-time kinds folded with any runtime-committed kinds) plus the committed
version it reflects. `createAdapterStore(graph, backend, { reconciled })` issues
**no** queries and validates writes against that snapshot, so kinds committed at
runtime remain writable without re-verifying. If you already hold a verified
store and only need to swap the connection, `store.withBackend(freshBackend)`
returns an equivalent store bound to the new connection with no re-verify.
The `getCommittedSchemaVersion` probe is your read-your-writes seam: one
round-trip, far cheaper than the full verified open (which also reconciles the
schema and checks index materialization), and re-verify only fires when the
version actually moves. Skip the probe only if your schema changes exclusively
during a deployment that also clears the cache. Otherwise reads may use the
stale schema snapshot, and the write fence rejects managed writes until the
cache is refreshed.
### Read Replica Separation
Route heavy graph queries to read replicas while writes go to primary.
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Primary for writes
const primaryPool = new Pool({
connectionString: process.env.PRIMARY_DATABASE_URL,
max: 10,
});
const primaryDb = drizzle(primaryPool);
const primaryBackend = createPostgresBackend(primaryDb);
const primaryStore = createStore(graph, primaryBackend);
// Replica for reads
const replicaPool = new Pool({
connectionString: process.env.REPLICA_DATABASE_URL,
max: 50, // Higher pool for read-heavy workloads
});
const replicaDb = drizzle(replicaPool);
const replicaBackend = createPostgresBackend(replicaDb);
const replicaStore = createStore(graph, replicaBackend);
// Route based on operation
export const stores = {
write: primaryStore,
read: replicaStore,
};
// Usage
async function searchDocuments(query: string) {
// Read from replica
return stores.read
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 10))
.select((ctx) => ctx.d)
.execute();
}
async function createDocument(data: DocumentInput) {
// Write to primary
return stores.write.nodes.Document.create(data);
}
```
**When to use:**
- Heavy read workloads (semantic search, graph traversals)
- Write/read ratio is heavily skewed toward reads
- Need to scale read capacity independently
**Considerations:**
- Replication lag means reads may be slightly stale
- Don't use replica for read-after-write scenarios
- Monitor replication lag in production
### Multi-Tenant Architecture
Four approaches for multi-tenant deployments, each with different tradeoffs.
#### Option 1: Shared tables with tenant isolation (simplest)
```typescript
import { defineNode, defineGraph } from "@nicia-ai/typegraph";
// Include tenantId in your node schemas
const Document = defineNode("Document", {
schema: z.object({
tenantId: z.string(),
title: z.string(),
content: z.string(),
}),
});
// Always filter by tenant in queries
function createTenantQuery(store: Store, tenantId: string) {
return {
searchDocuments: (query: string) =>
store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.tenantId.eq(tenantId).and(d.embedding.similarTo(queryEmbedding, 10)))
.select((ctx) => ctx.d)
.execute(),
createDocument: (data: Omit) => store.nodes.Document.create({ ...data, tenantId }),
};
}
// Middleware extracts tenant and creates scoped API
function withTenant(req: Request) {
const tenantId = req.headers.get("x-tenant-id")!;
return createTenantQuery(store, tenantId);
}
```
#### Option 2: Separate `graph_id` per tenant (divergent schemas, one database)
Every TypeGraph row is keyed by `graph_id`, and so is the committed schema. Two
graphs with different `id`s coexist in one database with **independent schemas** —
tenant A can declare kinds tenant B has never heard of, and neither sees the
other's nodes, edges, or kind namespace.
```typescript
function tenantGraph(tenantId: string) {
return defineGraph({
id: `tenant_${tenantId}`, // the isolation boundary
nodes: { Document: { type: Document } },
edges: {},
});
}
// Each tenant commits — and evolves — its own schema, in the same database.
const [store] = await createStoreWithSchema(tenantGraph("acme"), backend);
```
What `graph_id` isolates:
- **Kinds** — the kind namespace is per `graph_id`. Declaring `Invoice` in one
graph does not create it in another.
- **Data** — nodes and edges are filtered by `graph_id` on every read and write.
- **Schema version and evolution** — each graph owns its committed schema
document and version, so tenants migrate independently.
This is the cheap alternative to N physical databases when you want divergent
per-tenant schemas without N connections. See
[Graph identity and the kind namespace](/schemas-stores#graph-identity-and-the-kind-namespace).
**Caveat 1 — index names are database-global.** Materialized SQL index names
are derived from `(kind, fields, shape)` and are **not** namespaced by `graph_id`,
because a SQL index name is a database-global identifier. Two graphs declaring
the same kind name *and* the same index therefore resolve to one physical index:
- **Same shape** — the second graph reuses the first graph's index. For a
graph-scoped index this is safe: the index is keyed by `graph_id` (or
`graph_id, kind`), so each graph still gets its own region of the index.
- **Different shape** (say one `unique`, one not) — materialization fails loudly
with a signature-drift error instead of silently sharing a mismatched index.
Rename the declaration, or drop the existing index and retry.
**Caveat 2 — a unique index must stay graph-scoped.** `scope` decides which
TypeGraph system columns prefix the index key:
| `scope` | Key prefix | Unique constraint applies |
|---------|-----------|---------------------------|
| `"graphAndKind"` (default) | `(graph_id, kind)` | per kind, per graph |
| `"graph"` | `(graph_id)` | per graph |
| `"none"` | *(none)* | **across every graph in the table** |
A `unique` index declared with `scope: "none"` omits `graph_id` from the key, so
the database enforces that value as unique across **all** graphs sharing the
table — one tenant's row will block another tenant's insert. That is a real
cross-tenant effect, and it holds whether or not two graphs share the physical
index. Keep unique indexes on the default `"graphAndKind"` (or `"graph"`) scope
in a multi-graph database; reserve `scope: "none"` for non-unique indexes where
you deliberately want one index spanning every graph.
Subject to those two rules, per-`graph_id` isolation holds: reads and writes
stay filtered by `graph_id`, and the coupling is confined to physical index
reuse.
#### Option 3: Schema per tenant (PostgreSQL)
```typescript
import { sql } from "drizzle-orm";
async function createTenantStore(tenantId: string) {
const schemaName = `tenant_${tenantId}`;
// Create schema if not exists
await pool.query(`CREATE SCHEMA IF NOT EXISTS ${schemaName}`);
// Run migrations in tenant schema
await pool.query(`SET search_path TO ${schemaName}`);
await pool.query(generatePostgresMigrationSQL());
await pool.query(`SET search_path TO public`);
// Create Drizzle instance with schema
const db = drizzle(pool, { schema: { schemaName } });
const backend = createPostgresBackend(db);
return createStore(graph, backend);
}
// Cache tenant stores
const tenantStores = new Map();
async function getTenantStore(tenantId: string): Promise {
if (!tenantStores.has(tenantId)) {
tenantStores.set(tenantId, await createTenantStore(tenantId));
}
return tenantStores.get(tenantId)!;
}
```
#### Option 4: Database per tenant (strongest isolation)
```typescript
interface TenantConfig {
id: string;
databaseUrl: string;
}
async function createTenantStore(config: TenantConfig) {
const pool = new Pool({ connectionString: config.databaseUrl });
await pool.query(generatePostgresMigrationSQL());
const db = drizzle(pool);
const backend = createPostgresBackend(db);
return {
store: createStore(graph, backend),
close: () => pool.end(),
};
}
// Connection manager with LRU eviction
class TenantConnectionManager {
private stores = new Map Promise }>();
private maxConnections = 100;
async getStore(tenantId: string): Promise {
if (!this.stores.has(tenantId)) {
if (this.stores.size >= this.maxConnections) {
await this.evictOldest();
}
const config = await fetchTenantConfig(tenantId);
this.stores.set(tenantId, await createTenantStore(config));
}
return this.stores.get(tenantId)!.store;
}
private async evictOldest() {
const [oldestId, oldest] = this.stores.entries().next().value;
await oldest.close();
this.stores.delete(oldestId);
}
}
```
**Comparison:**
| Approach | Isolation | Complexity | Scaling | Cost |
| ------------------- | --------------- | ---------- | --------------------------- | ------- |
| Shared tables | Low (row-level) | Low | Single DB | Lowest |
| Schema per tenant | Medium | Medium | Single DB, separate schemas | Low |
| Database per tenant | High | High | Independent DBs | Highest |
**When to use each:**
- **Shared tables**: SaaS with many small tenants, cost-sensitive
- **Schema per tenant**: Moderate isolation needs, PostgreSQL only
- **Database per tenant**: Enterprise customers requiring data isolation, compliance requirements
## Next Steps
- [Quick Start](/getting-started) - Basic setup and first graph
- [Semantic Search](/semantic-search) - Vector embeddings and similarity
- [Performance](/performance/overview) - Optimization strategies
# Limitations
> Known constraints and backend-specific limitations
This page documents TypeGraph's known limitations and constraints.
## Backends Without Atomic Transactions
Some runtimes cannot hold a multi-statement database session and therefore
cannot offer atomic transactions:
- **Cloudflare D1** — the D1 binding has no interactive transaction
primitive (`D1Database.batch(...)` is transactional but batch-only).
- **`drizzle-orm/neon-http`** — Neon's HTTP driver issues each statement as
an independent request; there is no session to bind a transaction to.
Cloudflare **Durable Objects** SQLite is *not* in this list: a store backed
by `drizzle(ctx.storage)` is auto-detected as `transactionMode: "do-sqlite"`,
reports `capabilities.execution.interactiveTransactions: true`, and is fully atomic. An
`AdapterStore` created from that backend also exposes the adapter-only
`store.withTransaction` and `tx.sql` surfaces. See
[Backend Setup](/backend-setup#cloudflare-durable-objects-sqlite).
These backends report `capabilities.execution.interactiveTransactions: false`. Read-only
`store.batch(...)` still runs, but each query may use an independent connection
and observe a different database snapshot. (Whether the queries nonetheless
reuse one connection is up to the adapter — the no-transaction path hands each
query the same backend object.) Note this is a difference of degree, not of
kind: on PostgreSQL, `batch()`'s implicit transaction runs at the default
read-committed isolation, so queries there can also observe interleaved commits.
Write behavior depends on how the Store was constructed. A schema-managed Store
fuses its schema fence into a write's own statement when the write fuses, and
fails closed for writes that need the transaction-scoped schema or constraint
fence otherwise — see
[The guard every fused write shares](#the-guard-every-fused-write-shares)
below for which writes fuse and which refuse. A raw `createStore()` /
`createAdapterStore()` without a reconciled snapshot still has no interactive
transaction boundary. `store.transaction(fn)` refuses with a typed capability
error rather than pretending to provide rollback; direct backend writes remain
raw. Eligible operations that use a certified atomic SQL program can still be
available on these roots, but that transport guarantee is separate from the
interactive transaction capability.
These backends cannot honor the `isolationLevel` option on
`store.transaction(...)`; the method refuses before invoking its callback, so
the collection-read snapshot recipe documented elsewhere does not apply here.
```typescript
// On a raw D1 / neon-http Store, this refuses before the callback runs.
await store.transaction(async (tx) => {
await tx.nodes.Person.create({ name: "Alice" });
});
```
**If you require atomicity or schema-version fencing, branch on the capability:**
```typescript
if (store.capabilities.execution.interactiveTransactions) {
await store.transaction(async (tx) => {
/* atomic */
});
} else {
// Use independent operations, or a supported certified atomic operation.
const person = await store.nodes.Person.create({ name: "Alice" });
const company = await store.nodes.Company.create({ name: "Acme" });
await store.edges.worksAt.create(person, company, { role: "Engineer" });
}
```
If you need atomic writes from an edge runtime, use
`drizzle-orm/neon-serverless` (WebSocket-backed Pool) instead of
`drizzle-orm/neon-http`.
### Four kinds of write atomicity
TypeGraph distinguishes an interactive transaction, a static adapter batch, a
certified atomic SQL program, and an authoritative one-statement command. `store.transaction(...)` is the
interactive Store API: it pins a session and groups the callback's operations.
A static batch is adapter-internal (such as D1 `batch()` or a multi-row
insert); it is not a public Store transaction and cannot make arbitrary Store
calls atomic. A certified atomic SQL program is a closed ordered statement
sequence whose transport preserves result slots and parameters and rolls back
primary and sidecar writes when a later statement fails. An authoritative command is a single `commands.execute` write
whose database statement returns the decision it made. It can provide a safe
transactionless create/found path only when the backend has a durable arbiter.
Operational Identity, single-edge claim/cardinality enforcement, and any
undeclared dynamic `matchOn` convergence that may write still require an
interactive transaction and fail closed on a backend that cannot provide one.
Outside the native durable-convergence envelope, an all-live
`ifExists: "return"` endpoint batch is read-only and can return from its
set-oriented root read without a transaction. Inside the native envelope, the
authoritative upsert program runs before the Store knows every identity is
live. It preserves the logical `"found"` result in one exchange, but may take
incumbent-row locks and produce write amplification. Eligible direct edge
batches on bundled roots are a separate exception: their closed native program
carries the claim sidecars inside one atomic exchange. A
declared edge `matchIdentity` persists
a canonical endpoint/property key and has a unique database arbiter; eligible
root `getOrCreateByEndpoints` calls can therefore use the authoritative
one-statement command. The durable identity does not make unrelated Store
operations, claims, or history/revision side effects transactionless.
### The guard every fused write shares
Every static batch and every certified atomic program asserts the active
schema version inside the very statement that writes, never as a preceding
check — the fused create's `WHERE … is_active` predicate, or the program's
leading `schema_fence` CTE. A stale version makes that statement match zero
rows, so the write commits nothing, and the store re-reads and reports
`StaleVersionError` instead of writing against a version that already moved
on. This is what lets a `"batch"`-tier backend (`capabilities.execution.unitOfWork === "batch"`
— Cloudflare D1's `batch()`, Neon HTTP's `transaction(queries)`, which fix
every statement before the first one runs and commit them together with no
session in between) run schema-managed creates, updates and deletes, and
bulk writes at all: the fence travels inside the one exchange it can hold,
instead of needing a session to hold it separately. A singleton node update,
`upsertById`, or delete fuses the same way as a create, through a one-entry
certified atomic program, whenever its kind carries no declared unique
constraint — except a node delete, which fuses even when the kind DOES
carry one, because the atomic delete program releases that claim in the
same statement. A singleton edge update or delete fuses the same way
(`EdgeCollection` has no `upsertById`).
A write that needs more than that one guarded statement — because it must
read a value it wrote earlier in the same write, hold an interactive
callback open across round trips, maintain Operational Identity's closure,
hold history's per-graph lock across a whole write cascade, or hold one
transaction across a schema commit's compare-and-swap — refuses on a
`"batch"`-tier backend with `BATCH_WRITE_UNSUPPORTED`, naming which of those
it needed:
| `reason` | What it needs |
| --- | --- |
| `interactive-callback` | Hold an interactive callback transaction open across several round trips (`store.transaction(fn)`). |
| `constraint-needs-probe` | Read a value it wrote earlier in the same write before deciding what to write next (a declared constraint's probe-then-write). |
| `identity` | Read and write Operational Identity's closure across several round trips inside one held transaction. |
| `history` | Hold the per-graph write lock and clock open across a whole write cascade (`history: true` / `revisionTracking: true`). |
| `schema-commit` | Hold one transaction across its compare-and-swap read and its activating write (`commitSchemaVersion` / `setActiveVersion`). |
A write that simply cannot fuse — an ineligible write kind, a singleton
create/update/`upsertById` on a kind with a declared unique constraint, a
tombstone-resurrection write a supplied id falls through to, or a derived
backend — refuses with `SCHEMA_WRITE_FENCE_UNSUPPORTED` instead and carries
no `batchRefusal` reason: that gate has no proven need to name, only its own
plain limitation.
See [`BATCH_WRITE_UNSUPPORTED`](/errors#batch_write_unsupported) for where
each reason surfaces in an error's `details`.
## libsql Single-Connection Transactions
For local `@libsql/client` connections (`file:` paths and `file::memory:`),
`createLibsqlBackend` frames transactions with raw `BEGIN IMMEDIATE`/`COMMIT`
statements on the client's single stable connection. It deliberately avoids
`client.transaction()`, which hands the client's connection to the transaction
and lazily opens a new one afterwards — for an in-memory database that new
connection is a fresh, empty database
([tursodatabase/libsql-client-ts#229](https://github.com/tursodatabase/libsql-client-ts/issues/229)).
In-memory databases therefore work for all operations, including transactions.
Remote Turso connections (`libsql://`, `http(s)://`) run each transaction on
its own stream via the driver.
The trade-off of a single connection: a store-level operation awaited from
**inside** a `store.transaction` callback (on the root store, rather than the
`tx` context) can never run — the open transaction occupies the backend's
serialized execution slot until it completes — so the backend rejects it with
a `ConfigurationError` instead of deadlocking.
```typescript
// ✅ In-memory works, including transactions
const client = createClient({ url: "file::memory:" });
// ❌ Root-store access inside a transaction callback throws
await store.transaction(async (tx) => {
await store.nodes.Person.find(); // ConfigurationError — use tx.nodes
await tx.nodes.Person.find(); // ✅ transaction-scoped access
});
```
## Recursive Traversal Depth
Variable-length traversals use two depth caps and an explicit cycle policy:
1. Unbounded traversals (no `maxHops` option) are capped at 10 hops.
2. Explicit `maxHops` values are validated up to 1000 hops (`maxHops: >1000` throws).
3. Cycle prevention is on by default. To skip cycle checks for speed, opt into
`cyclePolicy: "allow"` (which may revisit nodes across hops).
This prevents runaway queries while still supporting deep, intentionally bounded traversals.
```typescript
// Implicitly limited to 10 hops
store
.query()
.from("Person", "p")
.traverse("reportsTo", "e")
.recursive()
.to("Person", "manager");
// Explicit limits up to 1000 are honored
store
.query()
.from("Person", "p")
.traverse("reportsTo", "e")
.recursive({ maxHops: 200 }) // honored
.to("Person", "manager");
// Explicit limits above 1000 throw
store
.query()
.from("Person", "p")
.traverse("reportsTo", "e")
.recursive({ maxHops: 2000 }) // throws
.to("Person", "manager");
```
The unbounded-traversal limit is defined as `MAX_RECURSIVE_DEPTH`:
```typescript
import { MAX_RECURSIVE_DEPTH } from "@nicia-ai/typegraph";
// MAX_RECURSIVE_DEPTH = 10
```
## Connection Management
Managed Store factories own their local SQLite or PGlite connection, and their
`store.close()` method releases it. The local backend factories
`createLocalSqliteBackend` and `createLocalPgliteBackend` likewise expose an
owned backend whose `close()` releases its resources.
Bring-your-own adapter factories leave connection ownership with you. For
`createSqliteBackend`, `createPostgresBackend`, and `createLibsqlBackend`, you
are responsible for:
1. **Creating and configuring** the database connection
2. **Implementing connection pooling** for production use
3. **Closing connections** when done
```typescript
import Database from "better-sqlite3";
import { drizzle } from "drizzle-orm/better-sqlite3";
import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// You manage the connection
const sqlite = new Database("app.db");
sqlite.exec(generateSqliteMigrationSQL());
const db = drizzle(sqlite);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
// You close the connection
sqlite.close();
```
For production deployments, use connection pooling:
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Maximum connections
});
const db = drizzle(pool);
const backend = createPostgresBackend(db);
```
In the bring-your-own example above, `store.close()` leaves the supplied driver
open. Close that driver or pool through its own API.
## Predicate Serialization
Where predicates in unique constraints cannot be serialized. If you use schema
serialization for versioning or migration, predicates are stored as
`"[predicate]"` and cannot be reconstructed.
```typescript
// This predicate works at runtime...
unique({
name: "email_unique_when_active",
fields: ["email"],
where: (props) => props.status.isNotNull(),
});
// ...but serializes as:
// { "where": "[predicate]" }
```
**Workaround:** For full schema serialization support, avoid predicates in unique constraints.
Use application-level validation instead.
## Vector Search Backend Requirements
Vector and hybrid search work across all primary backends via a pluggable
`VectorStrategy`. Each backend advertises its capabilities through
`backend.capabilities.vector` (`{ supported, metrics, indexTypes, maxDimensions }`):
| Backend | Requirement | Metrics |
|---------|-------------|---------|
| PostgreSQL | pgvector extension (HNSW / IVFFlat) | cosine, l2, inner_product |
| SQLite | sqlite-vec extension (`vec0` KNN) | cosine, l2 |
| libSQL / Turso | built-in native engine (DiskANN); nothing to load | cosine, l2 |
| D1 | Not supported | — |
Note that `inner_product` is PostgreSQL-only — sqlite-vec and libSQL support
cosine and l2 only.
Using vector predicates on unsupported backends throws `UnsupportedPredicateError`:
```typescript
try {
await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.similarTo(queryVector, 10))
.select((ctx) => ctx.d)
.execute();
} catch (error) {
if (error instanceof UnsupportedPredicateError) {
// Vector search not available on this backend
}
}
```
## Query Builder Type Inference
Complex query chains may occasionally require explicit type annotations when TypeScript cannot
infer the full type. This is rare but can occur with deeply nested selects or unions.
```typescript
// If type inference fails, add explicit type
const results = await store
.query()
.from("Person", "p")
.select((ctx) => ({
name: ctx.p.name as string, // Explicit annotation
}))
.execute();
```
## Bulk Operation Limits
Bulk operations (`bulkCreate`, `bulkInsert`, `bulkUpsertById`, `bulkDelete`) have practical limits based on your database:
| Database | Recommended Batch Size |
|----------|----------------------|
| SQLite | 500-1000 items |
| PostgreSQL | 1000-5000 items |
For larger datasets, batch your operations:
```typescript
const BATCH_SIZE = 1000;
for (let i = 0; i < items.length; i += BATCH_SIZE) {
const batch = items.slice(i, i + BATCH_SIZE);
await store.nodes.Person.bulkCreate(batch);
}
```
### Native node bulk eligibility
Bundled PostgreSQL roots using a recognized session-capable driver, Neon HTTP,
Cloudflare D1, and libSQL roots can use one schema-fenced native atomic program
for schema-managed `nodes.bulkInsert()` and `nodes.bulkCreate()` calls when the
node has no Operational Identity, history, or revision work. The program can
compose fulltext/vector projections with the complete supported uniqueness and
disjointness claim set for every member. Session-capable
PostgreSQL executes that program on one pinned
transaction; Neon HTTP, D1, and libSQL submit one transport batch. Advertised
same-kind or hierarchy-wide uniqueness claims, disjointness claims, and mixed
families are acquired in canonical order; compatibility reads preserve rows
written under legacy claim axes. Claim-free members may participate alongside
claimed members. IDs may be generated, caller-supplied, or mixed, and
`bulkCreate()` returns rows in input order. This is an internal optimization,
not a general Store batch API. Identity-enabled nodes, history/revision
tracking, a member beyond the executor's declared claim-input budget, missing
schema-fence support, and other unsupported shapes fail closed to the existing
transaction or fallback behavior.
The transport inventory for the supported libSQL root records one client
`batch` submission and zero client `execute` calls for both generated-ID
claim-free batches, multiple-claim batches, cross-scope claims, and
claim-plus-projection batches. This is a measured submission count, not a
wall-clock RTT benchmark;
fallback paths are intentionally not assigned a latency claim.
On D1, claim work is chunked inside the same submission rather than imposing a
batch-wide ceiling. Each member has 87 claim-input binds after its row and
fence: a canonical claim costs six, each legacy hierarchy-wide uniqueness
probe costs nine, and each legacy disjointness probe costs six. Custom
executors should call the exported `atomicNodeClaimInputCost()` owner rather
than reproduce this formula. A member beyond that bound retains the portable
behavior.
Direct `edges.bulkInsert()` and `edges.bulkCreate()` calls on those same roots
use one schema-fenced native program when history and revision capture are
disabled. Declared durable match identities and `one`, `unique`, or
`oneActive` cardinality are maintained inside that exchange; any endpoint,
identity, or cardinality refusal rolls the whole call back. Transaction-scoped
stores, derived backends, custom backends without the corresponding exact-root
semantic registration, and dynamic get-or-create convergence retain the
interactive path.
Direct edge `bulkDelete()` calls use the same exact-root exchange and refuse a
foreign-kind ID atomically. Restricted node `bulkDelete()` also releases every
unique or disjoint claim owned by rows it tombstones in the same program, while
enforcing live connected edges in SQL. Identity, projection, history,
revision, cascade, and disconnect shapes retain their transaction path. `bulkUpsertById()`
remains a resolved mutation set because it must read and schema-validate a
database preimage before its writes are known. Bundled serverless roots can
submit an eligible distinct-ID, live-row resolved set as one native exchange
after that read. Bundled session-capable PostgreSQL can bind the same program
to the exact collection-opened, caller-supplied, or adopted transaction; this
is a bounded statement sequence on the pinned session, not one network
exchange. Update-only sets use a guarded update; sets containing both
fresh creates and updates include a terminal database assertion that rolls the
whole exchange back when any guarded postimage is absent. Repeated IDs,
resurrections, temporal changes, claims, edge sidecars (including durable edge
match identity), history/revision capture, ordinary derived backends, and
unregistered sessions use the interactive path. On D1's 100-parameter budget,
each native statement carries at most 17 node mutations or 6 edge mutations.
Larger eligible sets are chunked inside the same atomic transport submission;
each chunk has its own terminal postimage assertion, so one refusal rolls every
sibling chunk back rather than weakening the set contract. A D1 submission is
bounded to 512 node members or 187 edge members; larger sets fail closed to the
portable path instead of building an unbounded request. Other backends derive
their statement width from their declared bind budget and retain an absolute
512-member submission ceiling. The
operation returns an explicit `unsupported` verdict before issuing program SQL;
the Store never infers fallback safety from a missing result. Once a session
program starts, a savepoint preserves the surrounding transaction for typed
refusal diagnosis.
Node `bulkReplaceById()` avoids that structural preimage read by accepting only
complete replacement documents and distinct IDs. On an eligible bundled root,
the complete call—including claim ownership changes and fulltext/vector
sidecars—uses one atomic transport submission. Live rows retain their stored
validity windows; tombstones receive a freshly stamped window. Operational
Identity and history/revision capture use the portable path. Custom backends
must register and semantically certify the independent `replaceNodes` family;
transport registration or another node family is not evidence for replacement.
Eligible singleton `update()` and `delete()` calls reuse those same registered
families. Plain or projected node updates, unconstrained non-durable-identity edge updates,
all direct edge deletes, and plain restricted node deletes remain two-exchange
operations—one authoritative read/gate and one atomic mutation—because
TypeGraph must validate merged update properties and must preserve the rule
that a missing delete fires no operation hooks. This removes explicit
transaction transport from the eligible shape; it does not turn claims, edge
sidecars, temporal, captured, derived-backend, or caller-transaction writes into
autocommit operations.
That singleton update path uses optimistic convergence: the mutation asserts
the row preimage it read and retries a moved preimage up to four times. Under
sustained same-row contention it can throw `DatabaseOperationError` where an
interactive transaction would have waited to serialize the writers. This
applies to eligible `update()` calls and the live-row leg of `upsertById()` on
registered exact-root atomic transports. Caller transactions and other
ineligible shapes continue to use the serialized transaction path. Applications
using an atomic root should retry the operation when sustained contention can
move the row throughout all four attempts.
### One `bulkUpsertById` batch cannot hand a constrained value between rows
`bulkUpsertById` applies items in order for the purpose of deciding each row's
final props, but it groups the writes: every create in the batch runs before
every update. A batch where one item **releases** a constrained value and a later
item **claims** it therefore fails, where the same operations applied one at a
time succeed.
- Nodes: releasing and re-claiming a `unique` constraint value in one batch
throws `UniquenessError` — the claiming create is checked while the releasing
row still reserves the value.
- Edges: ending the lone `oneActive` edge from a source while creating its
replacement throws `CardinalityError`, for the same reason.
Bulk semantics are set-like, not scripted — a batch states the rows you want,
not an order to reach them in — so this is a stated limitation rather than a
pending fix. It always surfaces as a typed error, never as a dropped write.
Split the handoff across two batches (release, then claim), or apply the
conflicting items one at a time — as sequential `upsertById` calls for nodes,
and as `update` then `create` for edges, which have no single-item upsert. See
[Data Sync](/data-sync#one-batch-cannot-hand-a-unique-value-from-one-row-to-another)
for the worked example.
## Graph Analytics Limits
TypeGraph ships focused algorithms on `store.algorithms.*` — shortest path
(weighted and unweighted), reachability, k-hop neighborhoods, degree, exact
weakly connected components, deterministic label propagation, and
global/personalized PageRank. See
[Graph Algorithms](/graph-algorithms) for the full API.
The following heavier analytics are **not** provided:
- Modularity-optimizing community detection such as Leiden or Louvain
- Centrality measures beyond degree (betweenness, closeness, eigenvector)
- Strongly connected components
- Topological sort
- Graph partitioning
For these use cases, export your data via `.query().traverse()` or
`store.subgraph()` and use a specialized library such as
[graphology](https://graphology.github.io/) in memory, or move to a
dedicated graph database.
## Single Database Deployment
TypeGraph is designed for single-database deployments. It does not support:
- Distributed storage across multiple databases
- Sharding
- Cross-database queries
- Replication coordination
For distributed graph workloads, consider a dedicated graph database.
## Temporal Query Limitations
Temporal queries (`asOf`, `includeEnded`) work correctly but have some constraints:
- Point-in-time queries cannot be combined with streaming (`.stream()`)
- `validFrom` defaults to the record's own creation timestamp when omitted, so `asOf` queries
work out of the box; an end boundary still requires an explicit `validTo` — an open `validTo`
means "still valid". A record written with a `validTo` at or before its own creation instant
is "born already ended" and stores no lower bound instead, so it reads back at every `asOf`
before that end
- Rows an **older library version** stored with a backwards window
(`valid_from > valid_to`) are readable at no coordinate, and upgrading does not rewrite
them. Making them observable is an explicit operator action: run
`repairInvertedValidityWindows({ relations: "live-and-recorded", mode: "apply" })` while
writers are stopped, then re-baseline any outstanding merge branches. See
[Repairing inverted validity windows](/schema-management#repairing-inverted-validity-windows)
- Clock skew between application servers can affect temporal accuracy
### Recorded / system time (`history: true`)
Recorded-time capture (`createStore(graph, backend, { history: true })`) and
`store.asOfRecorded(T)` add a second temporal axis with these constraints. Use
`createAdapterStore(..., { history: true })` instead when the application must
adopt a caller-owned transaction:
- **Opt-in, no backfill.** Capture only sees changes committed after it is
enabled; an entity that already exists is first recorded the next time it is
written. Enable it on a fresh graph for complete history.
- **TypeGraph-write capture.** Built-in capture records TypeGraph collection
writes only. Out-of-band database writes and row-returning raw SQL paths are
not captured into the recorded relations.
- **Reconstructing reads only.** A recorded view exposes point reads
(`getById` / `getByIds`), bounded deterministic `scan()` pages, `query()`,
`subgraph()`, and the graph algorithms. Broad filtered collection reads
(`find` / `count` / `findFrom`), `search`, and fulltext / vector predicates are
refused — those indexes reflect current state and cannot answer a
recorded-time query.
- **Transactional backend required.** Capture needs a backend with atomic
transactions and statement execution — the built-in SQLite / PostgreSQL
backends qualify. A custom backend must implement `executeStatement` (optional
on the `GraphBackend` interface, but required once `history: true` is set) or
enabling capture throws a `ConfigurationError` at write time. On an
`AdapterHistoryStore`, raw `tx.sql` is disabled under `history: true`; adopt
external transactions with `store.withRecordedTransaction(...)` instead of
`store.withTransaction(...)` (which is a compile error on a history store).
- **Reconstruction cost.** Recorded reads rebuild from the history relations and
are slower than live reads, most noticeably for full-graph subgraph /
algorithm reconstructions on PostgreSQL.
- **PostgreSQL capture requires `READ COMMITTED`.** Every captured commit
advances a single recorded-clock row for the graph. TypeGraph refuses
PostgreSQL `REPEATABLE READ` / `SERIALIZABLE` history-capture transactions
because snapshot isolation cannot safely allocate that per-graph recorded
clock inside the captured transaction. Omit the transaction isolation option,
or set it to `read_committed`.
- **Recorded anchors are per graph.** Each captured transaction advances a
fixed-width logical revision and pairs it with a non-decreasing physical
wall-time high-water mark. TypeGraph does not provide a cross-graph recorded
anchor. See
[Logical revision and physical time](/queries/temporal#logical-revision-and-physical-time).
- **The preview schema needs an offline migration.** Timestamp-only anchors and
PostgreSQL recorded relations using `timestamptz` predate numeric recorded
revisions and the `r1::` API encoding. Run
`migrateLegacyRecordedTime()` while writers are stopped, then use
`migrateRecordedAnchor()` for checkpoints held outside TypeGraph. See
[Migrating preview recorded time](/schema-management#migrating-preview-recorded-time).
## Schema Migration Constraints
Automatic migrations (`createStoreWithSchema`) only handle additive changes:
| Change Type | Auto-Migrated |
|-------------|---------------|
| Add new node type | Yes |
| Add new edge type | Yes |
| Add optional property | Yes |
| Add required property | No |
| Remove property | No |
| Rename type | No |
| Change property type | No |
Breaking changes throw `MigrationError` and require manual migration.
# Materializing External Event Logs
> How to project at-least-once event streams into TypeGraph without making TypeGraph an event-log product
External logs are the transport. TypeGraph is the typed, entity-resolved
materialization and merge layer.
Use this pattern when agents or integration runtimes already run on an event log
or stream: Electric Durable Streams, database changefeeds, message queues, or a
custom append-only feed. The log owns delivery, ordering, replay, and offsets.
TypeGraph owns the current graph, valid-time facts, recorded-time history, and
mergeable working copies. The sibling
[`agent-stream-graph`](https://github.com/nicia-ai/agent-stream-graph) package is
the reference implementation of this posture.
## The Shape of the Problem
External log consumers usually have three properties:
- **At-least-once delivery.** A change can be delivered more than once, especially
after a crash or reconnect.
- **Resume from a cursor.** The consumer persists the last source offset it has
safely processed.
- **Replay.** Reprocessing old events is normal: for recovery, backfills, or
rebuilding a derived graph.
That means a projector must be idempotent. Re-delivering the same source change
should converge on the same graph state, not create duplicates.
## Idempotent Projectors
Use stable source ids as TypeGraph ids whenever the source has them. For nodes,
that usually means `upsertById`. For edges, prefer `getOrCreateByEndpoints`.
Avoid `create` in a log projector unless the source event itself carries a
unique id you pass as the TypeGraph id.
```typescript
async function projectChange(
tx: TransactionContext,
change: Change,
) {
const issue = await tx.nodes.Issue.upsertById(
change.issueId,
{
title: change.title,
state: change.state,
},
{
validFrom: change.issueValidFrom,
onImmutableLowerBound: "preserve",
},
);
const actor = await tx.nodes.Actor.upsertById(change.actorId, {
name: change.actorName,
});
await tx.edges.changedBy.getOrCreateByEndpoints(
issue,
actor,
{ action: change.action },
{
ifExists: "update",
validFrom: change.relationshipValidFrom,
validTo: change.relationshipValidTo,
onImmutableLowerBound: "preserve",
},
);
}
```
The important rule is that the second delivery of the same change takes the same
code path and reaches the same row identities.
The `"preserve"` policy makes `validFrom` create/resurrection-only input for
both node and edge writes: a later revision updates props and `validTo` without
trying to rewrite the live row's start. Without it, the default `"refuse"`
policy raises `IMMUTABLE_VALIDITY_LOWER_BOUND` when a revision states a
different start. The edge also explicitly selects `ifExists: "update"`; the
default is `"return"`, which is right for create-once relationships but writes
neither revised props nor a closing `validTo` when the edge already exists.
### `matchOn` widens the identity key — don't reach for it by default
`getOrCreateByEndpoints` matches on the endpoints `(from, to)` alone unless you
pass `matchOn`. Endpoints-only is the **more** idempotent choice and is right for
most projectors: a re-delivered edge between the same two nodes converges on the
one existing edge regardless of how its properties drifted between deliveries.
`matchOn` adds the named property fields to the match key, so it *widens*
identity — two edges between the same endpoints are now distinct if they differ
on a matched field. Use it only when the relationship model genuinely allows
several parallel edges between one pair (say, one `changedBy` edge per distinct
`action`), and know the footgun: if a re-delivered change carries a **changed**
value in a matched field, it no longer matches the earlier edge and you get a
**second** edge instead of convergence. Reach for `matchOn` when the domain
needs the extra edges, not as a reflex.
Validity timestamps do not become part of this identity key. If the same
endpoints can have multiple application-time periods, include a stable period
or source-event identifier in the edge schema and in `matchOn`. This keeps a
re-delivery of one period convergent without collapsing a later period into the
same edge.
## Cursor Bookkeeping
A cursor is application state: the last source offset you have safely processed.
It should advance only at a source offset boundary, after every change in that
batch has been projected. Where the cursor lives — a row in your own relational
table, or a node in the graph — decides which guarantees you can get.
### Exactly-once with an adopted transaction
To commit the projected batch **and** the cursor as one unit, let the caller own
the transaction and adopt it with
[`store.withRecordedTransaction(externalTx, fn)`](/schemas-stores/#transaction-receipts).
The graph writes and your own cursor write land on the same connection inside the
same commit: either both persist or neither does, so the cursor can never advance
past a batch the graph did not durably record.
Two constraints make this the *only* sanctioned transactional recipe on a store
created with `createAdapterStore` or `createAdapterStoreWithSchema` and
`{ history: true }` — which the Transaction Receipts and Bitemporal sections
below both require:
- **Write your own tables through the external handle you passed in**, never
through `tx.sql`. Under history capture the typed transaction context omits
`sql` (raw SQL would bypass recorded-time capture); suppressed access reaches
a runtime guard and raises a
[`ConfigurationError`](/errors/#recorded-capture-guard-codes). The external
handle *is* the pinned connection, so writing your cursor row through it keeps
both layers in the one transaction.
- **`store.withTransaction()` — the non-recorded sibling — is a compile error on
a history store** (its `externalTx` argument is rejected against a message
type), and its runtime guard throws
`RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION`. It has no flush point before
the caller commits, so recorded-time capture could not seal. Use
`withRecordedTransaction` instead.
**Async drivers (Postgres / libsql)** open the boundary with `db.transaction`:
```typescript
const receipt = await db.transaction(async (dbTx) => {
const outcome = await store.withRecordedTransaction(dbTx, async (tx) => {
for (const change of batch.changes) {
await projectChange(tx, change);
}
});
// The cursor row goes through the external handle, in the same transaction.
await dbTx
.insert(streamCursors)
.values({ sourceId: batch.sourceId, offset: batch.endOffset })
.onConflictDoUpdate({
target: streamCursors.sourceId,
set: { offset: batch.endOffset },
});
return outcome.receipt;
}); // one COMMIT / ROLLBACK across both layers
```
**Synchronous `better-sqlite3`** cannot adopt an `async` transaction callback
(its driver rejects a promise-returning `db.transaction`), so the caller frames
the boundary by hand with `BEGIN IMMEDIATE` / `COMMIT` / `ROLLBACK` on the single
connection:
```typescript
await db.run(sql`BEGIN IMMEDIATE`);
try {
const { receipt } = await store.withRecordedTransaction(db, async (tx) => {
for (const change of batch.changes) {
await projectChange(tx, change);
}
await db.run(sql`
INSERT INTO stream_cursor (source_id, offset)
VALUES (${batch.sourceId}, ${batch.endOffset})
ON CONFLICT (source_id) DO UPDATE SET offset = excluded.offset
`);
});
await db.run(sql`COMMIT`);
// persist receipt.recorded as the offset's replay anchor — see below
} catch (error) {
await db.run(sql`ROLLBACK`); // graph writes and cursor roll back together
throw error;
}
```
The graph writes and your own statements share the caller's one pinned
connection. TypeGraph serializes the statements its collections issue; sequence
your own raw statements yourself (don't `Promise.all` them with graph writes) so
two queries never race on that connection.
For an adapter-backed materializer that runs against several backends, branch
on capability rather than message-matching: use
[`tx.sqlAvailability`](/recipes/#cross-store-transactions-drizzle--typegraph) to
decide whether raw SQL is usable inside `store.transaction`, and
[`isRecordedCaptureGuardError(error, code?)`](/errors/#recorded-capture-guard-codes)
to recognize a history-store guard when you catch one.
### At-least-once with a separate cursor store
When the runtime already owns checkpointing, or the backend cannot provide atomic
transactions (`backend.capabilities.execution.interactiveTransactions === false` — Cloudflare D1,
`drizzle-orm/neon-http`), keep the cursor outside the graph transaction. The
pattern is at-least-once plus idempotence: a crash after the graph writes but
before the cursor write replays the batch, which is safe precisely because the
projector converges.
This fallback requires a raw Store. A schema-managed Store refuses writes on a
non-transactional backend because it cannot hold the schema-version fence. Use a
transactional driver, or deliberately construct a raw Store and own schema/write
coordination yourself.
```typescript
for (const change of batch.changes) {
// Each successful projection may commit before a later projection or cursor write fails.
await projectChange(store, change);
}
await cursorStore.save({
sourceId: batch.sourceId,
offset: batch.endOffset,
});
```
This at-least-once path plus an idempotent projector is the workload TypeGraph is
built for. It is also the one that churns recorded history the hardest: every
re-delivery of a byte-identical change rewrites its row, allocating a fresh
recorded instant and a new history row per delivery. Enable
[`coalesceUnchangedUpserts: true`](/schemas-stores/#createstoregraph-backend-options)
on the store to suppress that. A node `upsertById` and an edge endpoint
get-or-create update perform no write, history row, or revision advance when
its validated props and requested window already equal the live row. Their bulk
forms have the same behavior. See
[Transaction Receipts](#transaction-receipts) for how a coalesced upsert reads on
a receipt.
Every captured transaction receives one versioned recorded instant: a strict
per-graph logical revision paired with a non-decreasing physical wall-time
high-water mark. High commit rates consume revisions without pushing the
timestamp beyond observed wall time. A backward clock correction holds the
physical component at its prior value until the clock catches up, preserving
cumulative diagonal checkpoint replay. Group changes by their durable
replay/checkpoint boundary so one addressable source position consumes one
recorded instant where practical. Cap transaction size independently: a source
may expose one coarse checkpoint for a very large initial sync, but that does
not make an unbounded transaction safe.
Recorded clocks are independent per graph, and there is no cross-graph
`recordedNow()` snapshot. See
[Logical revision and physical time](/queries/temporal#logical-revision-and-physical-time)
for the anchor encoding and replay semantics.
**Coalescing eliminates *re-delivery* churn, not replay cost.** The win is
scoped to re-delivery of the current value — the realistic at-least-once case,
where a change that was already applied arrives again (a crash-window replay, a
duplicate) and is value-identical to the live row. A full **replay-from-zero**
over the current state is different: if the stream contains in-place updates,
replaying `insert a=1 … update a=2` re-applies `a=1` over the live `a=2` — a
genuine backward change that writes — and then `a=2` restores it. Both writes
are correct (the replay faithfully re-walks each historical state), but
"coalescing makes replay free" holds only for streams whose rows never
supersede each other. It also leaves a spurious `a=2 → a=1 → a=2` band in the
live store's recorded history, stamped at replay time. To rebuild without
either cost, replay into a **fresh store** and publish it, rather than
re-applying the log over the current state.
**Historical ends need a historical start on the creating event.** When a fresh
store creates a row without a stated `validFrom`, TypeGraph uses the ingest
instant. A later replayed event whose historical `validTo` precedes that ingest
instant is therefore an `INVERTED_VALIDITY_WINDOW`, even if the source timeline
itself was ordered. An event-time decoder must emit `validFrom` on the event
that first creates each node or edge; `onImmutableLowerBound: "preserve"` then
lets later revisions carry their source bound without trying to move the stored
start.
### In-graph cursors and the receipt
A cursor can also live inside the graph as an ordinary node — convenient, and it
travels with the graph. But if you also use the transaction receipt (next
section) to detect a projector that dropped a change, an in-graph cursor
**corrupts that signal**: the receipt counts writes per transaction with no
attribution, so the cursor's own upsert is indistinguishable from the projector's
writes. A projector that drops a change in a transaction that also checkpoints an
in-graph cursor produces `writes.total === 1` from the cursor alone —
`writes.total > 0` no longer means "the projector wrote," the drop goes
undetected, and the cursor advances past the lost event.
Two ways out:
- **Scope the projector with `tx.measure`.** On a receipt-enabled context
(`transactionWithReceipt` or `withRecordedTransaction`),
[`tx.measure((scopedTx) => …)`](/schemas-stores/#scoped-receipts-txmeasure)
hands your callback a **scoped context** and returns a sub-receipt that counts
exactly the writes made **through that scoped context** (`scopedTx.nodes` /
`scopedTx.edges`). Run the projector through `scopedTx`; write the cursor
through the outer `tx` (or your own table). Attribution is by which context you
write through, not by timing, so the cursor's write counts only in the outer
receipt and the scope reflects the projector alone. This is what makes an
in-graph cursor and drop-detection composable; see the full loop below.
- **Keep the cursor in your own relational table.** The
[exactly-once recipe](#exactly-once-with-an-adopted-transaction) makes that
atomic anyway, and it keeps the belief graph pristine: an in-graph cursor node
still lands in every `asOfRecorded` reconstruction of the graph, so a consumer
that wants recorded-time reads to show only projected facts should keep the
cursor out of the graph entirely.
## Transaction Receipts
When you need to know what a projector did, use `store.transactionWithReceipt`
(TypeGraph owns the boundary) or `store.withRecordedTransaction` (you adopt an
open transaction); both return a `TransactionOutcome` with a `receipt`.
The receipt carries **two signals that deliberately disagree**, and a
materializer needs both. `receipt.writes.total` counts completed write intents at
the collection surface; `receipt.recorded` (on a `{ history: true }` store) is
the recorded commit instant this transaction allocated, or `undefined` when
nothing was captured or explicitly requested. The common, load-bearing case is
where they diverge:
| case | `writes.total` | `recorded` |
| --------------------------------------------- | -------------- | --------------- |
| projector wrote | `> 0` | defined |
| no-op delete of an absent key (a real intent) | `1` | **`undefined`** |
| coalesced upsert (value-identical, opt-in) | `1` | **`undefined`** |
| projector dropped the change | `0` | `undefined` |
| explicit recorded revision request | `0` | defined |
A no-op delete completes a write *intent* but captures nothing; a
[coalesced upsert](/schemas-stores/#createstoregraph-backend-options) is the same
shape by design. In both, `writes.total` counts (the method resolved) but
`recorded` is `undefined`. **An offset whose transaction reports
`recorded === undefined` must carry the prior anchor forward** — otherwise
replay-by-offset breaks at exactly the offsets where nothing changed.
Call `requestRecordedRevision()` when that offset must instead receive its own
anchor despite making no entity change.
Two counting rules bite materializers specifically, both worth internalizing
before you read `writes.total` as "the projector did work":
- **Bulk methods count by input length**, so `bulkCreate([])` contributes `0`. A
projector that filters a batch down to nothing and issues an empty bulk call
must not read as a writer.
- **A method that rejects counts `0`** — even on SQLite, where a failed statement
does **not** abort the surrounding transaction. A projector that swallows a
write error and commits can persist rows the receipt never counted, so do not
read the receipt as rows-affected in that scenario.
### The full materializer loop
Putting the pieces together: an adopted transaction for exactly-once cursors, a
`tx.measure`-scoped projector so a single dropped change is caught within a
multi-change batch — the outer receipt only tells you the whole batch wrote
nothing, whereas a `measure` scope attributes writes per change by having the
projector write through the scoped context it receives (any cursor written
through the outer `tx` stays out of that count) — `receipt.recorded` as the
per-offset replay anchor, and `writes.total === 0` on a non-delete change as the
drop signal.
`withRecordedTransaction` flushes recorded-time capture and resolves **before**
the caller's commit, so `outcome.receipt.recorded` is already known inside the
`db.transaction` callback. Write the cursor advance **and** its replay anchor
through `dbTx` there, in the same commit as the graph writes. Persisting the
anchor after the commit — as a separate step — would reopen the exactly-once gap
the adopted transaction exists to close: a crash between the commit and the
anchor write leaves the cursor advanced with no anchor, and that offset can never
be replayed.
If an accepted source position must be addressable even when the projector makes
no entity changes, request an explicit revision in the callback. Await the
`withRecordedTransaction` outcome, then insert the application-owned cursor and
returned anchor through the still-open native transaction before its commit:
```typescript
await db.transaction(async (dbTx) => {
const { receipt } = await store.withRecordedTransaction(dbTx, async (tx) => {
tx.requestRecordedRevision();
await projectBatch(tx, batch);
});
await dbTx.insert(cursors).values({
source: batch.source,
offset: batch.offset,
recorded: receipt.recorded,
});
});
```
```typescript
let lastAnchor: RecordedInstant | undefined = await loadLastAnchor(); // on resume
lastAnchor = await db.transaction(async (dbTx) => {
const outcome = await store.withRecordedTransaction(dbTx, async (tx) => {
for (const change of batch.changes) {
// The projector writes through the scoped context, so `projected`
// counts its writes alone — nothing else in the transaction.
const projected = await tx.measure((scopedTx) =>
projectChange(scopedTx, change),
);
// A non-delete change that wrote nothing was silently dropped.
if (
projected.receipt.writes.total === 0 &&
change.operation !== "delete"
) {
throw new DroppedChangeError(change); // rolls the whole batch back
}
}
});
// `recorded` is undefined when the batch captured nothing (all drops, no-op
// deletes, or coalesced upserts) — carry the prior anchor forward so replay
// by offset still resolves. Anchor comes from the receipt, never from a
// post-commit store.recordedNow() (see below).
const anchor = outcome.receipt.recorded ?? lastAnchor;
// Cursor and anchor commit atomically with the graph writes: no window where
// the cursor has advanced past an offset whose anchor was never persisted.
await dbTx
.insert(offsetAnchors)
.values({ sourceId: batch.sourceId, offset: batch.endOffset, recorded: anchor });
await dbTx
.insert(streamCursors)
.values({ sourceId: batch.sourceId, offset: batch.endOffset })
.onConflictDoUpdate({
target: streamCursors.sourceId,
set: { offset: batch.endOffset },
});
return anchor; // updates lastAnchor only once the transaction commits
});
```
**Take the replay anchor from `receipt.recorded`, never from a post-commit
`store.recordedNow()`.** `recordedNow()` is the graph-global recorded
high-water mark, advanced by **any** writer to the graph. Between your commit and
your read of it, a concurrent writer can advance it, and `asOfRecorded(that)`
then reconstructs a belief your stream never produced. The receipt hands you the
instant *this* transaction allocated; that is the only anchor that reconstructs
exactly what this offset materialized.
## Bitemporal Mapping
External streams usually carry domain time and delivery time. Keep those
separate:
- **Event time belongs in valid time.** If a source change says a fact became
true on January 1, pass that timestamp as `validFrom`; if it ended on January
31, pass `validTo`.
- **Ingest time is recorded time.** TypeGraph records when the graph committed
the write. Recorded time is allocated by the backend and cannot be backdated.
- **Backfills collapse recorded instants to now.** Replaying historical events
today writes historical valid-time facts with today's recorded-time anchors.
That is correct SQL:2011 bitemporal behavior, not a bug.
To replay by source offset, load the anchor you saved for that offset and read a
recorded-time view. The `receipt.recorded` you persisted is a branded
`RecordedInstant`, but round-tripping through your cursor table stores it as a
plain string — re-brand it with `asRecordedInstant` on the way back before
passing it to `asOfRecorded`:
```typescript
import { asRecordedInstant } from "@nicia-ai/typegraph";
const stored = await offsetAnchors.anchorFor(offset); // plain string from storage
const anchor = asRecordedInstant(stored); // validates + re-brands
const graphAtOffset = store.asOfRecorded(anchor);
const issue = await graphAtOffset.nodes.Issue.getById(issueId);
```
If the cursor table contains timestamp-only anchors from the recorded-time
preview, migrate the TypeGraph relations first and remap those cursor values
with `migrateRecordedAnchor({ backend, graphId, anchor: stored })`. See
[Migrating preview recorded time](/schema-management#migrating-preview-recorded-time).
That answers "what did the materialized graph know after offset X?" even if
later corrections changed or deleted rows. See
[Recorded time](/queries/temporal/#recorded-time-bitemporal) for the full view
surface.
### Refresh planner statistics after a large replay
A **custom** replay or backfill loop — one built from the projector recipes above
— runs its writes **inside a caller-provided transaction**, which never
auto-refreshes the query planner's table statistics: `ANALYZE` from another
connection cannot see rows that are still uncommitted, so the store deliberately
skips the automatic refresh it does after large autocommit bulk writes. Left
alone, the planner keeps pre-load row estimates and can pick an
order-of-magnitude-slower plan. After a large custom replay, refresh once:
```typescript
await replayEverything();
await store.refreshStatistics(); // once, after the bulk replay commits
```
The interchange path handles this for you: `importGraph` and `importGraphStream`
call `refreshStatistics()` once after the import commits (see
[Bulk Copy Between Stores](#bulk-copy-between-stores)), so a bulk copy needs no
manual refresh.
## Bulk Copy Between Stores
To copy a materialized graph into another store — most often a graph-merge
working copy — stream interchange directly from source to target with
`exportGraphStream` / `importGraphStream`. This is the same path
[graph-merge](/interchange/) uses internally, so a copy produces byte-identical
merge results, conflicts, and provenance to a native branch:
```typescript
import {
exportGraphStream,
importGraphStream,
} from "@nicia-ai/typegraph/interchange";
const result = await importGraphStream(
branch.store,
exportGraphStream(beliefStore, {
nodeKinds: ["Belief", "Claim"],
edgeKinds: ["supports"],
includeTemporal: true,
}),
{ onConflict: "update" },
);
if (!result.success) {
throw new Error(`copy failed with ${result.errors.length} import errors`);
}
```
Two option defaults are exactly right here and worth stating because they are not
obvious:
- **`includeDeleted` defaults to `false`, and the copy clones live state — it
does not synchronize deletions.** The exporter simply omits soft-deleted rows;
it cannot round-trip `deletedAt` at all (the wire format carries no deletion
flag). So a fact deleted on the source is merely *absent* from the stream: on a
fresh target it never appears, but on a populated target an existing live row
**stays live** — the copy never deletes it. If the target must reflect
deletions, apply them through your projector, not the bulk copy.
- **`includeTemporal` must be set to `true`** (it defaults to `false`). It is
what carries each fact's original `validFrom` / `validTo` across the copy;
without it the import re-stamps every fact with the *copy's* wall clock,
destroying valid-time fidelity in the merged branch.
`importGraphStream` preserves ids, routes existing rows through normal
`onConflict` handling, validates edge endpoints (`validateReferences` defaults to
`true`), and refreshes planner statistics once after the import commits.
## Cursor-Based Resumption and Electric
The examples above assume a per-change offset. **Electric does not provide one** —
every change in a `ShapeStream` catch-up batch shares the stream's
`lastOffset`. A cursor keyed on Electric's offset can therefore only advance at a
**batch boundary**, after the whole batch is projected. Advancing mid-batch is
unsafe: Electric's `read(after)` is strictly-after, so resuming from a
mid-batch offset permanently skips that batch's remaining changes. Project the
whole batch, then checkpoint the cursor once at its boundary.
# Multiple Graphs
> Using separate graph definitions for different domains in the same application
TypeGraph supports multiple graphs for applications that have distinct data domains that benefit from separate graph definitions.
## When to Use Multiple Graphs
Use separate graphs when you have:
- **Distinct domains**: A RAG system for documents and a business network for suppliers have different node types,
edge semantics, and query patterns
- **Independent lifecycles**: One graph might evolve rapidly while another is stable
- **Team ownership**: Different teams own different graphs, with separate schema review processes
- **Different retention policies**: Document chunks might be ephemeral while business relationships are long-lived
**Don't use multiple graphs** when:
- You need cross-graph queries or traversals (use a single graph with ontology relations instead)
- The domains are closely related (e.g., Users and Documents that Users author)
- You're trying to solve multi-tenancy (use tenant isolation patterns instead)
## Example: Documents and Business Network
A company needs two graphs:
1. **Documents graph**: Powers semantic search over internal documents
2. **Organization graph**: Tracks suppliers, partners, and contracts
### Defining the Graphs
```typescript
// graphs/documents.ts
import { z } from "zod";
import { defineNode, defineEdge, defineGraph, embedding } from "@nicia-ai/typegraph";
const Document = defineNode("Document", {
schema: z.object({
title: z.string(),
source: z.string(),
createdAt: z.string().datetime(),
}),
});
const Chunk = defineNode("Chunk", {
schema: z.object({
content: z.string(),
embedding: embedding(1536),
position: z.number().int(),
}),
});
const hasChunk = defineEdge("hasChunk");
export const documentsGraph = defineGraph({
id: "documents",
nodes: {
Document: { type: Document },
Chunk: { type: Chunk },
},
edges: {
hasChunk: { type: hasChunk, from: [Document], to: [Chunk] },
},
});
```
```typescript
// graphs/organization.ts
import { z } from "zod";
import { defineNode, defineEdge, defineGraph, subClassOf } from "@nicia-ai/typegraph";
const Organization = defineNode("Organization", {
schema: z.object({
name: z.string(),
domain: z.string().optional(),
}),
});
const Supplier = defineNode("Supplier", {
schema: z.object({
name: z.string(),
domain: z.string().optional(),
category: z.enum(["materials", "services", "logistics"]),
}),
});
const Partner = defineNode("Partner", {
schema: z.object({
name: z.string(),
domain: z.string().optional(),
partnershipLevel: z.enum(["bronze", "silver", "gold"]),
}),
});
const Contract = defineNode("Contract", {
schema: z.object({
title: z.string(),
value: z.number(),
startDate: z.string().datetime(),
endDate: z.string().datetime().optional(),
status: z.enum(["draft", "active", "expired"]).default("draft"),
}),
});
const supplies = defineEdge("supplies");
const hasContract = defineEdge("hasContract");
export const organizationGraph = defineGraph({
id: "organization",
nodes: {
Organization: { type: Organization },
Supplier: { type: Supplier },
Partner: { type: Partner },
Contract: { type: Contract },
},
edges: {
supplies: { type: supplies, from: [Supplier], to: [Organization] },
hasContract: { type: hasContract, from: [Organization], to: [Contract] },
},
ontology: [
subClassOf(Supplier, Organization),
subClassOf(Partner, Organization),
],
});
```
### Creating Stores
Both graphs can share the same database backend. Each graph's data is isolated by its `id`.
```typescript
// stores.ts
import { createStore } from "@nicia-ai/typegraph";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { drizzle } from "drizzle-orm/node-postgres";
import { Pool } from "pg";
import { documentsGraph } from "./graphs/documents";
import { organizationGraph } from "./graphs/organization";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
const db = drizzle(pool);
const backend = createPostgresBackend(db);
// Same backend, different stores
export const documentsStore = createStore(documentsGraph, backend);
export const organizationStore = createStore(organizationGraph, backend);
```
### Using the Stores
Each store is fully independent with its own typed API:
```typescript
// Semantic search in documents
async function searchDocuments(query: string, embedding: number[]) {
return documentsStore
.query()
.from("Chunk", "c")
.whereNode("c", (c) => c.embedding.similarTo(embedding, 10))
.select((ctx) => ({
content: ctx.c.content,
position: ctx.c.position,
}))
.execute();
}
// Business queries in organization
async function getActiveSuppliers(category: string) {
return organizationStore
.query()
.from("Supplier", "s")
.whereNode("s", (s) => s.category.eq(category))
.traverse("hasContract", "e")
.to("Contract", "c")
.whereNode("c", (c) => c.status.eq("active"))
.select((ctx) => ({
supplier: ctx.s.name,
contract: ctx.c.title,
value: ctx.c.value,
}))
.execute();
}
```
## Coordinating Across Graphs
Since cross-graph queries aren't supported, coordinate at the application level.
### Shared Identifiers
Use consistent IDs when entities relate across graphs:
```typescript
// When ingesting a supplier's documents, use the supplier ID as a reference
async function ingestSupplierDocument(
supplierId: string,
title: string,
content: string,
embedding: number[]
) {
// Store document with supplier reference in metadata
const doc = await documentsStore.nodes.Document.create({
title,
source: `supplier:${supplierId}`,
createdAt: new Date().toISOString(),
});
const chunk = await documentsStore.nodes.Chunk.create({
content,
embedding,
position: 0,
});
await documentsStore.edges.hasChunk.create(doc, chunk, {});
return doc;
}
// Later, find documents for a supplier
async function getSupplierDocuments(supplierId: string) {
return documentsStore
.query()
.from("Document", "d")
.whereNode("d", (d) => d.source.eq(`supplier:${supplierId}`))
.select((ctx) => ctx.d)
.execute();
}
```
### Application-Level Joins
Combine results from multiple graphs in your application:
```typescript
interface SupplierWithDocuments {
supplier: { name: string; category: string };
documents: Array<{ title: string }>;
}
async function getSupplierOverview(
supplierId: string
): Promise {
// Parallel queries to both graphs
const [supplier, documents] = await Promise.all([
organizationStore.nodes.Supplier.getById(supplierId),
getSupplierDocuments(supplierId),
]);
return {
supplier: {
name: supplier.name,
category: supplier.category,
},
documents: documents.map((d) => ({ title: d.title })),
};
}
```
### Event-Driven Sync
For loose coupling, use events to keep graphs in sync:
```typescript
// When a supplier is created, set up document ingestion
eventBus.on("supplier.created", async (event) => {
const { supplierId, name } = event.payload;
// Create a placeholder document node for future ingestion
await documentsStore.nodes.Document.create({
title: `${name} - Supplier Profile`,
source: `supplier:${supplierId}`,
createdAt: new Date().toISOString(),
});
});
// When a supplier is deleted, clean up related documents
eventBus.on("supplier.deleted", async (event) => {
const { supplierId } = event.payload;
const docs = await documentsStore
.query()
.from("Document", "d")
.whereNode("d", (d) => d.source.eq(`supplier:${supplierId}`))
.select((ctx) => ctx.d.id)
.execute();
for (const docId of docs) {
await documentsStore.nodes.Document.delete(docId);
}
});
```
## Separate Backends
For stronger isolation, use separate database connections:
```typescript
// Documents in PostgreSQL with pgvector for embeddings
const documentsPool = new Pool({
connectionString: process.env.DOCUMENTS_DATABASE_URL,
});
const documentsBackend = createPostgresBackend(drizzle(documentsPool));
export const documentsStore = createStore(documentsGraph, documentsBackend);
// Organization data in a separate database
const orgPool = new Pool({
connectionString: process.env.ORG_DATABASE_URL,
});
const orgBackend = createPostgresBackend(drizzle(orgPool));
export const organizationStore = createStore(organizationGraph, orgBackend);
```
**When to separate backends:**
- Different performance profiles (vector search vs. relational queries)
- Compliance requirements (PII in one database, analytics in another)
- Independent scaling needs
- Different backup/retention policies
## Schema Management
Each graph has independent schema versioning:
```typescript
import { createStoreWithSchema } from "@nicia-ai/typegraph";
// Each graph tracks its own schema version
const [documentsStore, docsSchemaResult] = await createStoreWithSchema(
documentsGraph,
backend
);
const [orgStore, orgSchemaResult] = await createStoreWithSchema(
organizationGraph,
backend
);
// Check migration status independently
if (docsSchemaResult.status === "migrated") {
console.log("Documents schema was migrated");
}
if (orgSchemaResult.status === "migrated") {
console.log("Organization schema was migrated");
}
```
## Inspecting What a Database Holds
Graphs sharing a backend are separated by `graph_id` inside TypeGraph's tables. Two reads answer the
questions an operator asks about that layout without depending on it: which graphs live in this
database, and how many rows one graph holds.
### `listGraphIds(backend, options?)`
Lists the graph ids that hold data, one bounded page at a time:
```typescript
import { listGraphIds } from "@nicia-ai/typegraph";
let after: string | undefined;
for (;;) {
const page = await listGraphIds(backend, { prefix: "tenant-", after, limit: 100 });
if (page.length === 0) break;
for (const graphId of page) console.log(graphId);
after = page.at(-1);
}
```
| Option | Meaning |
| -------- | ----------------------------------------------------------------------------------- |
| `prefix` | Only ids starting with this exact, case-sensitive text. `%` and `_` are not wildcards. |
| `after` | Exclusive cursor: only ids ordered after this one. Pass the last id of the previous page. |
| `limit` | Page size from 1 to 1000. Defaults to 100. Anything else throws `ConfigurationError`. |
Ids come back in byte order (UTF-8 code point order) on every backend, so `Tenant-x` sorts before
`tenant-a` on SQLite and PostgreSQL alike and a cursor resumes exactly where the last page ended,
whatever the database collation. The reserved deployment marker id that TypeGraph uses for
deployment-scoped contribution markers is never listed. A graph appears while it has nodes, edges or a
committed schema version, which are exactly the relations a default `store.clear()` empties. A
cleared graph therefore stops being listed even though `store.clear()` keeps its contribution
markers unless you pass `preserveContributionMaterializations: false`, and even when a
revision-tracked store reseeded its `recordedClock` row during the clear.
Each page walks graph ids by index seek, one seek per graph per relation, instead of reading every
row. The walk starts at the cursor or prefix and stops after the page, so a page costs about `limit`
seeks wherever it sits, however many graphs the database holds and however many rows they contain.
SQLite serves the seeks from the `graph_id`-leading primary keys, which are already in byte order.
PostgreSQL orders ordinary text indexes by the database collation, so it serves them from a
byte-ordered (`COLLATE "C"`) `graph_id` index that base-schema version 5 adds to `nodes`, `edges` and
`schema_versions`; see [Base-schema version 5](/backend-setup#base-schema-version-5-byte-ordered-graph_id-indexes-postgresql)
for what it costs and how to build it ahead of an upgrade. On a 20,000-graph, 50-rows-per-graph
PostgreSQL 18 database a page takes about 3 ms, where the same read took about 400 ms before the
index; at 200 graphs of 5,000 rows it is about 3 ms either way. A database without the index (its
base schema not adopted yet, or DDL managed by hand) lists the same ids by reading and de-duplicating
every row of those relations for each page, measured at 47 to 105 ms a page at these sizes. A backend that
declares no recursive traversal does the same. Use the listing for operator tooling, not on a request
path. Rows that exist only outside those relations, such as orphaned recorded history or contribution
markers, do not make a graph appear; `inspectGraphStorage` counts every relation.
The read runs in one read-only transaction where the backend supports it. It needs the backend's
catalog probes to tell a table that was never provisioned from an empty one, and throws
`ConfigurationError` on a custom backend that has none.
### `inspectGraphStorage(store)`
Counts one graph's rows in every relation that can hold them:
```typescript
import { inspectGraphStorage } from "@nicia-ai/typegraph";
await store.clear();
const { graphId, relations, totalRows } = await inspectGraphStorage(store);
const leftovers = relations.filter((relation) => relation.rows > 0);
// [{ relation: "contributionMaterializations", table: "typegraph_contribution_materializations", rows: 2 }]
```
`relations` lists every graph-scoped relation under its logical key (`nodes`, `edges`, `uniques`,
`edgeClaims`, `identityAssertions`, `recordedNodes`, `fulltext`, `schemaVersions`, and so on) with
the physical `table` it resolved to on this backend, so custom table names are reported as
configured. The graph's per-field vector tables come from its vector slots and the active vector
strategy and are reported as `vector:.`. A relation whose table the database never
provisioned counts as `0` rather than failing.
Use it to verify that `store.clear()` left nothing behind. Two relations can legitimately hold a
row after a clear, by design:
- `contributionMaterializations` is preserved unless you pass
`preserveContributionMaterializations: false`.
- `recordedClock` is reseeded inside the clear transaction on a store with live revision tracking
(without history).
Every other relation reads `0` after a clear, and other graphs in the same database are untouched.
#### Consistency of the counts
Each relation is counted by its own statement, so `relations` and `totalRows` describe one state of
the graph only when every statement read the same snapshot. The result carries a `consistency`
field that says whether they did:
| `consistency` | Meaning |
| --- | --- |
| `"snapshot"` | Every count came from one snapshot: `relations` and `totalRows` describe a state the graph was in. |
| `"per-statement"` | Each relation was counted independently. A write between two counts can leave the result describing a state that never existed together, for example rows in `nodes` beside an empty `schemaVersions`. |
The read asks for a read-only `repeatable read` transaction, but it does not trust the request: a
transaction wrapper can drop the isolation option, and a role or database can default the level.
The effective level is read on the counting session itself, inside the first count statement, so
the answer costs no extra round trip. What that gives on each backend:
- **SQLite (better-sqlite3, libSQL, and other drivers with interactive transactions):**
always `"snapshot"`. A SQLite transaction reads one snapshot whatever level was requested.
- **PostgreSQL (`pg`, `postgres-js`, PGlite):** `"snapshot"` when the session was observed at
`repeatable read` or `serializable`, which is what the request produces. `"per-statement"` when
it ran at `read committed`, which happens when a wrapper around `backend.transaction` does not
forward its options and the role or database defaults to `read committed`; the same wrapper
under a `repeatable read` default still reports `"snapshot"`, because the level is observed, not
requested. A backend that declares no session isolation read cannot be observed and reports
`"per-statement"`.
- **Backends without interactive transactions (Cloudflare D1, `neon-http`):** `"per-statement"`,
because there is no transaction to share a snapshot. The exception is a graph with at most one
provisioned relation, which is one statement and so trivially consistent.
The evidence proves the isolation of the session that ran the first count. A backend wrapper that
violates the transaction contract by handing the root pool through as its transaction backend can
run later counts on other sessions, which no observation on the first one can detect.
The read never refuses on a weaker level: it is a diagnostic. Treat `"per-statement"` counts as an
approximation. To verify a clear with them, make sure nothing else writes the graph while you read,
or read twice and compare.
## Shared Subgraph Helpers
When multiple graphs share a common set of node and edge types, you can write reusable
helpers that accept any store containing that shared subgraph. The `StoreProjection` utility
type makes this type-safe without coupling to a specific graph definition.
### Defining shared types and graphs
Start with the shared node and edge types, then define the graphs that use them:
```typescript
import {
createStore,
defineNode,
defineEdge,
defineGraph,
type Node,
type StoreProjection,
} from "@nicia-ai/typegraph";
const Document = defineNode("Document", {
schema: z.object({ title: z.string() }),
});
const Chunk = defineNode("Chunk", {
schema: z.object({ text: z.string() }),
});
const Comment = defineNode("Comment", {
schema: z.object({ text: z.string() }),
});
const hasChunk = defineEdge("hasChunk", { from: [Document], to: [Chunk] });
const aboutChunk = defineEdge("aboutChunk", { from: [Comment], to: [Chunk] });
const reviewGraph = defineGraph({
id: "review",
nodes: {
Document: { type: Document },
Chunk: { type: Chunk },
Comment: { type: Comment },
Label: { type: Label },
},
edges: { hasChunk, aboutChunk, hasLabel },
});
const catalogGraph = defineGraph({
id: "catalog",
nodes: {
Document: {
type: Document,
unique: [{ name: "title_unique", fields: ["title"], scope: "kind", collation: "binary" }],
},
Chunk: { type: Chunk },
Comment: { type: Comment },
Category: { type: Category },
},
edges: { hasChunk, aboutChunk, inCategory },
});
```
### Projecting a shared subgraph
Define a projection against either graph — it picks only the shared keys:
```typescript
type CoreStore = StoreProjection<
typeof reviewGraph,
"Document" | "Chunk" | "Comment",
"hasChunk" | "aboutChunk"
>;
```
### Writing a reusable helper
```typescript
async function addComment(
store: CoreStore,
chunk: Node,
text: string,
) {
const comment = await store.nodes.Comment.create({ text });
await store.edges.aboutChunk.create(comment, chunk);
return comment;
}
```
### Using across different graphs
The same `addComment` function works with any store whose graph includes the projected
nodes and edges — even if the graphs diverge on other types or unique constraints:
```typescript
const reviewStore = createStore(reviewGraph, backend);
const catalogStore = createStore(catalogGraph, backend);
await addComment(reviewStore, chunk, "needs revision");
await addComment(catalogStore, chunk, "good categorization");
```
The projection also works inside transactions — `TransactionContext` is structurally
assignable to `StoreProjection` for the same keys:
```typescript
await reviewStore.transaction(async (tx) => {
await addComment(tx, chunk, "transactional comment");
});
```
### What the projection strips
`StoreProjection` erases node constraint names, making constraint-based methods like
`findByConstraint` uncallable through the projection. This is intentional: unique
constraints are graph-registration-level details that typically differ between graphs
sharing the same node types. If you need constraint access, type the helper against a
specific `Store` instead.
## Caveats
**No cross-graph queries**: You cannot traverse from a node in one graph to a node in another. If you need this, consider:
- Merging the graphs into one with clear ontology separation
- Using application-level joins as shown above
**Separate ontology closures**: Each graph computes its own `subClassOf`, `implies`, etc. closures. Ontology relations
don't span graphs.
**Independent transactions**: A transaction in one store doesn't include the other. For cross-graph consistency, use
sagas or eventual consistency patterns.
**Shared tables**: When using the same backend, both graphs write to the same `typegraph_nodes` and `typegraph_edges`
tables, differentiated by `graph_id`. This is fine for most cases but means a database-level issue affects both
graphs.
## Next Steps
- [Multi-Tenant SaaS](./examples/multi-tenant) - Isolating data by tenant within a single graph
- [Schema Migrations](./schema-management) - Versioning and migrations
- [Integration Patterns](./integration) - More deployment strategies
# Provenance and Retraction
> Track source lineage for derived facts, retract bad sources, and use recorded time to replay what the graph believed before and after the transition.
Provenance and Retraction is the TypeGraph subpath for source lineage and
belief transitions. It maps your ordinary graph kinds onto four roles:
- one or more retractable source node kinds with a boolean `retracted` flag
- a justification node that represents an AND support rule
- one or more derived fact node kinds
- two typed edges: premises point to justifications, and justifications derive facts
The API lives at `@nicia-ai/typegraph/provenance`:
```typescript
import { createRetractionCapability } from "@nicia-ai/typegraph/provenance";
const provenance = createRetractionCapability(store, {
source: { kind: "Source" },
justification: { kind: "Justification" },
fact: { kinds: ["Fact"] },
premiseOf: { kind: "premiseOf" },
derives: { kind: "derives" },
});
```
Use `source: { kinds: [...] }` when different source node kinds share the same
boolean retraction field:
```typescript
const provenance = createRetractionCapability(store, {
source: { kinds: ["ScannerSource", "VendorSource"] },
justification: { kind: "Justification" },
fact: { kinds: ["Vulnerability", "DeployDecision"] },
premiseOf: { kind: "premiseOf" },
derives: { kind: "derives" },
});
```
`store` must be created with `{ history: true }`. Retraction mutates graph row
currency, so TypeGraph-managed recorded capture is required:
```typescript
const [store] = await createStoreWithSchema(graph, backend, {
history: true,
});
```
For a complete runnable version, see
[Provenance Retraction](/examples/provenance-retraction).
## Graph shape
Define the roles as normal TypeGraph nodes and edges.
```typescript
const Source = defineNode("Source", {
schema: z.object({
label: z.string(),
retracted: z.boolean().default(false),
}),
});
const Fact = defineNode("Fact", {
schema: z.object({ label: z.string() }),
});
const TerminalFact = defineNode("TerminalFact", {
schema: z.object({ label: z.string() }),
});
const Justification = defineNode("Justification", {
schema: z.object({ label: z.string() }),
});
const premiseOf = defineEdge("premiseOf");
const derives = defineEdge("derives");
const graph = defineGraph({
id: "claims",
nodes: {
Source: { type: Source },
Fact: { type: Fact },
TerminalFact: { type: TerminalFact },
Justification: { type: Justification },
},
edges: {
premiseOf: { type: premiseOf, from: [Source, Fact], to: [Justification] },
derives: { type: derives, from: [Justification], to: [Fact, TerminalFact] },
},
});
```
A justification fires when all of its premise nodes are in the well-founded
support set. Sources are in support unless their `retracted` flag is true. Facts
enter support when at least one firing justification derives them.
Fact kinds only need to appear in `premiseOf.from` if they can support another
justification. Terminal facts can be listed in `fact.kinds` and `derives.to`
without being valid premise endpoints.
## Retraction
`retract(source)` sets the source flag, recomputes support from the current
provenance graph, and makes unsupported facts non-current. A transition only
touches facts reachable from the flipped sources, and closing a fact is a
belief-status change, not a domain delete: none of the fact's edges are
deleted (its `onDelete` behavior is not enforced), so `unRetract` restores the
fact exactly as it was.
```typescript
const before = await store.recordedNow();
const report = await provenance.retract({ kind: "Source", id: sourceId });
const after = await store.recordedNow();
const previous = before ? store.asOfRecorded(before) : undefined;
const current = after ? store.asOfRecorded(after) : undefined;
```
The report partitions facts relative to the retracted source:
- `died`: facts that were believed before and lost grounded support
- `survivedVia`: affected facts that still have a firing justification
- `unaffected`: previously believed facts outside the source's provenance
`unRetract(source)` clears the source flag, recomputes support, and reopens
facts that regain support.
Use `retractMany(sources)` or `unRetractMany(sources)` to change several source
flags in one recorded transaction:
```typescript
const report = await provenance.retractMany([
{ kind: "ScannerSource", id: scannerId },
{ kind: "VendorSource", id: vendorId },
]);
```
## Recorded time
Retraction uses TypeGraph-managed writes, so before and after states are visible
through recorded-time reads. On PostgreSQL, provenance transitions serialize
with TypeGraph-managed history writes on the same graph before computing and
applying fact currency. Capture is scoped to TypeGraph-managed writes; it does
not claim to observe out-of-band database mutations.
```typescript
const factBefore = before ? await store.asOfRecorded(before).nodes.Fact.getById(factId) : undefined;
const factAfter = after ? await store.asOfRecorded(after).nodes.Fact.getById(factId) : undefined;
```
Use `holding()` when you only need the current well-founded believed facts:
```typescript
const facts = await provenance.holding();
```
# Execute
> Running queries with execute(), paginate(), and stream()
Execute operations run your query and retrieve results. Use `execute()` for simple queries,
`paginate()` for cursor-based pagination, and `stream()` for processing large datasets.
## execute()
Run the query and return all results:
```typescript
const results = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ctx.p)
.execute();
// results: readonly Person[]
```
### Return Type
Returns a readonly array of the selected type:
```typescript
// TypeScript infers the shape from your selection
const results = await store
.query()
.from("Person", "p")
.select((ctx) => ({
name: ctx.p.name,
email: ctx.p.email,
}))
.execute();
// results: readonly { name: string; email: string | undefined }[]
```
## executeChecked(expectedSchemaVersion)
Check a cached reconciled schema while reading data in one SQL statement:
```typescript
const rows = await store.query()
.from("Person", "person")
.select((ctx) => ctx.person)
.executeChecked(store.reconciledSchema.version);
```
The active schema version and data come from the same statement snapshot. A mismatch throws
`SchemaChangedError` with `details.graphId`, `details.expected`, and `details.actual` before
calling the selector, even when the data query matches no rows. `undefined` means no active
schema; it is distinct from version zero. On mismatch, reload the reconciled schema, rebuild
the query against the reopened store, and retry. Retry in a new transaction if the old one
holds a repeatable-read snapshot.
This is an explicit alternative to a standalone `getCommittedSchemaVersion` probe on the first
relational query. It checks that statement only; subsequent request reads can observe later
commits. It neither locks the schema nor replaces write fences, and does not alter store-open
or application cache policies.
Checked reads fetch full rows and support relational traversals, ordering, offsets, and limits.
Recursive and relevance-ranked queries are refused with `ConfigurationError`; use a separate
probe for those. Named parameters must be bound as ordinary values before building the query.
Bundled SQLite and PostgreSQL backends provide the required `tableNames.schemaVersions` binding.
A custom backend without it is refused before executing SQL. A custom binding must name a
relation with the standard `graph_id`, `version`, and `is_active` columns and one active row
per graph, consistent with `getActiveSchema`.
## first()
Get the first selected result or `undefined`. An existing `limit(0)` remains empty;
`offset()` is preserved. Add `orderBy()` when the choice of first row must be deterministic.
Only the returned row is passed to the selector:
```typescript
const alice = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.email.eq("alice@example.com"))
.select((ctx) => ctx.p)
.first();
if (alice) {
console.log(alice.name);
}
```
## count()
Count matching SQL rows without fetching their data or running a `select()` callback.
`count()` and `exists()` are available before and after `select()`. They preserve
`groupBy()`, `having()`, `limit()`, and `offset()`: a grouped query counts groups,
and `limit(0).count()` returns zero. Traversals count match rows, so multiple
relationships can count the same node more than once. Bind named parameters as
concrete values before using these terminals:
```typescript
const activeCount = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.count();
// activeCount: number
```
## exists()
Check if any results exist:
```typescript
const hasActiveUsers = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.exists();
// hasActiveUsers: boolean
```
## Cursor Pagination
Use `first`/`after` for forward pages or `last`/`before` for backward pages.
Page sizes must be positive safe integers. Do not combine directions or add
query-level `limit()`/`offset()` to a paginated or streamed query; those bounds
are refused rather than silently discarded. Use `limit()`/`offset()` with
`execute()` for offset pagination.
For large datasets, cursor-based pagination is more efficient than `limit`/`offset`. It uses keyset
pagination which doesn't degrade as you go deeper.
### paginate()
Nullable sort values follow the same ordering across page boundaries as in `execute()`:
ascending order places missing values last, and descending order places them first. Forward and
backward cursors retain rows in both the missing-value and non-missing-value groups. Sort fields
do not have to appear in the selected result; pagination retains them internally for its cursors.
```typescript
const firstPage = await store
.query()
.from("Person", "p")
.select((ctx) => ({
id: ctx.p.id,
name: ctx.p.name,
}))
.orderBy("p", "name", "asc") // ORDER BY required
.paginate({ first: 20 });
```
### Pagination Result Shape
```typescript
{
data: readonly T[], // The actual results
hasNextPage: boolean, // More results available forward
hasPrevPage: boolean, // More results available backward
nextCursor: string | undefined, // Opaque cursor for next page
prevCursor: string | undefined, // Opaque cursor for previous page
}
```
### Forward Pagination
Use `first` and `after` to paginate forward:
```typescript
// Get first page
const page1 = await query.paginate({ first: 20 });
// Get next page using the cursor
if (page1.hasNextPage && page1.nextCursor) {
const page2 = await query.paginate({
first: 20,
after: page1.nextCursor,
});
}
```
### Backward Pagination
Use `last` and `before` to paginate backward:
```typescript
// Get last page
const lastPage = await query.paginate({ last: 20 });
// Get previous page
if (lastPage.hasPrevPage && lastPage.prevCursor) {
const prevPage = await query.paginate({
last: 20,
before: lastPage.prevCursor,
});
}
```
### Batch cursor pages with `page()`
`paginate()` executes immediately. Use `page()` to build the same cursor-bounded read without
executing it, so the page can compose with other independent reads in one `batchOnce()` statement:
```typescript
const people = store
.query()
.from("Person", "person")
.orderBy("person", "name")
.select((fields) => fields.person);
const [page, companies] = await store.batchOnce(() => [
people.page({ first: 20, after: cursor }),
store
.query()
.from("Company", "company")
.orderBy("company", "name")
.select((fields) => fields.company),
]);
```
The returned page has the same `PaginatedResult` shape as `paginate()`. A page read also has an
`execute()` method for independent execution. As with every `batchOnce()` member, all reads must
belong to the same graph and execution target, and the combined statement must fit the backend's
bind-parameter budget.
### Pagination Parameters
| Parameter | Type | Description |
| --------- | -------- | -------------------------------------------- |
| `first` | `number` | Number of results from the start |
| `after` | `string` | Cursor to start after (forward pagination) |
| `last` | `number` | Number of results from the end |
| `before` | `string` | Cursor to start before (backward pagination) |
### Pagination with Traversals
Pagination works with graph traversals:
```typescript
const employeesPage = await store
.query()
.from("Company", "c")
.whereNode("c", (c) => c.name.eq("Acme Corp"))
.traverse("worksAt", "e", { direction: "in" })
.to("Person", "p")
.select((ctx) => ({
id: ctx.p.id,
name: ctx.p.name,
role: ctx.e.role,
}))
.orderBy("p", "name", "asc")
.paginate({ first: 50 });
```
## Streaming
For very large datasets, use streaming to process results without loading everything into memory.
### stream()
```typescript
const stream = store
.query()
.from("Event", "e")
.select((ctx) => ctx.e)
.orderBy("e", "createdAt", "desc") // ORDER BY required
.stream({ batchSize: 1000 });
// Process results as they arrive
for await (const event of stream) {
console.log(event.title);
await processEvent(event);
}
```
### Batch Size
The `batchSize` option controls how many records are fetched per database query:
```typescript
// Smaller batches: Lower memory usage, more database queries
.stream({ batchSize: 100 })
// Larger batches: Higher memory usage, fewer database queries
.stream({ batchSize: 5000 })
// Default is 1000
.stream()
```
### Streaming with Processing
```typescript
async function exportAllUsers(): Promise {
const stream = store
.query()
.from("User", "u")
.whereNode("u", (u) => u.status.eq("active"))
.select((ctx) => ({
id: ctx.u.id,
email: ctx.u.email,
name: ctx.u.name,
}))
.orderBy("u", "id", "asc")
.stream({ batchSize: 500 });
let count = 0;
for await (const user of stream) {
await exportToExternalSystem(user);
count++;
if (count % 1000 === 0) {
console.log(`Exported ${count} users...`);
}
}
console.log(`Export complete: ${count} users`);
}
```
## Batch Execution
When independent reads must share one database round trip, use `store.batchOnce()`.
It embeds each read as a CTE and returns the independently typed results in input order. Fluent
queries preserve explicit ordering even when the sort field is not selected. The callback's scoped
builder creates batch-scoped composable graph reads without adding parallel `*Query` methods to the
executing Store API:
```typescript
const [people, neighbors, neighborhood] = await store.batchOnce((read) => [
store.query().from("Person", "p").select((ctx) => ctx.p),
read.neighbors(person, { edges: ["knows"], limit: 5 }),
read.subgraph(person.id, { edges: ["knows"], maxDepth: 2 }),
]);
```
The callback can also return `roots.map(...)`, a singleton, or an empty array. A nonempty batch is
one statement with no sequential fallback; an empty batch executes no SQL. At most 500 reads may be
planned, and the combined statement must fit the backend's bind-parameter budget. Every response is
materialized as JSON rather than streamed, so bound each member's result explicitly. Response size
is data-dependent; TypeGraph neither estimates it nor imposes a response-byte cap before execution.
For several independent subgraphs, use the runtime-array form to collapse their database round
trips into one statement:
```typescript
const subgraphs = await store.batchOnce((read) =>
roots.map((root) =>
read.subgraph(root.id, {
edges: ["knows"],
maxDepth: 2,
project: { nodes: { Person: ["name"] } },
}),
),
);
```
When compatible subgraphs have substantially overlapping neighborhoods and project meaningful payloads,
opt into shared traversal and hydration:
```typescript
const subgraphs = await store.batchOnce(
(read) =>
roots.map((root) =>
read.subgraph(root.id, {
edges: ["knows"],
maxDepth: 2,
project: { nodes: { Person: ["name", "profile"] } },
}),
),
{ shareSubgraphs: true },
);
```
The option groups only compatible subgraph reads and hydrates a shared entity once while preserving
an independent result object for every request. The default remains independent subgraph plans in
the same one statement. Sharing adds membership and reconstruction overhead, so enable it for
measured overlap and payload shapes rather than assuming it is universally faster. See
[shared subgraph examples](/performance/overview#choosing-shared-subgraphs) for overlapping
biographies, disjoint neighborhoods, and identity-only results with different tradeoffs.
Tuple members may use different roots, edge sets, depths, windows, and projections when a page
needs heterogeneous neighborhoods:
```typescript
const [social, employment] = await store.batchOnce((read) => [
read.subgraph(person.id, {
edges: ["knows"],
maxDepth: 2,
edgeWindows: { knows: { limit: 20 } },
}),
read.subgraph(person.id, {
edges: ["worksAt"],
maxDepth: 1,
project: {
nodes: { Company: ["name", "industry"] },
edges: { worksAt: ["role"] },
},
}),
]);
```
The one-statement guarantee reduces round trips, which is often valuable for remote databases. By
default it does not combine recursive plans or share hydration between overlapping subgraphs, and
it does not promise less database work than direct `store.subgraph()` calls. The explicit
`shareSubgraphs` option changes that planning choice for compatible subgraph members only.
Use `store.batch()` when the batch includes queued edge collection `batchFind*` reads or when
sequential execution is the intended connection profile.
`batch()` does not batch round trips. The portable guarantee is that at most one query is in flight
at a time — at least one statement each, and two for a query whose selective-field mapping falls
back after its statement has already run. On a SQL backend with transactions it frames them with
`begin`/`commit`, putting a networked one at N+2 round trips **at best**; Durable Objects use an
ambient storage transaction with no framing, and without transactions there is no framing at all.
Connection reuse is the adapter's business either way.
Whole-node, whole-edge, and spread selections detected during planning use a full fetch from
the start. A selector branch that depends on actual row values can still trigger the fallback.
It will not merge arbitrary promises or collection calls. Use fluent queries or the callback's
`read.neighbors()`, `read.countNeighbors()`, and `read.subgraph()` methods when independent result
shapes must share its one statement. Other alternatives are a `.traverse()` chain,
`store.neighbors()` or `store.countNeighbors()` (one statement each), `store.subgraph()` (2
statements on SQLite, 3 on PostgreSQL), or `getByIds()` /
`bulkFindByIndex()`, which are chunked rather than fixed-cost.
Direct `store.subgraph()` and batch-scoped `read.subgraph()` share validation, traversal,
projection, and result semantics. The call context selects the physical execution contract: the
direct read uses backend-tuned hydration, while the batch-scoped read is embedded into the batch's
single statement.
It is also **not** a snapshot: PostgreSQL defaults to read-committed isolation, so a later query can
observe a commit the earlier ones did not. When several reads need one stable snapshot, run
`tx.query()`, `tx.neighbors()`, `tx.countNeighbors()`, `tx.subgraph()`, or `tx.batchOnce()` inside
`store.transaction(fn, { isolationLevel: "repeatable_read" })`. These reads are bound to the open
transaction and see its earlier uncommitted writes. Transactions require a backend with interactive
transaction support; a history-enabled store on PostgreSQL additionally requires
`accessMode: "read_only"` for a read-only transaction.
```typescript
const [people, companies] = await store.batch(
store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })),
store
.query()
.from("Company", "c")
.select((ctx) => ({ id: ctx.c.id, name: ctx.c.name }))
.orderBy("c", "name", "asc")
.limit(5),
);
// people: readonly { id: string; name: string }[]
// companies: readonly { id: string; name: string }[]
```
Each query preserves its own projection, filtering, sorting, and pagination. Results are returned
as a typed tuple matching the input order.
Edge collection `batchFind*` methods also return `BatchableQuery` and can be mixed freely with
fluent queries — each still costs its own statement:
```typescript
const [skills, employer] = await store.batch(
store.edges.hasSkill.batchFindFrom(alice),
store.edges.worksAt.batchFindFrom(alice),
);
```
**vs `Promise.all`**: workload- and adapter-dependent in both directions. `Promise.all` overlaps its
queries against a pool with idle capacity, but it does not necessarily hold N connections, and
against a single client or a saturated pool it queues. `batch()` keeps at most one query in flight,
so it pays the sum of their latencies — but it can still come out ahead where connection
acquisition dominates. Measure rather than assume.
**vs `transaction()`**: `batch()` may open an internal transaction only to serialize its statements.
Use `transaction()` when reads must share an explicit isolation level or see writes made earlier in
the callback. Its context supports fluent and set-oriented reads; `tx.batchOnce()` still emits
exactly one statement.
See [Batch Query Execution](/schemas-stores#batch-query-execution) for full API reference.
## Prepared Queries
Prepared queries let you build and structurally validate a query's AST once — so a malformed query
fails fast, before the first `.execute()` — and execute it many times with different parameter
values.
### `param(name)`
Use `param()` to declare a named placeholder inside any predicate position:
```typescript
import { param } from "@nicia-ai/typegraph";
```
### `prepare()`
Call `.prepare()` on an executable query to build and validate the AST once. Returns a
`PreparedQuery` that can be executed with different bindings. The statement is compiled once into
a cached template and reused by every `.execute()` call — see
[Prepared query SQL compilation](#prepared-query-sql-compilation) below for how that stays fresh.
```typescript
const findByName = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.name.eq(param("name")))
.select((ctx) => ctx.p)
.prepare();
// Execute with different bindings
const alices = await findByName.execute({ name: "Alice" });
const bobs = await findByName.execute({ name: "Bob" });
```
### Parameterized Bounds
Parameters work anywhere a scalar value is accepted:
```typescript
const findByAge = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.age.between(param("minAge"), param("maxAge")))
.select((ctx) => ctx.p)
.prepare();
const youngAdults = await findByAge.execute({ minAge: 18, maxAge: 25 });
const seniors = await findByAge.execute({ minAge: 65, maxAge: 120 });
```
`prepared.execute(bindings)` validates bindings strictly: all declared parameters must be
provided, and unknown binding keys are rejected.
### Supported Positions
`param()` works with any scalar predicate:
| Predicate | Example |
| --------------------------- | ----------------------------------------- |
| `eq` / `neq` | `p.name.eq(param("name"))` |
| `gt` / `gte` / `lt` / `lte` | `p.age.gt(param("minAge"))` |
| `between` | `p.age.between(param("lo"), param("hi"))` |
| `contains` | `p.name.contains(param("substr"))` |
| `startsWith` / `endsWith` | `p.name.startsWith(param("prefix"))` |
| `like` / `ilike` | `p.email.like(param("pattern"))` |
`in()` and `notIn()` take a list-valued parameter — the **whole** list, not individual elements:
```typescript
const byIds = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.id.in(param("ids")))
.select((ctx) => ctx.p)
.prepare();
await byIds.execute({ ids: ["a", "b", "c"] });
await byIds.execute({ ids: ["d"] });
```
The list is bound as a single parameter that the database unpacks, so the compiled SQL text does
not depend on the list's length: one statement serves every arity, and a list of ten thousand ids
still costs one bound parameter rather than blowing past the engine's bind limit. An empty list is
valid — `in([])` matches nothing, `notIn([])` matches everything.
Every element must be the field's type, and numbers must be finite. A mixed list — `[1, "a"]` bound
against a number field — is rejected with a `ConfigurationError` before it reaches the database, on
every backend, as is `NaN` or `Infinity`. This matches the literal form, which already refuses a
mixed list, and it is what keeps the two backends in step: left unchecked, PostgreSQL would fail
casting while SQLite silently matched nothing.
:::caution
A `param()` sitting among the **elements** of a literal list — `p.name.in(["Alice", param("other")])` —
is rejected with an `UnsupportedPredicateError`. Bind the whole list instead.
:::
### Prepared Query SQL Compilation
`.prepare()` builds and validates the AST once. On a backend that can compile and run raw SQL text
(both the SQLite and PostgreSQL backends can), the statement is then compiled **once** into a cached
template and reused by every `.execute()` call.
The subtlety a cache like that has to survive is freshness: a "current" (live) read filters on
temporal validity as of the instant it runs, so caching a compiled statement that had a concrete
"now" baked into it would freeze that instant for the prepared query's entire lifetime — hiding
every row created after `.prepare()` from every subsequent call. The template therefore reserves the
read instant as a **placeholder** rather than a value, and each `.execute()` fills it with a fresh
instant alongside the call's own bindings. Nothing about the statement's text depends on either.
Two cases fall back to substituting parameters into the AST and compiling through the standard path
on every call — same results and the same freshness guarantee, without the cached-template fast
path:
- `executeRaw` is unavailable (a custom or async backend).
- The statement's execution semantics ride on the compiled SQL object rather than its text, which no
amount of `executeRaw` support changes. **Approximate vector search**
(`similarTo(..., { approximate: true })`) carries the engine's iterative-scan wrapper, and
`store.subgraph()` on PostgreSQL forces a custom plan for its id-array fetches. Flattening either
to cacheable text would drop the behavior it depends on, so both are excluded deliberately.
## Query Debugging
### toAst()
Get the query AST for inspection:
```typescript
const builder = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ctx.p);
const ast = builder.toAst();
console.log(JSON.stringify(ast, null, 2));
```
### compile()
Use `toSQL()` to render SQL for the Store's configured dialect without
executing it:
```typescript
const compiled = builder.toSQL();
console.log("SQL:", compiled.sql);
console.log("Parameters:", compiled.params);
```
For adapter and tooling authors, `builder.compile()` returns TypeGraph's
database-independent `CompiledSelectSql` fragment. It can be passed to a
`GraphBackend` or rendered explicitly with `renderSqlite()` or
`renderPostgres()`. It is intentionally not a Drizzle `SQL` object.
Useful for:
- Debugging query behavior
- Understanding performance characteristics
- Building custom query executors
## Ordering Requirements
Both `paginate()` and `stream()` require an `orderBy()` clause:
```typescript
// Required for pagination
.orderBy("p", "name", "asc")
.paginate({ first: 20 });
// Required for streaming
.orderBy("e", "createdAt", "desc")
.stream();
```
### Stable Ordering
Cursor pagination and streaming automatically append missing start-node identity keys: `id ASC`
for a single kind, or `kind ASC` and `id ASC` for a multi-kind source. Existing caller-specified
identity ordering is preserved. Offset pagination needs an explicit total ordering.
For deterministic offset pagination of one kind, include `id` in your ordering:
```typescript
.orderBy("p", "name", "asc")
.orderBy("p", "id", "asc") // Ensures stable ordering
```
## Real-World Examples
### Paginated API Endpoint
```typescript
async function listUsers(cursor?: string, limit = 20) {
const query = store
.query()
.from("User", "u")
.whereNode("u", (u) => u.status.eq("active"))
.select((ctx) => ({
id: ctx.u.id,
name: ctx.u.name,
email: ctx.u.email,
}))
.orderBy("u", "createdAt", "desc")
.orderBy("u", "id", "desc");
const result = cursor
? await query.paginate({ first: limit, after: cursor })
: await query.paginate({ first: limit });
return {
users: result.data,
nextCursor: result.nextCursor,
hasMore: result.hasNextPage,
};
}
```
### Batch Processing
```typescript
async function processAllOrders() {
const stream = store
.query()
.from("Order", "o")
.whereNode("o", (o) => o.status.eq("pending"))
.select((ctx) => ctx.o)
.orderBy("o", "createdAt", "asc")
.stream({ batchSize: 100 });
for await (const order of stream) {
try {
await fulfillOrder(order);
await store.nodes.Order.update(order.id, { status: "fulfilled" });
} catch (error) {
console.error(`Failed to process order ${order.id}:`, error);
}
}
}
```
### Infinite Scroll
```typescript
function useInfiniteUsers() {
const [users, setUsers] = useState([]);
const [cursor, setCursor] = useState();
const [hasMore, setHasMore] = useState(true);
async function loadMore() {
const result = await store
.query()
.from("User", "u")
.select((ctx) => ctx.u)
.orderBy("u", "name", "asc")
.paginate({ first: 20, after: cursor });
setUsers((prev) => [...prev, ...result.data]);
setCursor(result.nextCursor);
setHasMore(result.hasNextPage);
}
return { users, loadMore, hasMore };
}
```
## Next Steps
- [Order](/queries/order) - Ordering and limiting results
- [Shape](/queries/shape) - Output transformation
- [Overview](/queries/overview) - Query categories reference
# Filter
> Match constraints and completed-result filters
Filter operations reduce the result set based on property values. TypeGraph provides `whereNode()`
and `whereEdge()` for match constraints, plus `where()` for completed match rows.
## whereNode()
Filter nodes based on their properties:
```typescript
const engineers = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.role.eq("Engineer"))
.select((ctx) => ctx.p)
.execute();
```
### Parameters
```typescript
.whereNode(alias, predicateFunction)
```
| Parameter | Type | Description |
| ------------------- | ------------------------- | ---------------------------------------------- |
| `alias` | `string` | The node alias to filter (must exist in query) |
| `predicateFunction` | `(accessor) => Predicate` | Function that returns a predicate |
The predicate function receives a typed accessor for the node's properties.
## whereEdge()
Filter based on edge properties during traversals:
```typescript
const highPaying = await store
.query()
.from("Person", "p")
.traverse("worksAt", "e")
.whereEdge("e", (e) => e.salary.gte(100000))
.to("Company", "c")
.select((ctx) => ({
person: ctx.p.name,
company: ctx.c.name,
salary: ctx.e.salary,
}))
.execute();
```
### Parameters
```typescript
.whereEdge(alias, predicateFunction)
```
| Parameter | Type | Description |
| ------------------- | ------------------------- | ---------------------------------------------- |
| `alias` | `string` | The edge alias to filter (must exist in query) |
| `predicateFunction` | `(accessor) => Predicate` | Function that returns a predicate |
## Combining Predicates
### AND
Both conditions must be true:
```typescript
.whereNode("p", (p) =>
p.status.eq("active").and(p.role.eq("admin"))
)
```
### OR
Either condition can be true:
```typescript
.whereNode("p", (p) =>
p.role.eq("admin").or(p.role.eq("moderator"))
)
```
### NOT
Negate a condition:
```typescript
.whereNode("p", (p) =>
p.status.eq("deleted").not()
)
```
### Complex Combinations
Build complex logic with parenthetical grouping:
```typescript
.whereNode("p", (p) =>
p.status
.eq("active")
.and(p.role.eq("admin").or(p.role.eq("moderator")))
)
```
This evaluates as: `status = 'active' AND (role = 'admin' OR role = 'moderator')`
## Multiple Filters
Chain multiple `whereNode()` calls for AND logic:
```typescript
const activeManagers = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.whereNode("p", (p) => p.role.eq("Manager"))
.select((ctx) => ctx.p)
.execute();
```
This is equivalent to:
```typescript
.whereNode("p", (p) =>
p.status.eq("active").and(p.role.eq("Manager"))
)
```
## Filtering After Traversal
Filter nodes at any point in the query:
```typescript
const techCompanyEngineers = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.role.eq("Engineer"))
.traverse("worksAt", "e")
.to("Company", "c")
.whereNode("c", (c) => c.industry.eq("Technology"))
.select((ctx) => ({
person: ctx.p.name,
company: ctx.c.name,
}))
.execute();
```
`whereNode("c", ...)` constrains the traversal match itself. For a recursive traversal it applies
at every hop and prunes a branch as soon as a target fails. Use scoped-expression `where()` when
intermediate nodes may fail the condition but a later endpoint should still be returned:
```typescript
const activeEndpoints = await store
.query()
.from("Person", "start")
.traverse("knows", "edge")
.recursive({ maxHops: 5 })
.to("Person", "person")
.where((fields) => expr.eq(fields.person.active, expr.literal(true)))
.select((ctx) => ctx.person)
.execute();
```
On an optional alias, an ordinary comparison removes rows where that alias is absent. Use an
explicit null check when absent matches should remain.
## Common Predicates
Here are the most commonly used predicates. For complete reference, see [Predicates](/queries/predicates/).
### Equality
```typescript
p.name.eq("Alice"); // equals
p.name.neq("Bob"); // not equals
```
### Comparison
```typescript
p.age.gt(21); // greater than
p.age.gte(21); // greater than or equal
p.age.lt(65); // less than
p.age.lte(65); // less than or equal
p.age.between(18, 65); // inclusive range
```
### String Matching
```typescript
p.name.contains("ali"); // substring match
p.name.startsWith("A"); // prefix match
p.name.endsWith("ice"); // suffix match
p.email.like("%@example.com"); // SQL LIKE pattern
p.name.ilike("alice"); // case-insensitive LIKE
```
### Fulltext Search
For nodes with at least one field declared with `searchable()`, use the
node-level `$fulltext.matches()` for BM25-style ranked fulltext search.
See [Fulltext Search](/fulltext-search) for the full guide.
```typescript
d.$fulltext.matches("climate change", 20); // Top 20 by relevance
d.$fulltext.matches("quarterly earnings", 10, {
mode: "websearch", // Google-style syntax
});
```
Combine with any other predicate — fulltext composes with metadata
filters, graph traversal, and vector search:
```typescript
d.$fulltext.matches("climate", 20).and(d.tenantId.eq(tenant)).and(d.published.eq(true));
```
### Null Checks
```typescript
p.deletedAt.isNull(); // is null/undefined
p.email.isNotNull(); // is not null
```
### List Membership
```typescript
p.status.in(["active", "pending"]);
p.status.notIn(["archived", "deleted"]);
```
### Array Operations
```typescript
p.tags.contains("typescript");
p.tags.containsAll(["typescript", "nodejs"]);
p.tags.containsAny(["typescript", "rust", "go"]);
p.tags.isEmpty();
p.tags.isNotEmpty();
```
## Predicate Types by Field
The available predicates depend on the field type:
| Field Type | Key Predicates |
| -------------------------------- | --------------------------------------------------- |
| String | `eq`, `contains`, `startsWith`, `like`, `ilike` |
| Nodes with `searchable()` fields | `$fulltext.matches()` (node-level, not per-field) |
| Number | `eq`, `gt`, `gte`, `lt`, `lte`, `between` |
| Date | `eq`, `gt`, `gte`, `lt`, `lte`, `between` |
| Array | `contains`, `containsAll`, `containsAny`, `isEmpty` |
| Object | `get()`, `hasKey`, `pathEquals` |
| Embedding | `similarTo()` |
See [Predicates](/queries/predicates/) for complete documentation.
## Count and Existence Helpers
### Count Results
```typescript
const count: number = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.count();
```
### Check Existence
```typescript
const exists: boolean = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.email.eq("alice@example.com"))
.exists();
```
### Get First Result
```typescript
const alice = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.email.eq("alice@example.com"))
.select((ctx) => ctx.p)
.first();
if (alice) {
console.log(alice.name);
}
```
## Next Steps
- [Predicates](/queries/predicates/) - Complete predicate reference
- [Traverse](/queries/traverse) - Navigate relationships
- [Advanced](/queries/advanced) - Subqueries with `exists()` and `inSubquery()`
# Traverse
> Navigate relationships with traverse() and optionalTraverse()
Traversals let you navigate relationships in your graph. Instead of writing complex SQL joins,
describe the path you want to follow.
## Single-Hop Traversal
Follow one edge from a node to connected nodes:
```typescript
const employments = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.id.eq("alice-123"))
.traverse("worksAt", "e") // Follow worksAt edges
.to("Company", "c") // Arrive at Company nodes
.select((ctx) => ({
person: ctx.p.name,
company: ctx.c.name,
role: ctx.e.role, // Edge properties are accessible
}))
.execute();
```
## Parameters
### traverse()
```typescript
.traverse(edgeKind, edgeAlias, options?)
```
| Parameter | Type | Description |
|-----------|------|-------------|
| `edgeKind` | `string` | The edge kind to traverse |
| `edgeAlias` | `string` | Unique alias for referencing this edge |
| `options.direction` | `"out" \| "in"` | Traversal direction (default: `"out"`) |
| `options.expand` | `"none" \| "implying" \| "inverse" \| "all"` | Ontology edge expansion mode (default: `"inverse"`) |
| `options.from` | `string` | Fan-out from a different node alias |
### optionalTraverse()
```typescript
.optionalTraverse(edgeKind, edgeAlias, options?)
```
Uses the same options as `traverse()`, but returns optional edge/node values in the result context.
### to()
```typescript
.to(nodeKind, nodeAlias, options?)
```
| Parameter | Type | Description |
|-----------|------|-------------|
| `nodeKind` | `string` | The target node kind |
| `nodeAlias` | `string` | Unique alias for referencing this node |
| `options.includeSubClasses` | `boolean` | Include subclass kinds (default: `false`) |
## Direction
By default, traversals follow edges in their defined direction (from → to). Use `direction: "in"` to traverse backwards:
```typescript
// Edge definition: worksAt goes from Person → Company
// Forward: Find companies where Alice works
.from("Person", "p")
.traverse("worksAt", "e") // Person → Company
.to("Company", "c")
// Backward: Find people who work at Acme
.from("Company", "c")
.whereNode("c", (c) => c.name.eq("Acme"))
.traverse("worksAt", "e", { direction: "in" }) // Company ← Person
.to("Person", "p")
```
## Edge Properties
Edges can carry properties. Access them through the edge alias:
```typescript
const employments = await store
.query()
.from("Person", "p")
.traverse("worksAt", "e")
.to("Company", "c")
.select((ctx) => ({
person: ctx.p.name,
company: ctx.c.name,
role: ctx.e.role, // Edge property
salary: ctx.e.salary, // Edge property
startDate: ctx.e.startDate, // Edge property
}))
.execute();
```
### Edge Object Structure
Each edge provides these fields:
| Property | Type | Description |
|----------|------|-------------|
| `id` | `string` | Unique edge identifier |
| `kind` | `string` | Edge type name |
| `fromId` | `string` | ID of the source node |
| `toId` | `string` | ID of the target node |
| `meta.createdAt` | `string` | When the edge was created |
| `meta.updatedAt` | `string` | When the edge was last updated |
| `meta.deletedAt` | `string \| undefined` | Soft delete timestamp |
| `meta.validFrom` | `string \| undefined` | Temporal validity start |
| `meta.validTo` | `string \| undefined` | Temporal validity end |
| *schema props* | varies | Properties defined in edge schema |
### Filtering on Edge Properties
Use `whereEdge()` to filter based on edge values:
```typescript
const highPaying = await store
.query()
.from("Person", "p")
.traverse("worksAt", "e")
.whereEdge("e", (e) => e.salary.gte(100000))
.to("Company", "c")
.select((ctx) => ({
person: ctx.p.name,
company: ctx.c.name,
salary: ctx.e.salary,
}))
.execute();
```
## Multi-Hop Traversals
Chain traversals to follow multiple relationships:
```typescript
const projectTasks = await store
.query()
.from("Person", "person")
.whereNode("person", (p) => p.name.eq("Alice"))
.traverse("worksOn", "e1")
.to("Project", "project")
.traverse("hasTask", "e2")
.to("Task", "task")
.select((ctx) => ({
person: ctx.person.name,
project: ctx.project.name,
task: ctx.task.title,
}))
.execute();
```
Each hop starts from the previous node set and arrives at new nodes.
### Mixed Directions
Combine forward and backward traversals:
```typescript
const teamStructure = await store
.query()
.from("Person", "p")
.traverse("worksAt", "e1") // Forward: Person → Company
.to("Company", "c")
.traverse("manages", "e2", { direction: "in" }) // Backward: Person ← manages
.to("Person", "manager")
.select((ctx) => ({
employee: ctx.p.name,
company: ctx.c.name,
manager: ctx.manager.name,
}))
.execute();
```
## Optional Traversals
Use `optionalTraverse()` for LEFT JOIN semantics—include results even when the traversal has no matches:
```typescript
const peopleWithOptionalEmployer = await store
.query()
.from("Person", "p")
.optionalTraverse("worksAt", "e")
.to("Company", "c")
.select((ctx) => ({
person: ctx.p.name,
company: ctx.c?.name, // May be undefined if no employer
}))
.execute();
// Includes all people, even those without a worksAt edge
```
### Mixing Required and Optional
```typescript
const employeesWithOptionalManager = await store
.query()
.from("Person", "p")
.traverse("worksAt", "e1") // Required: must work at a company
.to("Company", "c")
.optionalTraverse("reportsTo", "e2") // Optional: might not have manager
.to("Person", "manager")
.select((ctx) => ({
employee: ctx.p.name,
company: ctx.c.name,
manager: ctx.manager?.name, // undefined for top-level employees
}))
.execute();
```
### Optional Edge Access
With optional traversals, the edge may be `undefined`:
```typescript
.select((ctx) => ({
person: ctx.p.name,
company: ctx.c?.name, // Node may be undefined
role: ctx.e?.role, // Edge may be undefined
salary: ctx.e?.salary,
}))
```
## Ontology-Aware Traversals
If your ontology defines edge implications, expand queries to include implying edges:
```typescript
// Ontology: implies(marriedTo, knows), implies(bestFriends, knows)
const connections = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.name.eq("Alice"))
.traverse("knows", "e", { expand: "implying" })
.to("Person", "other")
.select((ctx) => ctx.other.name)
.execute();
// Returns people connected via "knows", "marriedTo", or "bestFriends"
```
`expand: "implying"` is only reachable for endpoint-compatible implications:
every node kind an implying edge (e.g. `marriedTo`) allows on a side must be
assignable to a kind the implied edge (`knows`) allows on that same side.
`implies()` relations that don't satisfy this are rejected with a
`ConfigurationError` when the graph is built into a store, so an
`expand: "implying"` traversal can never fold in rows whose kind couldn't
actually satisfy the traversal's own endpoints. See
[Ontology → Edge Relationships](/ontology#edge-relationships) for details.
If your ontology defines inverse edge kinds, you can expand traversals to include inverse edges:
```typescript
// Ontology: inverseOf(manages, managedBy)
const relationships = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.name.eq("Alice"))
.traverse("manages", "e", { expand: "inverse" })
.to("Person", "other")
.select((ctx) => ({
name: ctx.other.name,
viaEdgeKind: ctx.e.kind,
}))
.execute();
// Traverses both "manages" and "managedBy"
```
You can combine both options:
```typescript
.traverse("knows", "e", { expand: "all" })
```
:::note[Default expansion mode]
The default expansion mode is `"inverse"`, meaning traversals automatically include inverse edge kinds
from your ontology. To opt out for a single traversal, pass `expand: "none"`. To change the default
for all traversals, set `queryDefaults.traversalExpansion` in `createStore` options.
:::
## Runtime-declared kinds
For kinds and edges added at runtime via [graph
extensions](/graph-extensions), use the string-keyed siblings
`fromDynamic` (covered on [Source](/queries/source#runtime-declared-kinds)),
`traverseDynamic`, `optionalTraverseDynamic`, and `toDynamic`:
```typescript
const rows = await store
.query()
.fromDynamic("Paper", "p")
.traverseDynamic("authoredBy", "a")
.toDynamic("Author", "u")
.whereNode("p", (p) => p.field("year").number().gte(2020))
.select((ctx) => ({ paper: ctx.p, author: ctx.u, edge: ctx.a }))
.execute();
```
Each method runtime-validates against the registry — typos throw
`KindNotFoundError`, and a `toDynamic` target that isn't a valid
endpoint for the current edge / direction throws `EndpointError`.
### The `.field()` discriminator
Predicate accessors on dynamic-declared aliases expose schema properties
through a `.field(name)` discriminator. `BaseFieldAccessor` methods
(`eq`, `isNull`, `in`, `notIn`) work directly. Type-specific predicates
sit behind one of:
| Discriminator | Returns |
| --- | --- |
| `.string()` | `StringFieldAccessor` (`gte`, `contains`, `like`, …) |
| `.number()` | `NumberFieldAccessor` (`gte`, `between`, …) |
| `.date()` | `DateFieldAccessor` |
| `.array()` | `ArrayFieldAccessor` |
| `.object()` | `ObjectFieldAccessor<...>` |
| `.embedding()` | `EmbeddingFieldAccessor` (`similarTo`) |
Each discriminator validates against the registered Zod schema at
query-build time and throws `TypeError` on mismatch:
```typescript
.whereNode("p", (p) => p.field("year").number().gte(2020)) // ✓
.whereNode("p", (p) => p.field("year").string().eq("2020")) // throws TypeError
.whereNode("p", (p) => p.field("yera").number().gte(2020)) // throws (unknown property)
// BaseFieldAccessor methods don't need a discriminator.
.whereNode("p", (p) => p.field("year").isNotNull()) // ✓
```
The same `.field()` API is available on edge accessors:
```typescript
.whereEdge("a", (e) => e.field("order").number().eq(1))
```
### Mixed typed and dynamic aliases
Typed and dynamic aliases interleave in one query. Each alias's
predicate accessor is resolved independently — typed aliases keep their
narrow accessors, dynamic aliases get `.field()`:
```typescript
const rows = await store
.query()
.from("Document", "d") // compile-time kind
.traverseDynamic("taggedWith", "e") // runtime edge
.toDynamic("Tag", "n") // runtime target
.whereNode("d", (d) => d.title.eq("the doc")) // typed: direct
.whereNode("n", (n) => n.field("label").string().eq("research")) // dynamic
.select((ctx) => ({ doc: ctx.d, tag: ctx.n }))
.execute();
```
A typed `traverse("knownEdge", "e")` followed by `toDynamic(target, "n")`
keeps `e` typed — `e.role.eq(...)` works without `.field()` because the
edge schema is known at compile time. Only aliases declared via
`fromDynamic` / `traverseDynamic` / `optionalTraverseDynamic` /
`toDynamic` go through the discriminator.
### Optional dynamic traversal
`optionalTraverseDynamic` is the LEFT-JOIN sibling. Source nodes
without a matching edge still surface, with the edge and target aliases
as `undefined`:
```typescript
const papersWithOptionalAuthor = await store
.query()
.fromDynamic("Paper", "p")
.optionalTraverseDynamic("authoredBy", "a")
.toDynamic("Author", "u")
.select((ctx) => ({
paperTitle: ctx.p.title,
authorName: ctx.u?.name, // undefined for orphan papers
order: ctx.a?.order,
}))
.execute();
```
## Real-World Examples
### Organizational Hierarchy
```typescript
const teamMembers = await store
.query()
.from("Person", "manager")
.whereNode("manager", (p) => p.name.eq("VP Engineering"))
.traverse("manages", "e")
.to("Person", "report")
.select((ctx) => ({
manager: ctx.manager.name,
report: ctx.report.name,
department: ctx.report.department,
}))
.execute();
```
### Social Graph
```typescript
const friends = await store
.query()
.from("Person", "me")
.whereNode("me", (p) => p.id.eq(currentUserId))
.traverse("follows", "e")
.to("Person", "friend")
.select((ctx) => ({
id: ctx.friend.id,
name: ctx.friend.name,
followedAt: ctx.e.createdAt,
}))
.orderBy("e", "createdAt", "desc")
.limit(50)
.execute();
```
### E-Commerce
```typescript
const orderDetails = await store
.query()
.from("Order", "o")
.whereNode("o", (o) => o.id.eq(orderId))
.traverse("contains", "e")
.to("Product", "p")
.select((ctx) => ({
product: ctx.p.name,
quantity: ctx.e.quantity,
unitPrice: ctx.e.unitPrice,
}))
.execute();
```
## Next Steps
- [Recursive](/queries/recursive) - Variable-length paths with `recursive()`
- [Filter](/queries/filter) - Filter nodes and edges with predicates
- [Shape](/queries/shape) - Transform output with `select()`
# Evolving Schemas in Production
> Step-by-step guide for safely evolving your graph schema across deployments
Your graph schema will change as your application grows. This guide covers how
to make those changes safely — from adding a field to renaming a node type.
For API reference, see [Schema Migrations](/schema-management). For evolving
the kind set itself **at runtime** (agent-induced kinds, plugin-supplied
kinds, multi-tenant kind sets), see [Graph Extensions](/graph-extensions).
## How Schema Evolution Works
When you call `createStoreWithSchema()`, TypeGraph:
1. Serializes your current graph definition
2. Compares it against the stored schema (by hash, then by diff)
3. **Safe changes** — auto-migrates and bumps the version
4. **Breaking changes** — throws `MigrationError` (or returns `status: "breaking"`)
The key insight: TypeGraph manages **schema metadata**, not data migration. When
you add an optional field, TypeGraph records that the schema now includes it. It
does not alter existing rows — Zod defaults handle that at read time.
## Safe Changes
These changes are backwards compatible and auto-migrate without intervention:
- Adding new node types
- Adding new edge types
- Adding optional properties (with defaults)
- Adding ontology relations
- Changing per-kind annotations (UI hints, audit policy, etc.)
- Changing graph-scoped annotations (display metadata, capabilities, etc.)
### Adding an Optional Property
```typescript
// Version 1
const Person = defineNode("Person", {
schema: z.object({
name: z.string(),
}),
});
// Version 2 — safe, auto-migrates
const Person = defineNode("Person", {
schema: z.object({
name: z.string(),
email: z.string().optional(),
}),
});
```
On startup, `createStoreWithSchema()` returns `status: "migrated"`. Existing
Person nodes return `email: undefined` — no data transformation needed.
### Adding a Node Type with Edges
```typescript
// Version 2 — add Company and worksAt in one deploy
const Company = defineNode("Company", {
schema: z.object({ name: z.string() }),
});
const worksAt = defineEdge("worksAt", {
schema: z.object({ role: z.string() }),
});
const graph = defineGraph({
id: "my_app",
nodes: {
Person: { type: Person },
Company: { type: Company },
},
edges: {
worksAt: { type: worksAt, from: [Person], to: [Company] },
},
});
```
This is a single safe migration. New node and edge types don't affect existing
data.
### Changing Annotations
The `annotations` field on `defineNode` and `defineEdge` is part of the canonical
schema, so any change bumps the schema version. Changes are classified as
`safe` — no data migration needed, only the schema document is updated.
```typescript
// Version 1
const Incident = defineNode("Incident", {
schema: z.object({ title: z.string() }),
annotations: {
ui: { titleField: "title", icon: "alert-triangle" },
},
});
// Version 2 — swap the icon, add audit policy
const Incident = defineNode("Incident", {
schema: z.object({ title: z.string() }),
annotations: {
ui: { titleField: "title", icon: "circle-alert" },
audit: { pii: false, retentionDays: 365 },
},
});
```
`getSchemaChanges()` reports each annotations-only change per kind:
```typescript
import { getSchemaChanges } from "@nicia-ai/typegraph/schema";
const diff = await getSchemaChanges(backend, graph);
for (const change of diff?.nodes ?? []) {
if (change.details.includes("Annotations")) {
console.log(`${change.kind}: annotations changed (${change.severity})`);
// → "Incident: annotations changed (safe)"
}
}
```
The hash is computed with stable sorted-key order at every depth, so
re-formatting the annotations object — or swapping sibling key order — does
not bump the version. Only structural or value changes do.
A few things worth knowing:
- Graphs that never set `annotations` produce identical canonical-form hashes
to graphs from before this field existed. Adoption requires no migration.
- The canonical form omits empty / default annotations, so absent,
explicit `undefined`, and explicit `{}` all hash identically — no migration
is triggered just by writing `annotations: {}`.
- Annotations values must be JSON-serializable (`bigint`, `function`, `Date`,
and other class instances are rejected at definition time).
See the [schemas-stores reference](/schemas-stores#per-kind-annotations) for the
full annotations contract.
### Rolling out graph-scoped annotations
Graph-scoped annotations use a top-level `SerializedSchema` field. During a
mixed-version rollout, an older schema writer can otherwise recommit a document
without a field it does not understand. Use this two-step deployment invariant:
1. Upgrade **every process that can write schema versions** to TypeGraph 0.54 or
newer, without adding graph annotations yet.
2. After no older schema writer remains, enable `defineGraph({ annotations })`
or `defineGraphExtension({ annotations })` and commit the safe schema change.
Readers may be upgraded independently, but the writer floor must be complete
before annotations are enabled. TypeGraph 0.54+ preserves unknown top-level
schema fields across parse-and-recommit cycles, so later additive metadata
slices follow the same rollout rule.
## Breaking Changes
These require explicit handling:
- Removing node or edge types
- Removing properties
- Adding required properties (no default)
- Renaming types or properties
TypeGraph will throw `MigrationError` by default. You have two options: fix
the schema to be backwards compatible, or use the expand-contract pattern.
## The Expand-Contract Pattern
For breaking changes, use a multi-deploy strategy. This is the same pattern
used in relational database migrations — deploy in phases so there's never a
moment where running code is incompatible with the schema.
### Renaming a Property
Rename `name` to `fullName` on Person in three deploys:
#### Deploy 1 — Expand: add the new property
```typescript
const Person = defineNode("Person", {
schema: z.object({
name: z.string(),
fullName: z.string().optional(), // New property, optional for now
}),
});
```
Safe migration. Then backfill existing data:
```typescript
const [store] = await createStoreWithSchema(graph, backend);
const people = await store.query(Person).execute();
for (const person of people) {
if (!person.properties.fullName) {
await store.nodes.Person.update(person.id, {
fullName: person.properties.name,
});
}
}
```
#### Deploy 2 — Switch: use the new property everywhere
Update all application code to read/write `fullName` instead of `name`. Both
properties still exist, so this deploy is safe.
#### Deploy 3 — Contract: remove the old property
```typescript
const Person = defineNode("Person", {
schema: z.object({
fullName: z.string(),
}),
});
```
This is a breaking change (removing `name`). Use `migrateSchema()` to force it:
```typescript
import { getSchemaChanges, migrateSchema } from "@nicia-ai/typegraph/schema";
const [store, result] = await createStoreWithSchema(graph, backend, {
throwOnBreaking: false,
});
if (result.status === "breaking") {
// We've already backfilled — safe to force migrate
const activeSchema = await backend.getActiveSchema(graph.id);
await migrateSchema(backend, graph, activeSchema!.version);
}
```
Two things `migrateSchema()` will not let you do by accident:
- **Drop a kind that still holds rows.** The commit is refused with a
`MigrationError` whose `details.reason` is `"kind-removal"`. Committing
would make those rows unreachable, and the next `materializeRemovals()`
would delete them — it re-derives removals by walking schema history, so
the drop is not reversible by putting the kind back. Export or delete the
rows first (see [Removing a Node Type](#removing-a-node-type)), or pass
`{ discardDroppedKindRows: true }` if losing them is the intent. Dropping
an *empty* kind needs no flag.
- **Erase kinds added at runtime.** `migrateSchema()` folds the persisted
graph extension into the graph you hand it, the same way
`createStoreWithSchema()` does, so passing your compile-time graph never
drops a kind that `evolve()` committed. To remove one of those
deliberately, use `removeKinds()` — it queues the cleanup rows that make
the removal reconcilable.
### Removing a Node Type
#### Deploy 1 — Stop creating new instances
Update application code to stop creating the deprecated node type. Existing data
remains.
#### Deploy 2 — Clean up references
Delete edges that reference the deprecated node type, then delete the nodes
themselves:
```typescript
// Delete all edges connected to deprecated nodes
const deprecated = await store.query(OldNode).execute();
for (const node of deprecated) {
await store.nodes.OldNode.delete(node.id);
}
```
#### Deploy 3 — Remove from schema
Remove the node type from `defineGraph()` and force migrate. Deploy 2 is what
makes this step legal: `migrateSchema()` refuses to drop a kind that still
holds rows, so if any remain you will get a `MigrationError` with
`details.reason === "kind-removal"` naming the kind and its row count rather
than silent data loss.
### Changing a Property Type
Change `age` from `z.string()` to `z.number()`:
#### Deploy 1 — Add the new property
```typescript
const Person = defineNode("Person", {
schema: z.object({
age: z.string(),
ageNumeric: z.number().optional(),
}),
});
```
#### Deploy 2 — Backfill and switch
```typescript
const people = await store.query(Person).execute();
for (const person of people) {
if (person.properties.ageNumeric === undefined) {
await store.nodes.Person.update(person.id, {
ageNumeric: parseInt(person.properties.age, 10),
});
}
}
```
#### Deploy 3 — Contract
Remove `age`, rename `ageNumeric` to `age` with the new type, and force migrate.
### Changing an Embedding Dimension
Switching embedding models usually changes the vector dimension (e.g.
`embedding(1536)` → `embedding(3072)`). The stored vectors are invalid under the
new dimension — they must be recomputed, not converted — so this is handled
out-of-band from the schema diff. Update the field's `embedding(N)` in the
schema, then call `store.reembedVectorField()`. It drops and recreates the
field's per-`(graphId, kind, field)` `tg_vec_*` storage at the new dimension and,
when you pass an `embed` callback, pages the kind's nodes and re-embeds them:
```typescript
const result = await store.reembedVectorField("Document", "embedding", {
embed: async (nodes) => {
const texts = nodes.map((node) => node.content); // schema fields are top-level
const vectors = await batchEmbed(texts); // your new model
return new Map(nodes.map((node, index) => [node.id, vectors[index]]));
},
});
// result.recreated === true, result.reembedded ===
```
Without an `embed` callback, the storage is recreated empty and you re-embed via
normal `update()` writes. Until a field is re-embedded at the new dimension, a
stray write at the **old** dimension throws `EmbeddingDimensionChangedError`.
## Pre-Deploy Schema Checks
Use `getSchemaChanges()` in CI to catch breaking changes before they reach
production.
### CI/CD Script
```typescript
import { getSchemaChanges } from "@nicia-ai/typegraph/schema";
async function checkSchema(backend: GraphBackend, graph: GraphDef) {
const diff = await getSchemaChanges(backend, graph);
if (!diff) {
console.log("No existing schema — first deploy");
return;
}
if (!diff.hasChanges) {
console.log("Schema unchanged");
return;
}
console.log("Schema changes detected:");
console.log(diff.summary);
for (const change of [...diff.nodes, ...diff.edges]) {
const icon =
change.severity === "safe"
? "[safe]"
: change.severity === "warning"
? "[warn]"
: "[BREAKING]";
console.log(` ${icon} ${change.details}`);
}
if (diff.hasBreakingChanges) {
console.error("Breaking changes require migration before deploy.");
process.exit(1);
}
}
```
### Staging Validation
Before deploying to production, run against a staging database that mirrors
production schema state:
```typescript
const [store, result] = await createStoreWithSchema(graph, stagingBackend);
switch (result.status) {
case "initialized":
console.log("Staging DB was empty — initialized");
break;
case "migrated":
console.log(
`Auto-migrated v${result.fromVersion} → v${result.toVersion}`,
);
console.log("Changes:", result.diff.summary);
break;
case "breaking":
console.error("Would break in production. Fix before deploying.");
process.exit(1);
break;
}
```
## Testing Schema Changes
### Unit Testing Migrations
Test that your migration code handles existing data correctly:
```typescript
import { createStoreWithSchema, defineGraph, defineNode } from "@nicia-ai/typegraph";
import { createTestBackend } from "./test-utils";
it("migrates name to fullName", async () => {
const backend = createTestBackend();
// Set up v1 with data
const graphV1 = defineGraph({
id: "test",
nodes: { Person: { type: PersonV1 } },
edges: {},
});
const [storeV1] = await createStoreWithSchema(graphV1, backend);
await storeV1.nodes.Person.create({ name: "Alice" });
// Migrate to v2 (expand phase)
const graphV2 = defineGraph({
id: "test",
nodes: { Person: { type: PersonV2WithBothFields } },
edges: {},
});
const [storeV2, result] = await createStoreWithSchema(graphV2, backend);
expect(result.status).toBe("migrated");
// Run backfill
const people = await storeV2.query(PersonV2WithBothFields).execute();
for (const person of people) {
await storeV2.nodes.Person.update(person.id, {
fullName: person.properties.name,
});
}
// Verify
const updated = await storeV2.query(PersonV2WithBothFields).execute();
expect(updated[0].properties.fullName).toBe("Alice");
});
```
### Previewing Changes Without Applying
Use `getSchemaChanges()` to see what would change without modifying the database:
```typescript
import { getSchemaChanges } from "@nicia-ai/typegraph/schema";
const diff = await getSchemaChanges(backend, newGraph);
if (diff?.hasChanges) {
console.log("Pending changes:", diff.summary);
console.log("Breaking:", diff.hasBreakingChanges);
for (const change of diff.nodes) {
console.log(` ${change.severity}: ${change.details}`);
}
}
```
## Version History
TypeGraph preserves all schema versions in the `typegraph_schema_versions`
table. Only one version is active at a time.
```text
typegraph_schema_versions
├── version 1 (initial) ← inactive
├── version 2 (added email) ← inactive
├── version 3 (added Company) ← active
```
Access version history through the backend:
```typescript
// Get a specific version
const v1 = await backend.getSchemaVersion("my_app", 1);
console.log("V1 created at:", v1?.created_at);
// Get the active version
const active = await backend.getActiveSchema("my_app");
console.log("Current version:", active?.version);
```
## Summary: Change Classification
| Change | Classification | Auto-Migrated? |
| ------------------------------ | -------------- | -------------- |
| Add node type | Safe | Yes |
| Add edge type | Safe | Yes |
| Add optional property | Safe | Yes |
| Add ontology relation | Safe | Yes |
| Change kind annotations | Safe | Yes |
| Add required property | Breaking | No |
| Remove property | Breaking | No |
| Remove node/edge type | Breaking | No |
| Rename node/edge type | Breaking | No |
| Change property type | Breaking | No |
| Change onDelete behavior | Warning | Yes |
| Change unique constraints | Warning | Yes |
| Change edge cardinality | Warning | Yes |
| Change edge endpoint kinds | Warning | Yes |
| Remove allowed source-dependent endpoint pairs | Breaking | No |
## Rollback
If a deployment goes wrong, you can switch back to a previous schema version.
Version history is always preserved — `rollbackSchema()` simply changes which
version is active.
```typescript
import { rollbackSchema } from "@nicia-ai/typegraph/schema";
// Roll back to version 2
await rollbackSchema(backend, "my_app", 2);
```
This does not delete newer versions. You can migrate forward again later.
## Migration Hooks
Use `onBeforeMigrate` and `onAfterMigrate` for observability — logging,
metrics, and alerts during schema migrations:
```typescript
const [store, result] = await createStoreWithSchema(graph, backend, {
onBeforeMigrate: (context) => {
console.log(`Migrating ${context.graphId} v${context.fromVersion} → v${context.toVersion}`);
console.log("Changes:", context.diff.summary);
},
onAfterMigrate: (context) => {
console.log(`Migration complete: v${context.toVersion}`);
metrics.increment("schema_migrations_total");
},
});
```
For data transformations (backfill scripts), run them explicitly after store
creation rather than inside hooks. This gives you control over retries and
error handling:
```typescript
const [store, result] = await createStoreWithSchema(graph, backend);
if (result.status === "migrated" && result.toVersion === 3) {
// Backfill fullName from name for the expand phase
const people = await store.query(Person).execute();
for (const person of people) {
if (!person.properties.fullName) {
await store.nodes.Person.update(person.id, {
fullName: person.properties.name,
});
}
}
}
```
## Reclaiming Removed Embedding Storage
Embeddings live in per-`(graphId, kind, field)` tables (`tg_vec_*`), provisioned
by the privileged migrator (`createStoreWithSchema`, or `evolve()` for a
runtime-added field). When you remove an `embedding()` field from a **surviving**
kind, the schema change commits fast but the field's now-orphaned vector table
remains until you reconcile it. `store.materializeRemovals()` drops it — and
clears its durable contribution marker so a later re-add re-provisions cleanly
(this is the same pass that cleans up storage for fully removed kinds):
```typescript
const result = await store.materializeRemovals();
for (const reclaimed of result.reclaimedVectorFields) {
// → { kind: "Document", fieldPath: "embedding", status: "reclaimed" }
console.log(`Dropped vector table for ${reclaimed.kind}.${reclaimed.fieldPath}`);
}
```
The pass is idempotent and derived from immutable schema history, so re-running
it lists the same removed fields and the underlying `DROP ... IF EXISTS` is a
no-op on subsequent calls.
## Current Limitations
- **No automatic data transformation.** TypeGraph tracks schema metadata
changes but does not transform existing rows. Use backfill scripts (or
`onAfterMigrate` hooks) for data migration.
- **No rename detection.** Renaming a property looks like a removal + addition.
Use the expand-contract pattern instead.
- **Schema-level only.** Migrations operate on the graph definition, not on
underlying database tables. TypeGraph's storage tables are
schema-agnostic (nodes and edges are stored as JSON properties), so
"schema migration" means updating the schema document that TypeGraph
tracks, not running `ALTER TABLE`.
# Schema Migrations
> Schema versioning, migration, and lifecycle management
For a practical guide on evolving schemas across deployments, see
[Evolving Schemas in Production](/schema-evolution).
## When Do You Need Schema Management?
As your application evolves, your graph schema changes:
- **Adding features**: New node types, new properties, new relationships
- **Refactoring**: Renaming types, changing property formats
- **Deploying safely**: Ensuring schema changes don't break running applications
Without schema management, you'd face:
- No way to know if the database matches your code
- Silent failures when property names change
- Manual migration scripts for every deployment
TypeGraph's schema management:
1. **Stores the schema in the database** alongside your data
2. **Detects changes** between your code and the stored schema
3. **Auto-migrates safe changes** (adding types, optional properties)
4. **Blocks breaking changes** until you handle them explicitly
## How It Works
TypeGraph stores your graph schema in the database, enabling version tracking,
safe migrations, and runtime introspection.
When you create a store with `createStoreWithSchema()`, TypeGraph:
1. Creates the base tables if the database is fresh (auto-bootstrap)
2. Serializes your graph definition to JSON
3. Compares it with the stored schema (if any)
4. Returns the result so you can act on it
## Schema Lifecycle
When you create a store, TypeGraph can automatically manage schema versions:
```typescript
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const [store, result] = await createStoreWithSchema(graph, backend);
switch (result.status) {
case "initialized":
console.log(`Schema initialized at version ${result.version}`);
break;
case "unchanged":
console.log(`Schema unchanged at version ${result.version}`);
break;
case "migrated":
console.log(`Migrated from v${result.fromVersion} to v${result.toVersion}`);
break;
case "pending":
console.log(`Safe changes pending at version ${result.version}`);
break;
case "breaking":
console.log("Breaking changes detected:", result.actions);
break;
}
```
## Basic vs Managed vs Verified Store
TypeGraph provides three ways to create a store, each suited to a
different deployment role:
### Basic Store (No Schema Management)
Use `createStore()` when you manage schema versions yourself:
```typescript
import { createStore } from "@nicia-ai/typegraph";
const store = createStore(graph, backend);
// No schema versioning or write fence - you handle migrations manually
```
Because a basic Store has no committed schema-version metadata, its writes do
not participate in the schema-version fence. Direct backend writes have the
same raw semantics. Use this mode only when the application accepts
responsibility for quiescing writers around schema changes.
:::caution[Fulltext requires the managed store]
`createStore()` is attach-only. If the graph has `searchable()` fields,
use `createStoreWithSchema()` (below) at boot — it durably materializes
the fulltext storage. Bare `createStore()` throws
`StoreNotInitializedError` on the first fulltext operation.
:::
### Managed Store (Automatic Schema Management)
Use `createStoreWithSchema()` for automatic version tracking:
```typescript
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const [store, result] = await createStoreWithSchema(graph, backend, {
autoMigrate: true, // Auto-apply safe changes (default: true)
throwOnBreaking: true, // Throw on breaking changes (default: true)
onBeforeMigrate: (context) => {
console.log(`Migrating ${context.graphId} from v${context.fromVersion} to v${context.toVersion}`);
},
onAfterMigrate: (context) => {
console.log(`Migration complete: v${context.toVersion}`);
},
});
```
### Verified Store (Zero-DDL Attach With Verification Gate)
Use `createVerifiedStore()` at runtime when the application runs under a
least-privilege, DML-only database role and a separate privileged step
has already advanced the schema. It is the runtime counterpart of
`createStoreWithSchema()`: a synchronous-semantics attach that **issues
no DDL** and fails fast if the database is not at the same schema
version as the code graph.
```typescript
import { createVerifiedStore } from "@nicia-ai/typegraph";
// Runtime — least-privilege, DML-only role. Zero DDL.
const [store, result] = await createVerifiedStore(graph, backend);
// result.status === "unchanged" on success.
```
It throws:
- `BaseSchemaMigrationError` if deployment-wide base storage is missing,
stale, or newer than the running library. Its details report
`installedVersion`, `requiredVersion`, and `reason`.
- `ConfigurationError` if no schema has been initialized (run the
privileged migration step first).
- `MigrationError` if the persisted schema is behind the code graph by
**any** pending change (safe or breaking) — the least-privilege
runtime cannot migrate.
- `StoreNotInitializedError` if the schema is current but the
runtime-contribution markers (e.g. fulltext) are missing/stale.
The attach itself can succeed on a non-transactional or custom backend. On a
backend whose `capabilities.execution.unitOfWork` is `"batch"` (Cloudflare
D1, Neon HTTP), a fused write commonly succeeds — see
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares)
for which writes fuse — and a write that cannot fuse throws
`ConfigurationError` with `details.code === "SCHEMA_WRITE_FENCE_UNSUPPORTED"`,
or, for a proven need such as an interactive callback or a schema commit, a
typed error naming `BATCH_WRITE_UNSUPPORTED` under `details.batchRefusal`.
On any other backend that provides neither an interactive transaction nor
the schema-write fence, every managed write throws `ConfigurationError` with
`details.code === "SCHEMA_WRITE_FENCE_UNSUPPORTED"`. Reads remain available.
If you only need the check without building a Store (e.g. a readiness
probe), call `assertSchemaCurrent(backend, graph)` directly — it returns
the same `SchemaValidationResult` or throws the same errors.
:::note[Database privileges]
Only `createStoreWithSchema()` runs DDL. `createStore()` is a
synchronous zero-I/O attach; `createVerifiedStore()` is a SELECT-only
attach (zero DDL — reads the base-schema marker, active graph schema, and
contribution markers, nothing else). Graph-template registration and
instantiation are also DML-only; instantiation copies the source graph's
graph-local activation markers while deployment-scoped physical attestations
remain shared by the database. A target can therefore be reopened by
`createVerifiedStore()` from a later serverless isolate. To run the application
under a least-privilege, DML-only role, do the
privileged migration step once with `createStoreWithSchema(graph, adminBackend)`
to adopt and stamp the current base schema before using the template APIs at
runtime. See
[Database roles & least privilege](/backend-setup#database-roles--least-privilege)
for the canonical breakdown.
:::
### Which Stores are schema-managed?
A Store is schema-managed when it carries committed schema metadata:
`store.introspect().schemaVersion !== undefined`. The following paths create or
preserve that state:
- `createStoreWithSchema()` and `createAdapterStoreWithSchema()`
- `createVerifiedStore()` and `createVerifiedAdapterStore()`
- `createAdapterStore(..., { reconciled })` with a cached reconciled snapshot
- Stores returned by `evolve()` and Stores rebound from an already-managed Store
Managed writes acquire a transaction-scoped fence and revalidate that version
before changing graph data. On the official SQLite and PostgreSQL backends this
prevents a stale Store write from landing across a schema commit. A custom or
non-transactional backend fails closed on the first managed write that cannot
fuse the fence into its own statement — see
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares)
for which writes fuse and which refuse.
`createStore()` and `createAdapterStore()` without `{ reconciled }` are raw,
unversioned attaches. Their writes—and calls made directly through a backend—do
not participate in the fence. `store.clear()` deletes the graph's schema rows
and resets that Store to the same raw state; reopen it through a managed factory
before resuming writes when the versioned guarantee is required.
### Store lifetime after a schema commit
Managed Stores are immutable schema snapshots. A schema-changing operation such
as `evolve()` returns the Store for the resulting schema; it does not update the
instance on which it was called. Switch immediately to the returned Store for
all subsequent work in the same request:
```typescript
const evolved = await store.evolve(extension);
await evolved.getNodeCollectionOrThrow("Paper").create({ title: "..." });
```
For a long-lived local handle, pass a `StoreRef` and use either the return value
or the updated `ref.current` after the call:
```typescript
import type { StoreRef } from "@nicia-ai/typegraph";
const ref: StoreRef = { current: store };
const evolved = await ref.current.evolve(extension, { ref });
// `ref.current === evolved`; do not resume through the pre-evolve Store.
await ref.current.getNodeCollectionOrThrow("Paper").create({ title: "..." });
```
Capturing `ref.current` once at request entry is safe only for requests that do
not change the schema. The ref also cannot observe commits made by another
process or isolate. Before reusing a cross-request cache, compare the cached
Store's `introspect().schemaVersion` (or its reconciled snapshot version) with
`getCommittedSchemaVersion()`, then run `createVerifiedStore()` or
`createVerifiedAdapterStore()` when the version changes. The
[per-request connection recipe](/integration#per-request-connections-cache-the-verified-store)
shows the complete single-flight cache pattern.
## Schema Validation Results
The validation result indicates what happened during store initialization:
| Status | Meaning |
| ------------- | -------------------------------------------------- |
| `initialized` | First run - schema version 1 was created |
| `unchanged` | Schema matches stored version - no changes |
| `migrated` | Safe changes auto-applied, new version created |
| `pending` | Safe changes detected but `autoMigrate` is `false` |
| `breaking` | Breaking changes detected, action required |
The `initialized` and `migrated` results also include
`committedRow: SchemaVersionRow`, the schema row that was just written. Most
applications only need the version fields shown above, but integrations that
build schema metadata can use `committedRow` without issuing another
`getActiveSchema` read.
## Safe vs Breaking Changes
### Safe Changes (Auto-Migrated)
These changes are backwards compatible and can be auto-migrated:
- Adding new node types
- Adding new edge types
- Adding optional properties with defaults
- Adding new ontology relations
### Breaking Changes (Require Manual Action)
These changes require manual migration:
- Removing node or edge types
- Renaming node or edge types
- Changing property types
- Removing properties
- Changing cardinality constraints to be more restrictive
- Removing allowed endpoint pairs from a source-dependent edge
### Endpoint Pair Changes
[Source-dependent targets](/core-concepts#source-dependent-targets) are part of
the serialized schema. The `targetKindsBySource` field preserves the allowed
pairs alongside the source and target kind lists, so export/import and schema
round trips retain the restriction. For compile-time declarations, reordering
map entries or target arrays does not change the schema hash. Persisted runtime
extension documents also contribute to the hash and retain their array order.
Narrowing a target map is breaking even when the overall source and target kind
sets remain unchanged. For example, changing an edge from allowing every
`Employee`/`Student` to `Department`/`Course` combination to allowing only
`Employee → Department` and `Student → Course` removes two pairs. Existing rows
using those pairs need migration before adopting the narrower schema.
Adding allowed pairs, or changing the representation without removing any pairs,
is nonbreaking. Runtime extension changes have an additional empty-kind check
when tightening endpoints; see [extension edges](/graph-extensions#edges).
## Handling Breaking Changes
When breaking changes are detected:
```typescript
const [store, result] = await createStoreWithSchema(graph, backend, {
throwOnBreaking: false, // Don't throw, inspect instead
});
if (result.status === "breaking") {
console.log("Breaking changes detected:");
console.log("Summary:", result.diff.summary);
console.log("Required actions:");
for (const action of result.actions) {
console.log(` - ${action}`);
}
// Option 1: Fix your schema to be backwards compatible
// Option 2: Force migration (data loss possible!)
// import { migrateSchema } from "@nicia-ai/typegraph/schema";
// await migrateSchema(backend, graph, currentVersion);
}
```
### Pre-flighting before you commit
Both checks below are **SELECT-only** — no DDL, no writes — so a least-privilege
runtime can decide what to do *before* it hits the privileged migration wall:
```typescript
import { classifySchemaChanges } from "@nicia-ai/typegraph/schema";
// Cheapest: does this need the privileged path at all?
// (true when the schema is behind, and when nothing is committed yet)
if (await store.requiresMigration()) {
// Route to the privileged bootstrap instead of failing mid-request.
}
// Or get the three-way decision:
const diff = await store.schemaChanges();
const classification =
diff === undefined ? "uninitialized" : classifySchemaChanges(diff);
// "identical" | "additive" | "incompatible"
```
### Classifying a failure
If a commit does fail, branch on the structured outcome rather than the message
text, which is free to be reworded in any release:
```typescript
import { MigrationError } from "@nicia-ai/typegraph";
try {
await commitSomething();
} catch (error) {
if (error instanceof MigrationError) {
switch (error.details.reason) {
case "schema-behind": {
// The runtime can't migrate. `diff` says whether it's safe to proceed.
const additive = error.details.diff?.hasBreakingChanges === false;
break;
}
case "breaking-change": {
break;
}
case "kind-removal": {
// The commit would drop a kind that still holds rows. Narrowing on
// `reason` makes `droppedKinds` non-optional — the details type is a
// discriminated union, so each reason carries exactly its own payload.
const { nodes, edges } = error.details.droppedKinds;
console.error("still populated:", [...nodes, ...edges]);
break;
}
// "no-active-version" | "version-not-found"
}
}
}
```
`details.reason` is a stable discriminant (the `MIGRATION_FAILURE_REASONS`
union), and `details.diff` carries the same structured diff — with per-change
`severity` — that `getSchemaChanges` returns, so you never need a second query
to decide.
## Schema Introspection
### What Does This Database Already Have?
`getActiveSchema` returns the committed schema document — the same JSON stored
in `typegraph_schema_versions.schema_doc`, parsed into a `SerializedSchema`.
Read it instead of querying that table by hand:
```typescript
import { getActiveSchema, isSchemaInitialized, type SerializedSchema } from "@nicia-ai/typegraph";
// Check whether this graph has been committed at all
const initialized = await isSchemaInitialized(backend, "my_graph");
const schema: SerializedSchema | undefined = await getActiveSchema(backend, "my_graph");
if (schema) {
console.log("Version:", schema.version);
console.log("Nodes:", Object.keys(schema.nodes)); // ["Person", "Company"]
console.log("Edges:", Object.keys(schema.edges)); // ["worksAt"]
}
```
These are exported from both the package root and the
`@nicia-ai/typegraph/schema` subpath. Reach for `getCommittedSchemaVersion`
instead when you only need the version number — for example, to invalidate a
cached schema across isolates.
### Previewing Pending Changes
```typescript
import { getSchemaChanges } from "@nicia-ai/typegraph/schema";
const diff = await getSchemaChanges(backend, graph);
if (diff?.hasChanges) {
console.log("Pending changes:", diff.summary);
console.log("Is backwards compatible:", !diff.hasBreakingChanges);
}
```
## Manual Migration
For full control over migrations:
```typescript
import { initializeSchema, migrateSchema, rollbackSchema, ensureSchema } from "@nicia-ai/typegraph/schema";
// Initialize schema (first run only)
const row = await initializeSchema(backend, graph);
console.log("Created version:", row.version);
// Migrate to new version. Folds the persisted graph extension into `graph`
// first, and refuses (MigrationError, reason "kind-removal") if the commit
// would drop a kind that still holds rows.
const newVersion = await migrateSchema(backend, graph, currentVersion);
console.log("Migrated to version:", newVersion);
// Rollback to a previous version
await rollbackSchema(backend, "my_graph", 1);
console.log("Rolled back to version 1");
// Or use ensureSchema for automatic handling
const result = await ensureSchema(backend, graph, {
autoMigrate: true,
throwOnBreaking: true,
});
```
## Migrating Legacy Embedding Storage
Embeddings now live in per-`(graphId, kind, field)` typed tables
(`tg_vec___`), provisioned by `createStoreWithSchema` (the
privileged migrator) at boot. This replaces the single shared
`typegraph_node_embeddings` table. New deployments need no action — the per-field
tables are materialized by `createStoreWithSchema`, which the legacy migration
below also relies on having run.
Deployments that already hold rows in the legacy table run a one-time, idempotent
cutover with `migrateLegacyEmbeddings()`, exported from the package root:
```typescript
import { migrateLegacyEmbeddings } from "@nicia-ai/typegraph";
// `backend` is the post-cutover backend, wired with its VectorStrategy.
const result = await migrateLegacyEmbeddings({ backend });
console.log("Rows migrated:", result.migrated);
console.log("Per field:", result.perField);
console.log("Skipped (dimension mismatch):", result.skippedDimensionMismatch);
console.log("Legacy table existed:", result.legacyTablePresent);
```
The run re-inserts every legacy embedding into per-field storage and is a clean
no-op on a fresh install or a re-run (`legacyTablePresent: false`). A non-empty
`skippedDimensionMismatch` flags `(kind, field)` slots that held mixed dimensions
and need a deliberate re-embed at a single dimension — see
[`reembedVectorField`](/schema-evolution#changing-an-embedding-dimension).
The vector and hybrid query API (`.similarTo()`, `store.search.vector`,
`store.search.hybrid`) is storage-transparent and unchanged by this cutover.
## Migrating Preview Recorded Time
The initial recorded-time preview stored timestamps directly in
`recorded_from`, `recorded_to`, and the graph clock. Versioned anchors now keep
the durable string API while recorded relations compare numeric revisions.
**Stop writers and run the one-time migration before enabling `history: true`
with the new library version.** `createStoreWithSchema` and
`createVerifiedStore` validate the recorded table shapes during an async open
and reject an unmigrated preview schema before returning a store:
```typescript
import {
deleteLegacyRecordedAnchorMap,
migrateLegacyRecordedTime,
migrateRecordedAnchor,
} from "@nicia-ai/typegraph";
const result = await migrateLegacyRecordedTime({ backend });
console.log(result.graphs, result.anchors);
// Translate anchors stored in an application-owned checkpoint table.
const upgraded = await migrateRecordedAnchor({
backend,
graphId: "event-materializer",
anchor: oldTimestampOnlyAnchor,
});
await checkpoints.replaceAnchor(oldTimestampOnlyAnchor, upgraded);
// Do this only after every external checkpoint for the graph is upgraded.
await deleteLegacyRecordedAnchorMap({
backend,
graphId: "event-materializer",
dropWhenEmpty: true,
});
```
The bundled SQLite and PostgreSQL backends provide the recorded-relation DDL needed by this
rewrite. A custom backend that created the preview schema must implement
`backend.recordedTableDdl(tableNames)` before running `migrateLegacyRecordedTime`; otherwise the
migration throws `UnsupportedBackendCapabilityError` with
`details.capability: "recordedTableDdl"`. The callback is invoked for the temporary and final name
sets so the backend, rather than TypeGraph's portable entrypoint, remains the owner of
dialect-specific table and index DDL.
When the engine names primary-key constraints, each callback result must name the constraint for
both name sets or for neither. A one-sided declaration throws `ConfigurationError` with
`details.code: "RECORDED_DDL_CONSTRAINT_NAME_MISMATCH"` before the replacement tables are
published. See
[`recordedTableDdl` in the backend contract](/backend-setup#recorded-table-migration-ddl-recordedtableddl)
when adapting this migration to a custom backend.
The migration dense-ranks distinct legacy commit timestamps independently per
graph, preserving their exact total order. It rewrites the recorded relations
and clock atomically and retains a durable old-anchor mapping so downstream
stores can migrate separately. Re-running it after the cutover is a no-op.
`migrateRecordedAnchor` also accepts an already-versioned `r1` anchor, making a
mixed old/new checkpoint pass idempotent.
The synchronous `createStore` factory is an attach-only, zero-I/O path, so it
cannot inspect table shapes during construction. If used with `history: true`,
an unmigrated schema still fails loudly on the first recorded operation. Prefer
one of the async factories above at application startup when early schema
verification matters.
The old allocator may have pushed a hot graph's physical timestamp ahead of
real wall time. Migration preserves that value because lowering it would put
the clock behind recorded relation boundaries. New commits advance the logical
revision normally, while the physical component remains pinned until wall time
catches up. During that window, diagonal reads use the inherited future valid
time; recorded-only ordering and replay remain exact.
The mapping is graph-scoped: the same timestamp can correspond to different
revisions in different graphs. Keep writers stopped for the schema rewrite, and
delete mapping rows only after every external checkpoint for that graph has
been translated. `dropWhenEmpty: true` atomically drops the mapping table when
the deleted graph was the final one. Without that option, the empty table is
retained intentionally and can be dropped by your normal migration tooling.
## Repairing Inverted Validity Windows
Older library versions could store a row whose validity window runs backwards
(`valid_from > valid_to`). Such a row is readable at **no** coordinate at all:
`asOf(t)` needs `valid_from <= t < valid_to`, and backwards bounds admit no `t`.
The write paths no longer produce one — a write that stamps a lower bound the
caller did not state now stores no bound rather than an inverting one, see
[Open-left rows](/queries/temporal#open-left-rows-validfrom-is-undefined) — but
**upgrading rewrites nothing**. Rows already stored that way keep their window
and stay invisible until an operator repairs them, which is deliberate: an
upgrade that silently made previously-invisible rows appear in historical
queries would be the worse surprise.
`repairInvertedValidityWindows` is that explicit action. It has two modes:
`report` counts and writes nothing, `apply` normalizes the rows it counted to
`valid_from = NULL` ("ended at T, start unknown").
```typescript
import { repairInvertedValidityWindows } from "@nicia-ai/typegraph";
// Diagnose. `report` reads through `execute`, a required backend member, so it
// runs against ANY backend — including a history-capturing one and one with no
// statement-execution support.
const report = await repairInvertedValidityWindows({
backend: anyBackend,
relations: "live-and-recorded",
mode: "report",
});
// report.counts.recordedNodes === undefined means NOT SCANNED, never "clean".
// report.atomic === false means the counts came from per-relation snapshots.
// Repair, with writers stopped. On a history-enabled store pass the RAW backend
// you constructed it from: the repair mints no revision by design, and the
// capture wrapper refuses raw statements.
await repairInvertedValidityWindows({
backend: rawBackend,
relations: "live-and-recorded",
mode: "apply",
});
```
If `tableNames` is supplied, it patches `backend.tableNames`; unstated relation
names keep the backend's configured values. A partial override never sends the
other relations back to TypeGraph's built-in defaults.
`relations` is **required**, and `"live-and-recorded"` is the recommended scope.
Repairing only the live axis leaves the recorded twin carrying the inverted
window, which re-materializes the invisible row at any `asOfRecorded`
coordinate — the same defect one axis over. `"live"` is right in exactly two
cases: the store captures no history and the `recorded_*` tables do not exist
(scanning them is then an error, not a no-op), or you are deliberately keeping
the recorded axis as an audit record of the pre-repair state and accept that
historical `asOfRecorded` reads keep returning the invisible shape.
What an operator must know before running it:
1. **Run `apply` with writers stopped**, the same guidance
`migrateLegacyRecordedTime()` carries. A concurrent window-bearing update
may fence its write on the validity lower bound it read, so a repair landing
in between can make the peer's first `UPDATE` match no row. Store node and
edge updates re-read and re-judge against the repaired bound; interchange
records a per-row target-changed error instead of claiming the row was
written. `report` needs no quiescing: it scans in a read-only transaction
(`BEGIN` rather than SQLite's writer-reserving `BEGIN IMMEDIATE`, and
`BEGIN … READ ONLY` on PostgreSQL), so it cannot write itself.
2. **Repaired rows become visible** at `asOf` coordinates before their end. That
is the point, and it is a read-visibility change to historical queries.
3. **Outstanding `base@V` merge tokens are invalidated** for repaired rows —
`valid_from` is part of the base content fingerprint, so a merge whose base
token predates the repair fails its precondition afterwards. Quiesce merges,
repair, then re-baseline branches.
4. **The repair mints no revision and bumps no `version`**, and does not move
`updated_at`. It normalizes a storage convention for rows that were never
observable at any coordinate; it is not a logical write. That is why `apply`
is run against the raw backend, and why bypassing recorded-time capture here
is intended rather than a workaround.
5. **`apply` refuses when a scanned relation stores non-canonical bounds**
(SQLite only — PostgreSQL stores `timestamptz`, so a scanned relation always
reports `nonCanonical: 0`). SQLite compares the bounds as text, so a
non-canonical value cannot be classified without a timestamp semantics this
repair does not own. The refusal is total: the whole call is rejected before
any row is updated, so `apply` never repairs the rows it understood and skips
the rest. `report` still counts them, in `nonCanonical` — normalize those
bounds, or narrow the call with `graphId`, and re-run.
6. **On a backend without transactions the call still runs**, per relation, and
says so with `report.atomic === false`: the counts may span snapshots, and a
crash mid-`apply` can leave the live axis repaired and the recorded axis not.
Re-run — each statement is idempotent and convergent, and a later `report`
proves it converged.
7. **Repair before exporting a legacy graph.** An exported inverted row is
refused per row on re-import, so an unrepaired graph does not round-trip.
The statement touches only rows the library mis-stored, so it is empty on a
healthy graph and needs no batching. If a report returns a count large enough to
worry about, narrow the call with `graphId` and run it per graph.
## Schema Serialization
Schemas are stored as JSON documents with computed hashes for fast comparison:
```typescript
import { serializeSchema, computeSchemaHash } from "@nicia-ai/typegraph/schema";
// Serialize a graph definition
const serialized = serializeSchema(graph, 1);
// Compute hash for comparison
const hash = computeSchemaHash(serialized);
```
The serialized schema includes:
- Graph ID and version
- All node types with their Zod schemas (as JSON Schema)
- All edge types with endpoints and constraints
- Complete ontology relations
- Uniqueness constraints and delete behaviors
## Version History
TypeGraph maintains a history of all schema versions:
```text
typegraph_schema_versions
├── version 1 (initial)
├── version 2 (added User node)
├── version 3 (added email property) ← active
└── ...
```
Only one version is marked as "active" at a time. Previous versions are
preserved for auditing and potential rollback.
## Best Practices
### 1. Use Managed Stores in Production
```typescript
// Production: Use schema management
const [store, result] = await createStoreWithSchema(graph, backend);
// Development: Basic store is fine for rapid iteration
const store = createStore(graph, backend);
```
### 2. Check Migration Status on Startup
```typescript
async function initializeApp() {
const [store, result] = await createStoreWithSchema(graph, backend);
if (result.status === "breaking") {
console.error("Database schema incompatible with application!");
console.error("Run migrations before deploying this version.");
process.exit(1);
}
if (result.status === "migrated") {
console.log(`Schema auto-migrated to v${result.toVersion}`);
}
return store;
}
```
### 3. Preview Changes Before Deployment
```typescript
import { getSchemaChanges } from "@nicia-ai/typegraph/schema";
// In your CI/CD pipeline or migration script
const diff = await getSchemaChanges(backend, graph);
if (diff?.hasChanges) {
console.log("Schema changes detected:");
console.log(diff.summary);
if (!diff.isBackwardsCompatible) {
console.error("Breaking changes require manual migration!");
process.exit(1);
}
}
```
### 4. Add Properties with Defaults
When adding new properties, always provide defaults to ensure backwards
compatibility:
```typescript
// Good: Optional with default
const User = defineNode("User", {
schema: z.object({
name: z.string(),
// New property with default - safe migration
status: z.enum(["active", "inactive"]).default("active"),
}),
});
// Bad: Required without default - breaking change
const User = defineNode("User", {
schema: z.object({
name: z.string(),
status: z.enum(["active", "inactive"]), // No default!
}),
});
```
# Semantic Search
> Vector embeddings and similarity search for AI-powered retrieval
TypeGraph supports semantic search using vector embeddings, enabling you to find
semantically similar content using embedding models like OpenAI, Sentence Transformers,
CLIP, or any model that produces fixed-dimension vectors.
## Overview
Traditional search relies on exact keyword matching. Semantic search understands
meaning—"machine learning" matches documents about "neural networks" and "AI algorithms"
even without those exact words.
**Key capabilities:**
- Store embeddings as node properties alongside your graph data
- Find the k most similar nodes using cosine, L2, or inner product distance
- Combine semantic similarity with graph traversals and standard predicates
- Automatic vector indexing for fast approximate nearest neighbor search
## Use Cases
### Retrieval-Augmented Generation (RAG)
Build context-aware AI applications by retrieving relevant documents before
generating responses:
```typescript
async function ragQuery(question: string): Promise {
const questionEmbedding = await embed(question);
const context = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.embedding.similarTo(questionEmbedding, 5, {
metric: "cosine",
minScore: 0.7,
})
)
.select((ctx) => ({
title: ctx.d.title,
content: ctx.d.content,
}))
.execute();
return await llm.chat({
messages: [
{
role: "system",
content: `Answer based on this context:\n${context.map((d) => d.content).join("\n\n")}`,
},
{ role: "user", content: question },
],
});
}
```
### Semantic Document Search
Find documents by meaning rather than keywords:
```typescript
const results = await store
.query()
.from("Article", "a")
.whereNode("a", (a) =>
a.embedding
.similarTo(queryEmbedding, 20)
.and(a.category.eq("technology"))
)
.select((ctx) => ctx.a)
.execute();
```
### Image Similarity
Use CLIP or similar vision models for image search:
```typescript
const similarImages = await store
.query()
.from("Image", "i")
.whereNode("i", (i) => i.clipEmbedding.similarTo(queryImageEmbedding, 10))
.select((ctx) => ({
url: ctx.i.url,
caption: ctx.i.caption,
}))
.execute();
```
### Product Recommendations
Recommend products based on embedding similarity:
```typescript
const recommendations = await store
.query()
.from("Product", "p")
.whereNode("p", (p) =>
p.embedding
.similarTo(referenceProductEmbedding, 10)
.and(p.inStock.eq(true))
)
.select((ctx) => ctx.p)
.execute();
```
## Database Setup
Vector search requires database-specific extensions for storing and querying
high-dimensional vectors efficiently.
### PostgreSQL with pgvector
[pgvector](https://github.com/pgvector/pgvector) is the recommended extension
for PostgreSQL. It provides:
- Native `vector` column type
- HNSW and IVFFlat indexes for fast approximate nearest neighbor search
- Support for cosine, L2, and inner product distance
**Installation:**
```sql
-- Install the extension (requires superuser or database owner)
CREATE EXTENSION vector;
```
**Docker setup:**
```yaml
services:
postgres:
image: pgvector/pgvector:pg16
environment:
POSTGRES_PASSWORD: password
POSTGRES_DB: myapp
ports:
- "5432:5432"
```
**TypeGraph migration enables vector support:**
```typescript
import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Generates DDL including `CREATE EXTENSION IF NOT EXISTS vector;`.
// It does NOT create a single embeddings table — each embedding field gets
// its own typed `vector(N)` table, provisioned by `createStoreWithSchema`
// (the privileged migrator) at boot (see Storage Layout below).
const migrationSQL = generatePostgresMigrationSQL();
```
### SQLite with sqlite-vec
[sqlite-vec](https://github.com/asg017/sqlite-vec) provides vector search
for SQLite. It offers:
- `vec_f32` type for 32-bit float vectors
- Cosine and L2 distance functions
:::caution[sqlite-vec requires a native (better-sqlite3) connection]
sqlite-vec is a loadable C extension. TypeGraph loads it through
better-sqlite3's `loadExtension` in `createLocalSqliteBackend`, so it only
applies to the **local, native** SQLite backend. It does **not** apply to the
**libSQL / Turso** backend (`createLibsqlBackend`): `@libsql/client` does not
expose `loadExtension`, and libSQL ships its **own** native vector engine
(`F32_BLOB`, `vector_distance_cos`, `vector_top_k`) which is a different API
than sqlite-vec. See [libSQL / Turso](#libsql--turso-native-vectors) below.
:::
**Installation:**
```bash
npm install sqlite-vec
```
**Loading the extension:**
```typescript
import Database from "better-sqlite3";
import * as sqliteVec from "sqlite-vec";
const sqlite = new Database("myapp.db");
sqliteVec.load(sqlite);
```
**Limitations:**
- sqlite-vec does not support inner product distance
- Use `cosine` or `l2` metrics only
### libSQL / Turso (native vectors)
The **libSQL / Turso** backend (`createLibsqlBackend`) does **not** use
sqlite-vec. libSQL has a built-in vector engine — no extension to load — so
vector and hybrid search work out of the box on local files, embedded
replicas, and remote Turso databases:
```typescript
import { createClient } from "@libsql/client";
import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql";
const client = createClient({ url: "libsql://my-db.turso.io", authToken: "..." });
const { backend } = await createLibsqlBackend(client);
// backend.capabilities.vector?.supported === true
```
Under the hood it stores embeddings as `F32_BLOB` and searches with
`vector_distance_cos` / `vector_distance_l2`, with optional approximate
nearest-neighbor (DiskANN) indexes via `libsql_vector_idx` + `vector_top_k`.
Supported metrics are `cosine` and `l2` (no `inner_product`), matching the
sqlite-vec feature set.
One caveat specific to DiskANN: `vector_top_k` is a table function with no
filter pushdown, so the liveness filter every search applies (only
non-deleted nodes may rank — see below) runs *after* ANN retrieval.
TypeGraph over-fetches 4× `limit` neighbors to leave headroom; if more than
3×`limit` of those neighbors are filtered out, fewer than `limit` results
return. pgvector and sqlite-vec apply the filter inside the index scan and
do not share this bound.
### Supported Distance Metrics
| Metric | PostgreSQL | SQLite (sqlite-vec) | libSQL / Turso | Description |
|--------|------------|---------------------|----------------|-------------|
| `cosine` | `<=>` | `vec_distance_cosine` | `vector_distance_cos` | Cosine distance (1 - similarity). Best for normalized embeddings. |
| `l2` | `<->` | `vec_distance_l2` | `vector_distance_l2` | Euclidean distance. Good for unnormalized vectors. |
| `inner_product` | `<#>` | Not supported | Not supported | Negative inner product. For maximum inner product search (MIPS). |
## Storage Layout & Maintenance
Each embedding field is stored in its own typed, graph-scoped table named
`tg_vec___`, carrying that field's fixed dimension
(pgvector `vector(N)`, libSQL `F32_BLOB(N)`, sqlite-vec `vec0`). The privileged
migrator (`createStoreWithSchema`, and `evolve()` for runtime-added fields)
provisions each table plus a durable contribution marker at boot; the runtime
hot path then asserts the marker (a cached SELECT) and never issues DDL, so a
least-privilege, DML-only role can read and write embeddings. An embedding
write against an un-provisioned slot throws `StoreNotInitializedError` rather
than lazily creating the table; vector reads (`store.search.vector`,
`store.search.hybrid`, and query-builder `.similarTo()` predicates) compile
straight to SQL, so they surface the engine's missing-relation error instead —
use `createVerifiedStore` to catch both at attach. See [Database roles & least
privilege](/backend-setup#database-roles--least-privilege). Graph-scoping means
several graphs in one database can declare the same `kind`+`field` at different
dimensions without collision. This is transparent to queries — `.similarTo()`,
`store.search.vector`, and `store.search.hybrid` read it for you.
### Deleted nodes never rank
Every facade search (`store.search.vector` / `fulltext` / `hybrid`) computes
its top-k over live nodes only: the search SQL constrains candidates to
non-deleted node ids, so a stale embedding or fulltext row — one whose node
was tombstoned by a writer that bypassed the store's cleanup — can neither
surface in results nor crowd live rows out of the top-k. You always get
`limit` results when at least `limit` live matches exist (on libSQL DiskANN,
subject to the over-fetch bound above).
### Changing an embedding dimension
Switching embedding models usually changes the vector dimension. Stored vectors
can't be reinterpreted at a new dimension, so a stray write at the old
dimension throws `EmbeddingDimensionChangedError`. Update the field's
`embedding(N)` declaration, then recompute the stored vectors with
`store.reembedVectorField()`, which recreates the field's storage at the new
dimension:
```typescript
// embedding(1536) → embedding(3072): recreate storage and re-embed in batches.
// `embed` receives a page of nodes and returns a Map from node id to vector.
await store.reembedVectorField("Document", "embedding", {
embed: async (nodes) => {
const vectors = await batchEmbed(nodes.map((node) => node.content));
return new Map(nodes.map((node, index) => [node.id, vectors[index]]));
},
});
// → { recreated: true, reembedded: }
```
Between the declaration change and the `reembedVectorField()` call, the slot
is in a deliberate limbo: boot (`createStoreWithSchema` / `evolve()`) detects
that the provisioned storage no longer matches the declared shape, warns, and
leaves it untouched — it never recreates the table implicitly, because that
would silently drop every stored vector. Embedding writes to the field fail
with a `StoreNotInitializedError` whose reason is `stale` (its message points
here) until `reembedVectorField()` recreates the storage and re-stamps its
durable marker.
Without an `embed` callback the storage is recreated empty and you re-embed via
normal `update()` writes.
### Reclaiming removed embedding fields
Removing an embedding field from a kind that still exists orphans its
`tg_vec_*` table. `store.materializeRemovals()` reclaims it — it drops per-field
tables for embedding fields no longer in the active schema and reports them in
`reclaimedVectorFields`:
```typescript
const { reclaimedVectorFields } = await store.materializeRemovals();
// → [{ kind: "Document", fieldPath: "embedding", status: "reclaimed" }]
```
The active schema is the source of truth, so a removed-then-re-added field is
never dropped. The pass is idempotent.
### Migrating from the legacy shared table
Earlier versions stored every embedding in a single shared
`typegraph_node_embeddings` table. If you have existing data there, run the
one-time, idempotent `migrateLegacyEmbeddings()` utility to copy it into the new
per-field tables (new deployments need no action):
```typescript
import { migrateLegacyEmbeddings } from "@nicia-ai/typegraph";
const result = await migrateLegacyEmbeddings({ backend });
// → { migrated, perField, skippedDimensionMismatch, legacyTablePresent }
```
## Schema Design
### Defining Embedding Properties
Use the `embedding()` function to define vector properties with a specific dimension:
```typescript
import { defineNode, embedding } from "@nicia-ai/typegraph";
import { z } from "zod";
const Document = defineNode("Document", {
schema: z.object({
title: z.string(),
content: z.string(),
embedding: embedding(1536), // OpenAI ada-002 dimension
}),
});
const Image = defineNode("Image", {
schema: z.object({
url: z.string(),
caption: z.string().optional(),
clipEmbedding: embedding(512), // CLIP ViT-B/32 dimension
}),
});
```
### Common Embedding Dimensions
| Model | Dimensions | Use Case |
|-------|------------|----------|
| all-MiniLM-L6-v2 | 384 | Fast, lightweight text embeddings |
| CLIP ViT-B/32 | 512 | Image-text multimodal |
| BERT base | 768 | General text embeddings |
| OpenAI ada-002 | 1536 | High-quality text embeddings |
| OpenAI text-embedding-3-small | 1536 | Efficient, high-quality |
| OpenAI text-embedding-3-large | 3072 | Maximum quality |
| Cohere embed-v3 | 1024 | Multilingual support |
### Optional Embeddings
Embedding properties can be optional for gradual population:
```typescript
const Article = defineNode("Article", {
schema: z.object({
title: z.string(),
content: z.string(),
embedding: embedding(1536).optional(),
}),
});
// Create without embedding
const article = await store.nodes.Article.create({
title: "Draft Article",
content: "...",
});
// Add embedding later via background job
await store.nodes.Article.update(article.id, {
embedding: await generateEmbedding(article.content),
});
```
### Multiple Embeddings per Node
Nodes can have multiple embedding fields for different purposes:
```typescript
const Product = defineNode("Product", {
schema: z.object({
name: z.string(),
description: z.string(),
imageUrl: z.string(),
// Text embedding for description search
textEmbedding: embedding(1536).optional(),
// Image embedding for visual similarity
imageEmbedding: embedding(512).optional(),
}),
});
```
## Storing Embeddings
Embeddings are stored when creating or updating nodes:
```typescript
// Using OpenAI
import OpenAI from "openai";
const openai = new OpenAI();
async function generateEmbedding(text: string): Promise {
const response = await openai.embeddings.create({
model: "text-embedding-ada-002",
input: text,
});
return response.data[0].embedding;
}
// Store with embedding
const embedding = await generateEmbedding("Machine learning fundamentals");
await store.nodes.Document.create({
title: "ML Guide",
content: "Machine learning fundamentals...",
embedding: embedding,
});
```
### Batch Embedding
For bulk operations, batch your embedding API calls:
```typescript
async function batchEmbed(texts: string[]): Promise {
const response = await openai.embeddings.create({
model: "text-embedding-ada-002",
input: texts,
});
return response.data.map((d) => d.embedding);
}
// Process in batches
const documents = await fetchDocumentsWithoutEmbeddings();
const batchSize = 100;
for (let i = 0; i < documents.length; i += batchSize) {
const batch = documents.slice(i, i + batchSize);
const embeddings = await batchEmbed(batch.map((d) => d.content));
await store.transaction(async (tx) => {
for (const [index, doc] of batch.entries()) {
await tx.nodes.Document.update(doc.id, {
embedding: embeddings[index],
});
}
});
}
```
## Querying
### Basic Similarity Search
Use `.similarTo()` to find the k most similar nodes:
```typescript
const queryEmbedding = await generateEmbedding("neural networks");
const similar = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.embedding.similarTo(queryEmbedding, 10) // Top 10 most similar
)
.select((ctx) => ({
title: ctx.d.title,
content: ctx.d.content,
}))
.execute();
```
### Approximate retrieval for `.similarTo()` (opt-in)
By default `.similarTo()` ranks with an exact distance scan — correct at any
scale, and index-served by the PostgreSQL planner where the plan shape
allows. When a kind declares an ANN index (`embedding(n)` defaults to
`hnsw`), you can opt the predicate into the engine's native approximate
retrieval:
```typescript
const similar = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.embedding
.similarTo(queryEmbedding, 10, { approximate: true })
.and(d.status.eq("published")),
)
.select((ctx) => ctx.d)
.execute();
```
This is a semantic change, never applied silently: results are subject to
the index's recall. Composed predicates constrain the ANN candidate set —
exactly on pgvector and sqlite-vec, bounded by over-fetch on libSQL
DiskANN. A kind declared with `indexType: "none"` keeps the exact scan even
with the opt-in — a declared degradation, since there is no ANN structure
for `approximate` to opt into and the results are exactly what was asked
for.
**A `metric` override that differs from the field's declared metric is
refused when `approximate: true` is stated.** Every engine materializes
metric-specific ANN structures — `vec0` bakes `distance_metric` into the
virtual table, libSQL's DiskANN index is built with `metric=…`, pgvector's
index carries a per-metric operator class — so an ANN structure only
retrieves under the metric it was built for. Retrieving by the declared
metric and re-scoring under the override returns the declared metric's
neighbors wearing the override's scores: the wrong rows, silently. The two
options state something that cannot both hold, so the call throws a
`ConfigurationError` naming both metrics rather than quietly serving the
exact scan:
```typescript
// Refused: the HNSW index is built for cosine.
d.embedding.similarTo(queryEmbedding, 10, {
approximate: true,
metric: "l2",
});
```
`details` carries `nodeKind`, `fieldPath`, `requestedMetric`,
`declaredMetric`, and `indexType`. Omit `metric` (or pass the declared one)
to keep approximate retrieval, or drop `approximate` to scan exactly under
the overriding metric. A slot declared `indexType: "none"` is not refused:
there is no ANN structure to be bound to a metric.
Note the deliberate asymmetry with the facade. `store.search.vector` and
`store.search.hybrid` refuse **every** metric override that differs from
the declared one, whether or not `approximate` was stated — their rule is
broader because vector storage is built for the declared metric and the
facade is the guided surface. The query builder's *exact* path stays wider
on purpose: an exact scan computes any metric over the stored vectors
correctly, and nothing was stated there that the engine cannot honor. Only
the silent half — the combination that cannot be served — is closed here.
### Scoped facade search: filters, pagination, subclasses
`store.search.vector` (and `fulltext` / `hybrid`) accept a `where`
predicate, an `offset`, and `includeSubClasses` — all compiled into the
search statement itself, so the engine ranks only eligible rows. A filter
never costs you results: you get `limit` hits whenever `limit` matching
nodes exist (on libSQL DiskANN, subject to the over-fetch bound above).
```typescript
// Top 10 most similar *published* documents, second page.
const hits = await store.search.vector("Document", {
fieldPath: "embedding",
queryEmbedding,
limit: 10,
offset: 10,
where: (d) => d.status.eq("published"),
});
// Search a kind and all of its subClassOf descendants; per-kind results
// merge into one globally ordered ranking. Kinds that don't declare the
// embedding field are skipped.
const acrossKinds = await store.search.vector("Content", {
fieldPath: "embedding",
queryEmbedding,
limit: 10,
includeSubClasses: true,
});
```
The `where` predicate is compiled by the same query compiler as
`store.query()` — property predicates behave identically, use the same
declared indexes, and apply the same current-read semantics (tombstoned
nodes and nodes outside their validity window never rank). Kinds expanded
via `includeSubClasses` must share one declared metric: scores from
different metrics cannot merge into one ranking (and a per-call `metric`
cannot bridge the gap — each kind's storage is validated against its
declared metric), so mixed-metric expansions throw; search those kinds
separately.
### Choosing a Distance Metric
```typescript
// Cosine similarity (default) - best for normalized embeddings
d.embedding.similarTo(queryEmbedding, 10, { metric: "cosine" })
// L2 (Euclidean) distance - for unnormalized embeddings
d.embedding.similarTo(queryEmbedding, 10, { metric: "l2" })
// Inner product - for maximum inner product search (PostgreSQL only)
d.embedding.similarTo(queryEmbedding, 10, { metric: "inner_product" })
```
**When to use each:**
- **Cosine**: Most common choice. Works well with normalized embeddings
(OpenAI, Sentence Transformers). Focuses on direction, not magnitude.
- **L2**: Use when vector magnitude matters. Good for detecting exact
duplicates.
- **Inner product**: For MIPS (maximum inner product search). Useful when
embeddings encode both relevance and importance in magnitude.
### Minimum Score Filtering
Filter results below a similarity threshold:
```typescript
const highQualityMatches = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.embedding.similarTo(queryEmbedding, 100, {
metric: "cosine",
minScore: 0.8, // Only results with similarity >= 0.8
})
)
.select((ctx) => ctx.d)
.execute();
```
The `minScore` parameter filters results using **similarity** (not distance):
- **Cosine**: 1.0 = identical, 0.0 = orthogonal. Typical thresholds: 0.7-0.9
- **L2**: Maximum distance to include (lower = more similar)
- **Inner product**: Minimum inner product value
:::note[Similarity vs Distance]
While the underlying database operators use distance (where 0 = identical for cosine),
`minScore` uses similarity semantics for intuitive usage. TypeGraph converts internally:
`distance_threshold = 1 - minScore` for cosine.
:::
### Combining with Predicates
Semantic search integrates with all standard query predicates:
```typescript
const filteredSearch = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.embedding
.similarTo(queryEmbedding, 20)
.and(d.category.eq("technology"))
.and(d.publishedAt.gte("2024-01-01"))
.and(d.status.eq("published"))
)
.select((ctx) => ctx.d)
.execute();
```
### Combining with Graph Traversals
Search within graph relationships:
```typescript
// Find similar documents by authors I follow
const personalizedSearch = await store
.query()
.from("Person", "me")
.whereNode("me", (p) => p.id.eq(currentUserId))
.traverse("follows", "f")
.to("Person", "author")
.traverse("authored", "a", { direction: "in" })
.to("Document", "d")
.whereNode("d", (d) =>
d.embedding.similarTo(queryEmbedding, 10)
)
.select((ctx) => ({
title: ctx.d.title,
author: ctx.author.name,
}))
.execute();
```
## Best Practices
### Normalize Your Embeddings
Most embedding models produce normalized vectors (unit length). If yours doesn't,
normalize before storing:
```typescript
function normalize(vector: number[]): number[] {
const magnitude = Math.sqrt(vector.reduce((sum, v) => sum + v * v, 0));
return vector.map((v) => v / magnitude);
}
await store.nodes.Document.create({
title: "Example",
content: "...",
embedding: normalize(rawEmbedding),
});
```
### Use Consistent Embedding Models
Always use the same model for both storing and querying:
```typescript
// Bad: Mixing models
const docEmbedding = await embed("text-embedding-ada-002", content);
const queryEmbedding = await embed("text-embedding-3-small", query); // Different!
// Good: Same model throughout
const MODEL = "text-embedding-ada-002";
const docEmbedding = await embed(MODEL, content);
const queryEmbedding = await embed(MODEL, query);
```
### Handle Missing Embeddings
Not all nodes may have embeddings. Handle gracefully:
```typescript
// Only search nodes with embeddings
const results = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.embedding
.isNotNull()
.and(d.embedding.similarTo(queryEmbedding, 10))
)
.select((ctx) => ctx.d)
.execute();
```
### Choose Appropriate k Values
The `k` parameter (number of results) affects performance:
```typescript
// For RAG: Small k (3-10) for focused context
d.embedding.similarTo(query, 5)
// For exploration: Larger k with pagination
d.embedding.similarTo(query, 100)
```
### Index Considerations
Vector indexes (HNSW, IVFFlat) trade accuracy for speed:
- **Small datasets (< 10K)**: Exact search is fast enough
- **Medium datasets (10K-1M)**: HNSW provides good recall with fast queries
- **Large datasets (> 1M)**: Consider IVFFlat with appropriate parameters
TypeGraph creates HNSW indexes by default for optimal balance.
### Filtered vector search needs a node index on the filter field
Combining `similarTo` with a property predicate is the shape that
degrades first at scale — and the vector index is not the reason. The
candidates side (`d.category.eq(...)`) is a JSON property predicate
over the nodes table, and rows that carry an embedding field have LARGE
props: on PostgreSQL the predicate scan detoasts every row, so at 50k
documents (384-dim embeddings) the filter alone costs ~375ms regardless
of how the vector side is executed. SQLite pays the same class of cost
parsing large JSON props per row.
Declare a node index on the filter field and materialize it — the
candidates predicate becomes an index lookup:
```typescript
import { defineNodeIndex } from "@nicia-ai/typegraph/indexes";
const categoryIndex = defineNodeIndex(Document, { fields: ["category"] });
const graph = defineGraph({
id: "docs",
nodes: { Document: { type: Document } },
edges: {},
indexes: [categoryIndex],
});
await store.materializeIndexes();
```
Measured at 50k documents on PostgreSQL: the filtered exact search
drops from ~375ms to ~19ms and the filtered approximate search to
~20ms — a ~20× difference from one declared index. The `bench:vector`
lane tracks both forms (`vector:exact-filtered` before the index,
`vector:exact-filtered-postindex` after).
### Approximate search under selective filters
`approximate: true` combined with a highly selective property filter is
the shape where approximate means it. The index scan walks neighbors
best-first and keeps going until enough filtered rows surface
(TypeGraph applies pgvector's `hnsw.iterative_scan = strict_order`
automatically on transaction-capable Postgres drivers with pgvector
≥ 0.8 — the setting is transaction-scoped, so non-transactional
backends such as `neon-http` keep the plain bounded scan), but the
scan is still bounded by
pgvector's `hnsw.max_scan_tuples` (default 20,000). If the nearest rows
matching the filter live far from the query — a filter *correlated*
with embedding geometry, like "category X" when category X's documents
form their own distant cluster — the scan can exhaust its budget and
return plausible-but-distant rows. For filters independent of the
embedding space (the common case), filtered approximate recall stays
near 1.0. When the filter is known to be geometry-correlated and
selective, drop `approximate` (the exact path is index-assisted on the
candidates side by a node index on the filter field) or raise
`hnsw.max_scan_tuples`.
### Tuning recall per query with `efSearch`
pgvector's HNSW index searches a dynamic candidate list whose size is
the `hnsw.ef_search` GUC — **default 40**. That frontier caps how many
neighbors a single scan can surface, so on corpora past a few million
vectors recall@k flattens well below 1.0 at the default. TypeGraph
exposes it as a per-search `efSearch` knob on `store.search.vector` and
the vector half of `store.search.hybrid`:
```typescript
const hits = await store.search.hybrid("Document", {
limit: 20,
vector: {
fieldPath: "embedding",
queryEmbedding,
k: 80, // over-fetch 80 candidates from the vector side
efSearch: 240, // ~3× k — high-recall frontier for this query
},
fulltext: { query: "renewable energy" },
});
```
Sizing guidance:
- **Floor — `efSearch >= k`.** Hybrid over-fetches `k` candidates from
the vector side (default `4 * limit`). If `efSearch` is below `k` the
scan can't fill the candidate set, so the over-fetch silently
under-delivers — RRF papers over this on head queries (the fulltext
half covers the miss) but drops tail queries only the vector side
knows about.
- **Target — ~2–4× `k`.** On million-scale corpora this clears roughly
0.95 recall@10, versus ~0.82–0.85 at the default 40. Verify the curve
against your own corpus rather than hard-coding a multiplier.
- **Ceiling — 1000.** pgvector caps `hnsw.ef_search` at 1000; TypeGraph
rejects a larger `efSearch` with a clear error.
Because it's per-search, one connection pool can serve both a
latency-sensitive interactive path (omit `efSearch`, inherit the session
default) and a recall-sensitive batch/ETL path (raise it) — a session
GUC can't, a per-call override can.
**Mechanics and limits.** The override is applied transaction-locally
(`SET LOCAL hnsw.ef_search`) around the vector `SELECT`, so it never
leaks to the next query on a pooled connection. Omitting it preserves
today's behavior exactly — no transaction is opened. It applies to the
**Postgres HNSW** path only:
- **SQLite backends refuse it.** Neither `sqlite-vec` — whose `vec0` KNN
takes only `k`, the page size — nor `libsql-native`, whose DiskANN
`vector_top_k` fixes `search_l` at index-creation time, has a per-search
frontier to set. Supplying `efSearch` to either, on the vector path or
the hybrid path, is refused with `UnsupportedBackendCapabilityError`
(`details.capability` `vector.searchFrontierTuning`, `details.reason`
naming the engine's limitation) rather than searching as if the option
had not been passed.
- **Transaction-less Postgres drivers** (`drizzle-orm/neon-http`) can't
scope `SET LOCAL`, so a search that supplies `efSearch` is refused with
`UnsupportedBackendCapabilityError`. Use a transactional driver
(`node-postgres` / `neon-serverless` / `postgres-js`) to apply it.
- It tunes HNSW only. Supplying it for IVFFlat is refused with a
`ConfigurationError`; IVFFlat's analogous knob (`ivfflat.probes`) is not yet
exposed.
## Troubleshooting
### "Extension not found" errors
**PostgreSQL:**
```sql
-- Check if pgvector is installed
SELECT * FROM pg_extension WHERE extname = 'vector';
-- Install it
CREATE EXTENSION vector;
```
**SQLite:**
```typescript
// Ensure sqlite-vec is loaded before queries
import * as sqliteVec from "sqlite-vec";
sqliteVec.load(sqlite);
```
### "Inner product not supported" (SQLite)
sqlite-vec only supports `cosine` and `l2` metrics. Use one of those instead:
```typescript
// Instead of:
d.embedding.similarTo(query, 10, { metric: "inner_product" })
// Use:
d.embedding.similarTo(query, 10, { metric: "cosine" })
```
### Dimension mismatch errors
Ensure query embedding has the same dimension as stored embeddings:
```typescript
const Document = defineNode("Document", {
schema: z.object({
embedding: embedding(1536), // 1536 dimensions
}),
});
// Query embedding must also be 1536 dimensions
const queryEmbedding = await embed(text); // Verify this returns 1536-dim vector
```
### Slow queries
1. **Check index creation**: Vector indexes may not exist
2. **Reduce k**: Smaller k = faster queries
3. **Add filters**: Pre-filter with standard predicates before similarity search
4. **Consider approximate search**: HNSW indexes sacrifice some accuracy for speed
## Hybrid Search: Combining with Fulltext
Vector search excels at semantic similarity but misses exact matches —
proper nouns, SKUs, code identifiers, rare technical terms. **Hybrid
search** fuses vector and fulltext results with Reciprocal Rank Fusion
and typically beats either approach alone.
```typescript
// One query, both signals — fused with RRF at the SQL layer
const results = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext
.matches("renewable energy", 50)
.and(d.embedding.similarTo(queryVec, 50))
)
.select((ctx) => ctx.d)
.limit(10)
.execute();
```
For tunable per-source weights and RRF parameters, use the store-level
`store.search.hybrid()` API. See the [Fulltext Search guide](/fulltext-search)
for the complete hybrid workflow.
## API Reference
See the [Predicates documentation](/queries/predicates#embedding) for
complete API reference of the `similarTo()` predicate and related options.
See [Fulltext Search](/fulltext-search) for the `n.$fulltext.matches()`
predicate and `searchable()` schema brand.